Skip to main content
Upgrade
Loading account

LTX 2.3

LTX 2.3 creates videos from text or one opening image with configurable format, frame rate, and generated audio. Try LTX 2.3 on Nuzza to control the clip setup and review the complete audiovisual result.

Describe what you want to create with LTX 2.3

Loading inspiration gallery.

    LTX 2.3: Plan Picture and Sound on One Shot Timeline

    Editorial LTX 2.3 concept board comparing text-to-video and start-image workflows with shared controls

    Describe the scene, action, camera, mood, dialogue, ambience, and effects as one timed event.

    Editorial storyboard and waveform concept board aligning a camera move with dialogue, footsteps, and ambience

    Plan picture and sound on one shot timeline

    Describe the scene, action, camera, mood, dialogue, ambience, and effects as one timed event. Leave Generate Audio on when sound belongs in the request, or switch it off for a silent clip. Review speech, ambience, effects, timing, and synchronization rather than treating the toggle as a guarantee of usable audio.

    Try For Free
    Editorial endpoint diagram showing one kinetic sculpture moving from a closed start pose to a compatible open end pose

    Define a plausible path between two compositions

    Use one uploaded still to anchor the opening composition. If the destination matters, add an optional End Image URL; it is not a second local upload. Keep subject, scale, viewing axis, lighting, and ratio compatible, describe the connecting motion, and inspect the transition for geometry or identity drift.

    Try For Free
    Editorial format matrix comparing landscape and portrait boards with resolution, duration, and frame-rate tokens

    Set resolution, framing, duration, and frame rate

    Choose 6, 8, or 10 seconds; 1080p, 1440p, or 2160p; and 24, 25, 48, or 50 FPS. Text generation offers landscape or portrait, while image generation adds Auto framing. These controls define the request. Higher resolution or frame rate does not automatically fix motion, text, anatomy, composition, or sound timing.

    Try For Free
    Editorial shot map separating subject, foreground prop, background train, camera rail, lighting, and audio zones

    Organize the scene instead of stacking adjectives

    Write relationships: who is present, where objects sit in depth, what changes, how the camera moves, what motivates the light, and what should be heard. This creates concrete review points. Spatial, fine-detail, text, or sound instructions can still be missed, so adjust one important variable at a time.

    Try For Free

    How to Use LTX 2.3

    01

    Choose Text or Start Image

    Begin from a prompt, or upload exactly one opening image; add an optional End Image URL only when the image workflow needs a destination composition.

    02

    Write and Configure the Shot

    Describe subject, action, camera, light, and sound, then choose duration, resolution, ratio, frame rate, and whether generated audio should be included.

    03

    Generate and Review

    Inspect motion, crop, fine detail, text, endpoint transition, dialogue, effects, ambience, and timing before using the clip or changing one variable for another attempt.

    Best-Fit Workflows for LTX 2.3

    Editorial concept board timing a chef closing a metal oven with a latch click, room tone, and one short spoken line

    Stage one short event with planned sound cues

    Choose LTX 2.3 when a single shot needs visual direction and generated dialogue, ambience, or effects in the same request, because the current workspace combines an audio toggle with selectable resolution, frame rate, framing, and duration. Keep the cue list focused and review whether the picture and sound actually support each other.

    Editorial portrait concept board guiding a paper-fashion figure from a start pose to a compatible ending silhouette

    Guide a vertical motion beat toward an endpoint

    Choose the image workflow when a portrait still defines the opening and a compatible destination image is available by URL. LTX 2.3's current 9:16 or Auto framing and endpoint field make this a specific option for a controlled turn, reveal, or pose transition. Review silhouette, scale, identity, light direction, and the path between frames.

    How to Get More Reviewable LTX 2.3 Results

    Shot and Sound Timeline
    Recommended editorial storyboard with one gesture, one dolly move, and two aligned sound cues

    Use one action, one camera move, and two timed sound cues

    Avoid editorial storyboard overloaded with conflicting action, camera, dialogue, music, and effects cues

    Avoid competing actions, camera moves, and sound events

    Endpoint Compatibility
    Recommended editorial endpoint board preserving one sculpture, viewing axis, scale, framing, and light

    Match subject, scale, axis, ratio, and light across endpoints

    Avoid editorial endpoint board changing the sculpture, viewpoint, scale, environment, and lighting

    Avoid unrelated subjects, viewpoints, environments, and geometry

    Format and Review
    Recommended editorial review board checking crop, detail, text, motion sampling, and waveform timing

    Choose settings for delivery, then inspect the actual result

    Avoid editorial review board ignoring crop, detail, text, motion, and audio issues because settings are high

    Avoid treating 2160p or 50 FPS as a quality guarantee

    LTX 2.3 vs Google Veo 3.1 vs Kling 3.0

    Best For

    LTX 2.3

    Single audiovisual shots where delivery format, portrait framing, frame rate, and optional endpoint guidance matter.

    Google Veo 3.1

    Prompt-led or frame-guided clips that benefit from dedicated endpoint uploads, audio, a negative prompt, or a seed.

    Kling 3.0

    Multi-shot sequences and image-led scenes that need element references or broader prompt-guidance controls.

    Starting Inputs

    LTX 2.3

    Starts from a required prompt or exactly one uploaded start image plus a required prompt.

    Google Veo 3.1

    Starts from text, one uploaded start image, or a dedicated first-and-last-frame workflow.

    Kling 3.0

    Starts from text, one uploaded start image, or uploaded first and last frames; prompts are optional in current definitions.

    Duration

    LTX 2.3

    Offers 6-, 8-, and 10-second choices.

    Google Veo 3.1

    Offers 4-, 6-, and 8-second choices.

    Kling 3.0

    Offers 3 through 15 seconds.

    Resolution and Ratio

    LTX 2.3

    Offers 1080p, 1440p, or 2160p and landscape, portrait, or Auto framing where applicable.

    Google Veo 3.1

    Offers 720p, 1080p, or 4K and landscape, portrait, or Auto framing where applicable.

    Kling 3.0

    Offers 16:9, 9:16, or 1:1 for text generation; the current definition does not expose a resolution selector.

    Frame Rate and Advanced Controls

    LTX 2.3

    Offers 24, 25, 48, or 50 FPS; no seed, negative-prompt, multi-shot, or element controls are exposed.

    Google Veo 3.1

    Does not expose FPS selection, but provides negative-prompt and seed controls.

    Kling 3.0

    Provides multi-prompt, shot type, negative prompt, guidance scale, and element controls; no FPS selector is exposed.

    Audio

    LTX 2.3

    Generate Audio is available and on by default, with results requiring review for content and timing.

    Google Veo 3.1

    Generate Audio is available and on by default in the current workflows.

    Kling 3.0

    Generate Audio is available and on by default in the current workflows.

    Endpoint Handling

    LTX 2.3

    Image to Video accepts one uploaded start frame and an optional End Image URL.

    Google Veo 3.1

    Provides a dedicated workflow with uploaded first and last images.

    Kling 3.0

    Provides uploaded first-and-last-frame generation and an End Image URL in Image to Video.

    Choose When

    LTX 2.3

    Choose it when selectable FPS, 1440p or 2160p, generated audio, and a focused single-shot workflow fit the brief.

    Google Veo 3.1

    Choose it when a dedicated endpoint-upload workflow, negative prompt, seed, or 4K option matters more than FPS control.

    Kling 3.0

    Choose it when longer duration, multiple shots, square framing, or element-guided scenes matter more than explicit FPS and resolution choices.

    Choose LTX 2.3 for a focused audiovisual request with visible FPS and 1080p-to-2160p choices. Choose Google Veo 3.1 for a dedicated first/last-frame workflow, negative prompt, seed, or 4K option, and Kling 3.0 for longer multi-shot or element-guided direction. Nuzza keeps these current models and their visible controls in one composer, but it does not guarantee that one model will always produce the best result.

    FAQs about LTX 2.3

    What is LTX 2.3?

    LTX 2.3 is Nuzza's hosted presentation of the LTX-2.3 video model family from Lightricks. It creates short video from text or one start image and can request generated audio. Official family positioning includes improved detail, prompt interpretation, image-to-video motion, and audio, but each result remains probabilistic and needs review.

    What is the best AI video generator?

    The best AI video generator depends on the input, intended result, budget, and controls you need. LTX 2.3's Nuzza workflow is a strong choice because it is built around a focused text and start-image audiovisual workflow with explicit resolution and frame-rate choices.

    Is LTX 2.3 free to use?

    The page is free to browse. Running a task requires Credits, and the current cost is shown before you submit.

    What can I use LTX 2.3 for?

    LTX 2.3 is useful for audiovisual scene and portrait transition. Choose LTX 2.3 when a single shot needs visual direction and generated dialogue, ambience, or effects in the same request, because the current workspace combines an audio toggle with selectable resolution, frame rate, framing, and duration. Keep the cue list focused and review whether the picture and sound actually support each other.

    What inputs does LTX 2.3 support?

    Yes. Text to Video requires a prompt. Image to Video requires one uploaded start image and a prompt, and it can optionally use an End Image URL. The current Nuzza entry does not expose video input, audio input, multiple reference images, local model weights, or editing endpoints such as retake, extend, and reframe.

    Do I need to sign in to use LTX 2.3?

    You can browse this page without signing in. Submitting a generation, storing persistent inputs, checking task status, and saving a finalized result require a Nuzza account.

    What should I check in the video result?

    Watch the complete result and check subject identity, motion, camera continuity, frame edges, text, and audio synchronization. Compare important details with the source or prompt before publishing.

    What duration, resolution, framing, and FPS choices are available?

    The current workspace offers 6, 8, or 10 seconds; 1080p, 1440p, or 2160p; and 24, 25, 48, or 50 FPS. Text to Video offers 16:9 and 9:16. Image to Video adds Auto framing. These settings configure the request but do not guarantee detail, motion quality, text accuracy, or synchronization.

    Does LTX 2.3 generate audio?

    Yes. Generate Audio is available in both current workflows and is on by default. Describe the dialogue, ambience, music direction, or effects that belong in the shot, then review what was produced for content, timing, intelligibility, and synchronization. Turn the option off when you want a silent clip for a separate audio workflow.