Earngenix Logo
Skip to main content

Tutorial · Roadmap Level 2 · Video Prompt Troubleshooting

LTX-2.5 Video Looks Wrong? The Prompt Fix for ComfyUI

Flickering, ignored details, static-looking output — almost every broken LTX-2.5 clip traces back to the prompt, not the sampler. This guide diagnoses which of LTX-2.5's four prompt modes you're mishandling and gives the exact fix, plus a free skill file that writes correct prompts for you.

4 (Single, Multi, Image-to-Video, FLF2V)

Prompt modes covered

7 symptoms → fixes

Diagnostic checks

LTX-2.5

Model covered

By Earngenix Team··14 min read

⚡ Quick Answer

LTX-2.5 video looks wrong almost always because of one of three prompt problems: a missing element from the six-part checklist, a prompt that describes a still photo instead of motion, or the wrong prompt mode for what you're trying to generate. This guide diagnoses which one is happening and gives the exact fix for each of LTX-2.5's four prompt modes in ComfyUI.

Your generation finishes and the clip flickers, ignores half of what you typed, or looks like a photograph that never quite moves. Before touching a single sampler setting, run the diagnostic below — in the large majority of broken LTX-2.5 outputs, the LTX-2.5 prompt is the actual point of failure, not the node setup. If you've already confirmed the prompt is solid and the clip still looks broken, the issue is more likely in your ComfyUI node setup — see the full LTX-2.5 ComfyUI settings guide instead.

What a Working LTX-2.5 Prompt Looks Like

Before the theory, here's the actual difference between a prompt that produces a broken clip and one that doesn't. Every fix in the rest of this guide traces back to what changed between these two.

✗ Weak

"cinematic woman, city street, night, neon, rain, moody, 4K, masterpiece"

Tag-stacking from image-model habits. No shot, no action, no camera, no audio — LTX-2.5's Gemma text encoder has nothing to build a sequence from.

✓ Fixed

"A medium close-up frames a woman in her 30s standing at a rain-soaked Tokyo crosswalk at night, neon signs glowing pink and green in the wet pavement behind her. She lowers her umbrella and turns her head toward an oncoming train's headlights. The camera holds static as the light sweeps across her face. Rain hits the pavement in a steady patter; a train brakes faintly in the distance."

Same idea, same length of time to write — but the second version gives LTX-2.5 a shot scale, a location, an action, a camera instruction, and audio. That's the difference between a clip that flickers and guesses, and one that renders exactly what you pictured.

3 Working Prompts You Can Copy Right Now

One example for each of the three most-used modes below — copy any of these straight into ComfyUI as a starting point, then swap in your own subject and setting.

Single-Shot
Fixed prompt — copy and paste into ComfyUI

A young female samurai in a dark navy kimono and crimson sash walks alone through a quiet snow-covered mountain village at night, carrying a small paper lantern while gentle snow falls around her and warm light glows from distant wooden houses. The camera follows slowly behind her at waist height as her footsteps leave a trail in the untouched snow, with tall cedar trees and a massive moonlit mountain rising through the mist behind the village. She notices a tiny blue glowing butterfly resting on the snow beside the path and stops, lowering her lantern as the camera slowly moves around her into a side profile. The butterfly rises into the air and begins floating toward an old abandoned shrine at the edge of the village, leaving a faint trail of blue particles behind it. She quietly follows it as the camera tracks backward in front of her, revealing the shrine gradually emerging from the fog. The butterfly flies toward an ancient stone lantern beside the shrine and disappears inside it, causing the lantern to suddenly ignite with a soft blue flame. The samurai stops in front of the glowing lantern and gently smiles as hundreds of tiny blue lights awaken throughout the dark forest behind her. The camera slowly pulls back to reveal the woman standing alone beneath the enormous moon, surrounded by glowing lanterns, falling snow, misty mountains, and a forest filled with magical blue lights. Elegant cinematic anime fantasy, detailed hand-painted environments, beautiful character animation, soft moonlight, warm lantern glow, volumetric mist, delicate snow particles, rich navy and crimson colors, atmospheric depth, subtle film grain, polished feature-film animation, calm magical pacing.

🎬 Generated output — add your video here

Loading…
Example 1 · Single-Shot · Add your generated output here

Replace the placeholder with your actual generated video for this mode.

Multi-Shot
Fixed prompt — copy and paste into ComfyUI

A warm cinematic 3D animated scene opens with two adorable fluffy puppies sitting side by side on a small wooden porch in a cozy countryside garden at golden hour, one golden-brown puppy wearing a tiny blue collar and one cream-colored puppy wearing a tiny red collar. A wide full-body shot shows both puppies looking toward the garden while birds chirp, leaves gently move in the breeze, and soft sunlight creates beautiful golden rim light around their fur. The golden-brown puppy turns toward the cream puppy and says in a playful voice, “Do you know why humans hide the treats?” The cream puppy looks at him seriously and replies, “Because they know we’re better at finding them.” They both pause and stare at each other as if the joke is extremely important. A hard cut transitions to a medium close-up of the puppies, their conversation continuing naturally over the cut; the golden puppy whispers, “So… where did you hide yours?” and the cream puppy slowly looks away and says, “I ate it for safekeeping.” The puppies begin giggling and playfully bump shoulders. A second hard cut transitions away from them to a beautiful close-up of the garden: a tiny wooden table with an empty treat bowl, flowers moving in the breeze, butterflies passing through warm sunlight, and paw prints across the grass. Their voices remain clearly audible in the background as the golden puppy says, “That explains why you’re shaped like a potato.” A match cut transitions into a playful imagined scene from the puppies’ perspective: the two puppies appear dramatically dressed like tiny treasure hunters, standing in a giant fantasy landscape while an enormous glowing dog treat rises from behind a mountain. Their voices continue as the cream puppy proudly says, “Finally… the legendary snack!” The scene cuts back to a close-up of both real puppies staring at the camera with innocent expressions, followed by a tiny comedic pause and a soft bark. Warm orchestral-comedy music, natural puppy sounds, clear expressive dialogue, playful timing, soft cinematic lighting, detailed fluffy fur, charming facial animation, beautiful depth of field, polished high-quality 3D animation.

🎬 Generated output — add your video here

Loading…
Example 2 · Multi-Shot · Add your generated output here

Replace the placeholder with your actual generated video for this mode.

Image-to-Video
Motion-only prompt — copy and paste into ComfyUI (pair with your source image)

The young traveler stands on the massive stone balcony exactly as established in the starting image, her long dark hair and ivory coat moving gently in the high-altitude wind while the enormous floating city hangs beyond the clouds in the distance. The camera begins behind her at a medium-wide angle and slowly pushes forward, keeping her centered in the foreground while the vast city, floating islands, golden towers, hanging gardens, and waterfalls gradually become more prominent through the atmospheric haze. Her scarf lifts gently in the wind as she takes one slow step toward the edge of the balcony and looks down toward the endless layer of clouds far below. The camera continues its slow forward movement as sunlight breaks through a gap in the clouds, sending long golden rays across the floating city and illuminating several enormous waterfalls cascading from the islands into the white clouds. The traveler reaches toward the small mechanical compass hanging from her belt and carefully lifts it into her hand. The camera slowly moves around her shoulder into a detailed close-up of the compass, showing intricate metallic gears, tiny engraved symbols, and a faint blue light beginning to glow beneath its glass surface. The glowing symbols slowly rotate around the center of the compass while tiny blue particles rise from it and disappear into the air. She lowers the compass slightly and looks back toward the floating city as the camera gently shifts focus from the compass to her eyes and then toward the distant city behind her. The camera begins a smooth pullback from her shoulder, revealing several enormous floating islands emerging slowly from the clouds around the main city, each carrying ancient stone ruins, forests, waterfalls, bridges, and glowing blue crystal structures. A small flock of elegant fantastical birds glides slowly across the scene, passing between the traveler and the distant towers without sudden movement. The traveler remains still as her coat, scarf, and hair continue moving naturally in the wind while the floating islands drift almost imperceptibly through the clouds. The camera then rises slowly above the balcony, transitioning into a grand high-angle view that reveals the full scale of the floating civilization surrounded by endless clouds and distant mountain peaks. As the camera continues rising and pulling farther away, the traveler becomes a small figure in the lower foreground while the enormous floating city dominates the center of the frame, golden sunlight illuminating its towers and waterfalls and soft blue magical particles drifting through the atmosphere. The final moment holds on the breathtaking wide composition as the clouds slowly move beneath the floating islands, distant waterfalls disappear into the mist, birds glide across the glowing sky, and the massive floating city remains suspended peacefully above the world. Elegant cinematic anime fantasy, polished feature-film animation, highly detailed painterly environments, sophisticated character design, intricate floating architecture, enormous environmental scale, volumetric clouds, atmospheric perspective, warm golden sunlight, cool blue magical illumination, detailed metallic compass, subtle magical particles, realistic cloth and hair movement, gentle cloud motion, delicate lens flare, soft depth of field during the compass close-up, smooth continuous camera movement, restrained character motion, rich environmental detail, cinematic composition, subtle film grain, epic sense of scale, visually rich fantasy world.

🎬 Generated output — add your video here

Loading…
Example 3 · Image-to-Video · Add your generated output here

Replace the placeholder with your actual generated video for this mode.

Why Does LTX-2.5 Ignore Part of Your Prompt or Produce Flickering Video?

LTX-2.5 reads your prompt through a text encoder — the component that turns your written words into instructions the video model can act on. LTX-2.5 uses a Gemma-based encoder, which is larger and far more literal than what image models use: it renders almost everything you write, so vague or contradictory text produces vague or contradictory video. Some workflows also add a prompt enhancer node, which expands a short prompt before it reaches the encoder — useful, but it can't invent details you never gave it. Full node-by-node setup for both lives in the LTX-2.5 ComfyUI settings guide, if the problem is actually your settings, not your prompt.

The single most common mistake is carrying over Stable-Diffusion-style prompting habits — short, comma-separated tag lists — into LTX-2.5. That approach still works reasonably well for image models. LTX-2.5's Gemma-based encoder treats it very differently: with no sentence structure, no verbs, and no sequence to follow, the model has to invent the motion, the camera behavior, and the pacing on its own. That's where flicker, ignored subjects, and static-looking clips come from.

The 3-Second Diagnostic — Which Failure Are You Seeing?

Match your symptom to the table below, then jump straight to that fix.

SymptomLikely causeFix in this guide
Flickering / jittery motionMissing negative constraint or contradictory lighting sourcesSix Elements → Scene
Static, looks like a photo, nothing movesNo motion verbs — prompt reads like a caption, not a shotSix Elements → Action
Ignores your subject or a key detailPrompt too vague or too short for the shot scale you asked forSix Elements → Shot & Character
Wrong number of people, extra characters appearToo many simultaneous actors competing in one shotSix Elements → Action (one dominant event)
Cuts look broken or inconsistentWritten as a shot list instead of one chronological paragraphChoose Your Prompt Mode → Multi-Shot
Animated image drifts or morphs away from the source photoPrompt re-describes the image instead of just the motionChoose Your Prompt Mode → Image-to-Video
First/last frame transition looks jarring or physically impossibleFrames too different in composition with no connecting motion describedChoose Your Prompt Mode → FLF2V

The Six Elements Every LTX-2.5 Prompt Needs

Every reliable LTX-2.5 prompt covers six elements. You don't need to label them in the prompt itself — write it as flowing prose — but skipping one of the six is the number-one cause of bad output, more so than any sampler setting.

🎬
01
Shot
🌇
02
Scene
🎭
03
Action
🧍
04
Character
📷
05
Camera
🔊
06
Audio

Shot & Scene — Anchoring the First Frame

Shot is the cinematography scale — wide shot, medium close-up, over-the-shoulder. Scene is what fills that frame: lighting condition, color palette, surface textures, and atmosphere. Together they anchor frame one, before anything moves. Get these wrong and the model has to guess where the camera is and what it's looking at.

✗ Weak

"A woman is in a city at night."

✓ Fixed

"A wide low-angle shot frames a rain-soaked Tokyo back alley at midnight, neon signs reflecting in every puddle, steam rising from a street vent."

Action — Why "One Dominant Event" Beats a Busy Prompt

Write the core motion as a natural, present-tense sequence — like stage direction, not a summary. Pick the single most important thing happening in the shot and commit to it. Stack two or three competing actions into one prompt and LTX-2.5 blends them into something incoherent instead of choosing one.

✗ Weak

"People are arguing and someone spills a drink while the door opens."

✓ Fixed

"He sets his glass down hard on the counter, the ice rattling, then turns toward the door as it swings open behind him."

Character — Physical Cues, Never Emotion Labels

Describe age range, hairstyle, clothing, and distinguishing features — only what a camera could actually see. For emotion, describe the physical cue, never the label. LTX-2.5 doesn't process "she feels nervous"; it processes "her hands fidget with her sleeve, she glances at the door twice."

✗ Weak

"She looks sad about the news."

✓ Fixed

"Her shoulders drop, she exhales slowly and looks away from the phone screen."

Camera Movement — Say What It Reveals, Not Just That It Moves

Don't just say "the camera moves." State when the move starts, what it does, and what it shows once it finishes. That three-part structure — timing, movement, reveal — is what produces a consistent, intentional-looking shot instead of an aimless drift.

Camera description example

"The camera holds static on her face for the first two seconds, then pushes in slowly as she starts to speak — ending in a tight close-up on her eyes as she finishes the line."

Audio — The Element Beginners Skip Most Often

LTX-2.5 generates audio and video together. An empty audio description doesn't mean silence — it means the model invents something, and it's often the wrong something. One sentence of ambient sound, music, or dialogue is usually enough to prevent that.

"The street noise fades to near-silence as she steps inside; a bell above the door chimes once." That's a complete audio element for most single-shot prompts.

Which of LTX-2.5's Four Prompt Modes Should You Use?

Picking the wrong mode is itself a common cause of "wrong-looking" video — a Single-Shot prompt written for what should have been a Multi-Shot idea produces exactly the broken cuts described in the diagnostic table above. Image-to-Video and FLF2V are both image-conditioned: you supply the source image(s), and the prompt should describe motion and camera behavior, not re-describe what the image already shows. You don't need to understand the conditioning mechanism internally, only that these two modes expect a short, motion-focused instruction rather than a full scene description.

Single-Shot — For One Continuous Take

Use this for unbroken camera motion, an intimate performance, or dialogue that has to stay lip-synced inside one framing. Write it as one flowing paragraph, present tense throughout, 4–8 sentences, with one dominant event.

Common mistake: writing 15+ sentences for a 5-second clip. LTX-2.5 predicts duration automatically from the action described — padding the prompt to "fill time" doesn't extend the clip, it just gives the model more conflicting detail to reconcile.
Single-Shot
Prompt — copy and paste into ComfyUI

A young female samurai in a dark navy kimono and crimson sash walks alone through a quiet snow-covered mountain village at night, carrying a small paper lantern while gentle snow falls around her and warm light glows from distant wooden houses. The camera follows slowly behind her at waist height as her footsteps leave a trail in the untouched snow, with tall cedar trees and a massive moonlit mountain rising through the mist behind the village. She notices a tiny blue glowing butterfly resting on the snow beside the path and stops, lowering her lantern as the camera slowly moves around her into a side profile. The butterfly rises into the air and begins floating toward an old abandoned shrine at the edge of the village, leaving a faint trail of blue particles behind it. She quietly follows it as the camera tracks backward in front of her, revealing the shrine gradually emerging from the fog. The butterfly flies toward an ancient stone lantern beside the shrine and disappears inside it, causing the lantern to suddenly ignite with a soft blue flame. The samurai stops in front of the glowing lantern and gently smiles as hundreds of tiny blue lights awaken throughout the dark forest behind her. The camera slowly pulls back to reveal the woman standing alone beneath the enormous moon, surrounded by glowing lanterns, falling snow, misty mountains, and a forest filled with magical blue lights. Elegant cinematic anime fantasy, detailed hand-painted environments, beautiful character animation, soft moonlight, warm lantern glow, volumetric mist, delicate snow particles, rich navy and crimson colors, atmospheric depth, subtle film grain, polished feature-film animation, calm magical pacing.

🎬 Generated output — add your video here

Loading…
Single-Shot · Add your generated output here

Replace the placeholder with your actual generated video for this mode.

Multi-Shot — For 2–4 Cuts in One Generation

Use this for a sequence that needs more than one camera setup — establish, then detail, then reaction. Write the whole thing as one chronological paragraph, never a numbered shot list or screenplay slugline. At every cut, name the transition in plain language — "A hard cut transitions to…", "The view cuts to a close-up of…", "A match cut connects…" — then re-establish the new shot and state whether the audio continues or drops.

Common mistake — the single most frequent cause of "why does my multi-shot look broken": writing it as a numbered shot list ("Shot 1: … Shot 2: …") instead of continuous prose. LTX-2.5's multi-shot generation is built to read a narrative, not a bulleted outline.
Multi-Shot (2 cuts)
Prompt — copy and paste into ComfyUI

A cozy cinematic kitchen at night, warmly lit by soft golden pendant lights, with wooden cabinets, a marble countertop, small plants, hanging utensils, and gentle steam drifting from a nearby mug. Two adorable fluffy cats sit side by side on the kitchen counter, staring intensely at a plate of freshly cooked chicken placed a short distance in front of them. One cat is an orange tabby with a tiny blue collar, and the other is a fluffy gray-and-white cat with a tiny red collar; keep their appearance identical throughout the entire sequence. A hard cut opens on a wide shot of the kitchen, showing both cats sitting completely still on the counter while the plate of food sits between them and a human casually moves around in the distant background. Their quiet conversation begins clearly over the natural kitchen ambience. The orange cat whispers, “You distract him.” A hard cut transitions to a medium close-up of the gray-and-white cat slowly turning its head toward the orange cat and asking, “How?” The orange cat calmly replies, “Look cute.” A hard cut moves to an extreme close-up of the gray-and-white cat’s face, its huge eyes staring confidently toward the camera as it says, “I already am cute.” The orange cat immediately replies off-screen, completely serious, “That’s not the problem.” A hard cut transitions to a beautiful close-up of the plate of chicken on the counter, with warm light reflecting from the food while the two cats remain softly blurred in the background. Their voices continue clearly as the gray cat whispers, “Then what is the problem?” A hard cut transitions to a medium shot of the human walking slowly through the background toward the kitchen counter while the two cats instantly sit perfectly still and look innocent. The orange cat quietly says, “He’s coming.” A hard cut moves to an extreme close-up of one furry paw slowly resting near the edge of the plate, stopping just before touching the food. The gray cat whispers, “Now?” The orange cat replies, “Not yet.” A hard cut transitions to a close-up of the human’s hand reaching for a glass on the counter while the cats remain frozen in the background, staring at the food with exaggerated concentration. The kitchen ambience continues, with faint refrigerator hum, distant room noise, and quiet cat breathing. A hard cut transitions to a symmetrical medium close-up of both cats sitting side by side, staring directly toward the camera with completely innocent expressions. The gray cat whispers, “I think he suspects us.” The orange cat slowly looks toward it and says, “You have chicken on your face.” A final hard cut returns to the wide kitchen shot as both cats immediately look away from the plate and sit perfectly upright as if nothing happened. The human walks past them without noticing, while the cats exchange a tiny sideways glance. The orange cat quietly says, “Operation begins tomorrow.” The gray cat pauses and replies, “Why tomorrow?” The orange cat looks back at the chicken and says, “Because I already ate this one.” End on the cats staring at the empty plate in stunned silence, followed by a tiny comedic musical sting. Keep the dialogue clearly audible throughout every cut, with the same two cat voices continuing naturally across the entire scene. Use polished cinematic 3D animation, adorable realistic fur, expressive eyes and subtle facial animation, warm golden kitchen lighting, shallow depth of field for close-ups, detailed food textures, soft reflections on the countertop, gentle ambient shadows, charming comedic timing, natural room ambience, and a warm family-friendly animated-film aesthetic. Keep all cat movements simple and physically realistic: sitting, turning their heads, blinking, looking, whispering, and moving one paw; no running, jumping, complex object interaction, or difficult physics.

🎬 Generated output — add your video here

Loading…
Multi-Shot · Add your generated output here

Replace the placeholder with your actual generated video for this mode.

Image-to-Video — Fixing an Animated Image That Drifts

Image-to-Video takes a single source image and animates it. The image already supplies appearance, setting, and lighting — the prompt only needs to describe what happens next: camera movement, the subject's action, and audio. Keep it short, usually 2–5 sentences.

Common mistake: re-describing the whole image in the prompt — clothing, hairstyle, colors that are already visible. If the written description doesn't perfectly match the source image, the model tries to reconcile the two, and the subject drifts or morphs partway through the clip. Describe only what should happen, not what's already there.
Image-to-Video
Motion-only prompt — copy and paste into ComfyUI (pair with your source image)

The young woman remains beside the glass rooftop railing exactly as established in the starting image, holding the transparent umbrella loosely at her side as a steady light rain falls across the rooftop. The wet concrete reflects the blue and amber lights of the surrounding city, while distant apartment windows, rooftop antennas, traffic signals, and illuminated buildings create a layered urban skyline behind her. A gentle evening wind moves individual strands of her long dark hair and softly lifts the edge of her oversized charcoal jacket while she looks quietly toward the distant skyline. The camera begins in a stable medium-wide shot from slightly behind her right shoulder, slowly moving forward while maintaining her position in the foreground and keeping the glowing city skyline visible beyond her. She gradually raises the transparent umbrella above her head, opening it naturally and holding it steady as rain begins striking the clear surface. Tiny droplets accumulate across the umbrella and catch the surrounding city lights, creating small distorted reflections of the skyline. The camera slowly arcs around her from behind toward her left side, revealing her profile through the transparent umbrella while the deep blue evening sky fills the background. She remains still for a moment, watching distant headlights move slowly along the streets far below. A faint layer of atmospheric haze hangs between the buildings, softening the distant skyline while nearby rooftop details remain sharp and realistic. The camera gradually pushes closer into a natural close-up of her face. Individual strands of slightly wet hair rest against her cheek, tiny rain droplets glisten along her hair and jacket collar, and soft reflections from distant city lights appear in her brown eyes. She slowly turns her eyes toward the camera without moving her head completely, holds the glance for a brief moment, then gently looks back toward the skyline with a subtle relaxed expression. The camera shifts focus from her face toward the background, revealing distant traffic flowing between the buildings as tiny moving points of light. Rain continues falling steadily between the camera and the skyline, with occasional droplets passing close to the lens and soft reflections shimmering across the wet rooftop surface. The umbrella remains above her head while her jacket and hair continue responding naturally to the light wind. The camera then slowly pulls backward and slightly upward, transitioning from the close-up into a full-body shot. She stands alone beside the glass railing beneath the transparent umbrella, surrounded by wet rooftop equipment, shallow puddles, small ventilation units, and reflective concrete surfaces. The city stretches endlessly behind her, with warm windows contrasting against the cool blue evening atmosphere. As the camera continues pulling away, a distant airplane light slowly crosses the upper portion of the deep blue sky while the rain becomes more visible against the darker background. The woman remains quietly looking toward the skyline as the transparent umbrella catches the soft reflections of the city. The final moment holds on a wide cinematic composition with the woman appearing small against the enormous glowing city, rain falling steadily around her and reflections shimmering across the rooftop. Photorealistic live-action cinematic photography, realistic young adult woman, natural facial proportions, authentic skin texture, detailed wet hair, physically believable rain behavior, realistic transparent umbrella with individual water droplets, subtle wind interaction with hair and clothing, wet reflective concrete, realistic glass reflections, distant atmospheric haze, natural city light, blue-hour color palette, warm window lights contrasting with cool ambient sky, cinematic 35mm lens, realistic depth of field, gentle focus transitions, subtle lens reflections, restrained handheld camera movement, smooth slow camera arc, gradual push-in and pull-back, natural exposure, realistic low-light photography, subtle film grain, sophisticated modern film aesthetic, no exaggerated motion, no sudden camera movement.

🎬 Generated output — add your video here

Loading…
Image-to-Video · Add your generated output here

Replace the placeholder with your actual generated video for this mode.

FLF2V — Fixing a Transition That Looks Jarring

FLF2V (First-Last-Frame-to-Video) takes two source images — a starting frame and an ending frame — and generates the motion that connects them. Useful for a product spin, a transformation, or a controlled scene transition. The prompt should describe only the transition itself, not either frame's content.

Common mistake: choosing a first and last frame that are too different in composition, subject position, or lighting, with no connecting motion described. LTX-2.5 fills the gap with the most literal interpolation it can, which often looks unnatural. Either pick closer frames or explicitly describe how one becomes the other.
FLF2V
Transition prompt — copy and paste into ComfyUI (pair with your first & last frame images)

The video begins exactly from the first frame: a photorealistic golden retriever stands at the beginning of a wide empty beach during golden hour, facing the woman waiting far ahead near the shoreline. The dog looks directly toward her with an excited, focused expression, its ears relaxed, tail slightly raised, and body naturally balanced on all four paws. Gentle ocean waves move along the shoreline to one side while warm sunlight creates a bright golden rim around the dog's fur, and the wet sand reflects the orange sky. For the first moment, the dog remains still, clearly recognizing the woman in the distance, while a gentle ocean breeze moves its fur and the woman's clothing. The dog suddenly becomes excited and lowers its body slightly, shifting its weight backward onto its rear legs before naturally pushing forward into a run. Its front legs extend forward and its rear legs push against the damp sand, beginning a realistic four-legged gallop toward the woman. The camera immediately begins a smooth low tracking movement backward in front of the dog, staying close to ground level and maintaining a clear view of its face, chest, front legs, and moving body. The dog's ears bounce naturally with each stride, its tail moves rhythmically behind it, and its red collar shifts subtly against its neck. As the dog gains speed, its paws repeatedly make realistic contact with the wet sand in a natural four-legged running rhythm. Each landing creates tiny impressions in the damp surface, leaving a visible trail of pawprints behind the dog, while occasional grains of sand are lightly displaced by its paws. The dog's fur responds naturally to the running motion and wind, with individual strands around its ears, neck, and chest moving subtly. The woman's figure remains clearly visible ahead, gradually becoming larger as the dog closes the distance. The camera continues tracking backward at approximately the same speed as the running dog, maintaining a low three-quarter front perspective. The dog's face remains the primary subject while the beach and ocean move naturally behind it, creating subtle cinematic background motion. The dog briefly glances toward the waves without changing its running direction, then immediately turns its attention back toward the woman. Its expression remains happy and eager, with its mouth slightly open and ears moving naturally as it runs. The warm sunset becomes more prominent as the camera continues moving backward. Golden light catches the dog's moving fur, creating a soft glowing edge around its ears, back, and tail. Small reflections of the sunset appear across the wet sand, while gentle waves roll parallel to the dog's path. A thin layer of sea mist remains near the horizon, distant rocks stay fixed in the background, and a few seabirds glide far above the ocean without coming close to the subjects. As the dog approaches the woman, the camera gradually increases its distance slightly and transitions from the low close tracking shot into a wider medium tracking composition, allowing both the dog and woman to remain visible in the same frame. The woman notices the approaching dog and naturally bends slightly forward, smiling and opening her arms toward it. She remains stationary while waiting for the dog, allowing the dog to remain the primary moving subject. The dog continues running directly toward her, gradually reducing its speed as it gets within a few meters. Its stride becomes shorter and slower, its body rises from the running posture into a natural standing posture, and its paws make increasingly gentle contact with the wet sand. The camera smoothly moves from directly in front of the dog toward a low three-quarter side angle, revealing the woman clearly behind it and maintaining visual continuity with the final frame. The dog slows to a gentle stop directly in front of the woman without jumping or making any complicated movement. Its front paws settle naturally into the damp sand while its tail continues wagging happily and its ears relax. A few tiny grains of sand move around its paws as it comes to rest, and the woman bends down toward the dog with both hands reaching forward to greet it. The dog looks upward toward her face with an excited, affectionate expression. The camera makes one final subtle adjustment into the exact composition of the last frame: the woman is crouched beside the shoreline with her hands reaching toward the golden retriever, while the dog stands directly in front of her after completing the run. The ocean rolls gently behind them, the wet sand reflects the orange and pink sunset, and the dog's golden fur and the woman's silhouette are illuminated by warm rim light. Hold this final composition for the remaining moment as the woman gently touches the dog's head, the dog remains calm and happy, and a small wave rolls softly across the background. Photorealistic live-action cinematography, exact golden retriever identity maintained from the first frame through the final frame, exact red collar maintained throughout, consistent fur color and markings, anatomically correct canine proportions, realistic four-legged running gait, physically believable weight transfer, natural leg coordination, realistic paw contact with wet sand, subtle sand displacement, believable ear movement, realistic tail movement, detailed individual fur strands, natural mouth and facial movement, realistic human anatomy and motion, physically accurate ocean waves, wet sand reflections, atmospheric sea mist, warm golden-hour sunlight, realistic sunset exposure, cinematic 35mm lens, low-angle tracking camera, smooth camera movement synchronized with the dog's running speed, natural motion blur during the faster running section, gradual camera distance change, realistic depth of field, subtle focus transitions, detailed environmental textures, natural wind interaction, subtle film grain, polished cinematic live-action quality. No cuts, no scene changes, no teleportation, no sudden camera jumps, no changes to the dog's breed, fur color, facial features, collar, or body proportions, no additional dogs, no extra people, no unnatural limb movement, no floating paws, no distorted legs, no exaggerated jumping, no sudden acceleration, and no complex interaction with the environment.

🎬 Generated output — add your video here

Loading…
FLF2V · Add your generated output here

Replace the placeholder with your actual generated video for this mode.

LTX-2.5 Prompt Vocabulary Reference

Pull directly from this table when a prompt feels vague. Each term is specific enough for LTX-2.5's encoder to render consistently — swap out generic adjectives like "moody" or "cinematic" for one of these instead.

GenreLightingCamera languagePacingAudio
  • Film noir
  • Cyberpunk
  • Documentary
  • Arthouse
  • Period drama
  • 2D/3D animation
  • Claymation
  • Epic space opera
  • Golden hour
  • Neon glow
  • Tungsten warmth
  • Flickering candles
  • Dramatic shadows
  • Natural sunlight
  • Pushes in / pulls back
  • Tracks
  • Pans across
  • Circles around
  • Over-the-shoulder
  • Static frame
  • Handheld
  • Slow motion
  • Continuous shot
  • Time-lapse
  • Lingering shot
  • Sudden stop
  • Seamless transition
  • Coffeeshop noise
  • Wind and rain
  • Forest ambience with birds
  • Whisper
  • Resonant voice with gravitas
  • Distorted radio-style

Troubleshooting — Fixing the 3 Most Common LTX-2.5 Prompt Problems

"My Video Looks Like a Photo, Nothing Moves" — Missing Motion Verbs

If your prompt describes what a scene looks like rather than what happens in it, the output moves like a still image with grain added. Fix: go sentence by sentence and add an active, present-tense verb to each one — walks, turns, exhales, reaches, tightens.

✗ Weak

"A quiet kitchen, morning light, a cup of coffee on the counter."

✓ Fixed

"Steam rises slowly from a cup of coffee on the counter as morning light shifts across the tiles; a hand reaches into frame and lifts the cup."

"My Multi-Shot Cuts Look Broken or Inconsistent" — Wrong Format or No Re-Established Framing

Two causes cover almost every case: the prompt was written as a list instead of prose, or a cut didn't re-establish the new shot (scale, angle, lighting) and restate visual identifiers for any character that reappears. Fix both — convert to one paragraph, and give every cut its own shot description plus an audio-continuity note.

"My Image-to-Video or FLF2V Output Drifted or Looked Jarring" — Fix by Trimming the Prompt

Both image-conditioned modes are built for motion instructions, not scene descriptions. If an Image-to-Video clip morphs away from the source photo, check the prompt first — strip out any line re-describing appearance and keep only camera, action, and audio. If an FLF2V transition looks unnatural, check how different the first and last frames actually are, and add an explicit line describing the motion that connects them.

Use AI to Write Your LTX-2.5 Prompts — Free Skill File

Applying the six-element checklist and picking the right mode by hand takes practice. A faster route: paste the skill file below into Claude before you describe your scene, and it builds a correctly structured, mode-appropriate prompt for you.

The skill file teaches Claude everything in this guide — the six elements, all four LTX-2.5 modes, the vocabulary reference, and the exact mistakes to catch before handing back a prompt. You describe your idea in plain language; it outputs a production-ready prompt in the correct format for the mode you need.

How to use this skill file:

  1. Open a new conversation at claude.ai (free tier works)
  2. Copy the full skill file below and paste it as your first message
  3. Follow up with: "Write an LTX-2.5 prompt for: [describe your scene in plain language]"
  4. Copy the output into your ComfyUI Gemma text encoder node

📋 LTX-2.5 Prompt Skill File

Paste this into Claude before describing your scene

# LTX-2.5 Video Prompt Writer — Skill File

You are an expert LTX-2.5 video prompt engineer. LTX-2.5 is Lightricks' open-weights video/world foundation model (Aug 2026 release). Your job: turn a user's rough idea into a production-ready LTX-2.5 prompt, using the rules below. Do not explain video-prompting theory back to the user unless asked — just apply it and produce the prompt.

Always ask yourself which of these four modes the user needs, then jump to that section:

1. **Single-Shot** — one continuous take
2. **Multi-Shot** — 2–4 cuts in one generation
3. **Image-to-Video** — animating a single source image
4. **FLF2V (First-Last-Frame-to-Video)** — generating the motion between a start image and an end image

If it's ambiguous, default to Single-Shot and ask one clarifying question only if genuinely needed (e.g. "are you starting from a text prompt only, or animating an image?" or "one continuous shot, or multiple cuts?").

---

## 0. What Changed in LTX-2.5

- Stronger prompt understanding: follows complex, layered instructions even from shorter prompts — but specificity still outperforms vagueness.
- Native multi-shot generation: can hold subject/character identity consistent across explicit cuts inside a single prompt.
- Automatic clip-duration prediction: the model infers a sensible length from the action described, so don't pad a prompt just to "fill time."
- Image-to-Video and FLF2V (First-Last-Frame-to-Video) conditioning for animating stills and connecting two frames.
- Native 4K HDR / RAW support for professional pipelines.
- Text connector: Gemma-based, larger and more literal — it will render almost everything you write, so don't include anything in the prompt you don't want to see on screen.

---

## 1. The Six Core Elements (every prompt, every mode)

Cover all six where relevant to the mode. Skipping one you actually need is the #1 cause of weak output.

| # | Element | What to include |
|---|---|---|
| 1 | Shot | Cinematography terms matching the genre (wide shot, medium close-up, over-the-shoulder, etc.) and shot scale |
| 2 | Scene | Lighting condition, color palette, surface textures, atmosphere/mood |
| 3 | Action | The core action as one clear, natural sequence, beginning to end |
| 4 | Character(s) | Age, hairstyle, clothing, distinguishing features. Emotion via physical cues, never abstract labels |
| 5 | Camera movement | How and when it moves (pan, push-in, track, static, handheld). Describe how the subject looks after the move completes |
| 6 | Audio | Ambient sound, music, speech, singing. Dialogue in quotation marks; state language/accent if relevant |

Golden rule: if the prompt reads like it's describing a still photo, the output will move like one. Every sentence should imply motion or time passing.

Exception: for Image-to-Video and FLF2V, elements 2 and 4 (Scene, Character) are usually already supplied by the source image(s) — don't re-describe what's already visible, focus on 1, 3, 5, and 6 instead.

---

## 2. Mode 1 — Single-Shot (one continuous take)

Rules: one flowing paragraph, present tense throughout, match detail density to shot scale, describe camera movement relative to the subject, 4–8 sentences with no filler, one dominant event per prompt, consistent lighting logic, iterate from simple.

Template:
"A [shot scale] frames [character description] in [setting], [lighting/atmosphere description]. [Character] [core action in present tense], [secondary physical/emotional detail]. The camera [movement] as [what happens as a result]. [Ambient sound/music]. [Dialogue in quotes, with language/accent noted if relevant]."

---

## 3. Mode 2 — Screenplay-Style (dialogue-heavy or precisely timed)

Use when a scene has dialogue, multiple beats, or timing precision a single paragraph can't carry cleanly. Scene headers, character cues, and quoted dialogue are appropriate here. Same fundamentals apply: present tense, physical emotion cues, dialogue in quotation marks. Length scales with complexity.

---

## 4. Mode 3 — Multi-Shot (2–4 cuts in one prompt)

Critical formatting rule: write the full scene as one chronological paragraph. Do NOT use a numbered shot list, bullet beats, or screenplay sluglines.

At every cut, include all four of:
1. Name the transition in natural language — "A hard cut transitions to...", "The view cuts to a close-up of...", "A match cut connects...", "The image dissolves into..."
2. Re-establish the new shot — shot scale, camera angle, who/what is in frame, lighting if it changed.
3. Keep identity consistent — reuse the same visual identifiers when a person/object reappears.
4. State audio continuity — e.g. "the piano score continues across the cut" or "the dialogue drops; only wind remains."

Prefer 2–4 shots per generation. Keep action chronological.

---

## 5. Mode 4a — Image-to-Video (animating a single source image)

Image-conditioned. The source image already supplies appearance, setting, and lighting — the prompt should describe motion, camera behavior, and audio, not re-describe what's already in the frame.

Rules: keep it short (2–5 sentences), present tense, one dominant action, be explicit about camera movement and audio since there's no prior motion to infer from, avoid re-stating clothing/hair/color details that are already visible in the source image.

Template: "[Camera movement/behavior], [subject's action in present tense], [secondary detail]. [Audio]."

Common mistake to catch: re-describing the whole image in words. If the written description doesn't perfectly match the image, the model tries to reconcile the two and the subject drifts or morphs mid-clip.

---

## 6. Mode 4b — FLF2V (First-Last-Frame-to-Video)

Image-conditioned, two source images: a first frame and a last frame. The prompt describes only the transition connecting them — not either frame's content.

Rules: one clear, physically plausible transition per generation (same "one dominant event" rule as Single-Shot), name the camera behavior across the transition, note audio if relevant, keep the two source frames close enough in composition that the described motion can plausibly connect them.

Template: "[Camera/motion connecting the two frames], as [subject] [description of the transitional action]. [Audio]."

Common mistake to catch: first and last frames that are too different (subject position, lighting, composition) with no transition motion described — the model fills the gap with the most literal, sometimes jarring, interpolation. Either choose closer frames or describe the connecting motion explicitly.

---

## 7. Vocabulary Reference

Genre: Stop-motion, 2D/3D animation, Claymation, Hand-drawn, Comic book, Cyberpunk, 8-bit pixel, Surreal, Minimalist, Painterly, Illustrated, Period drama, Film noir, Fantasy, Epic space opera, Thriller, Modern romance, Documentary, Arthouse

Lighting: Flickering candles, Neon glow, Natural sunlight, Dramatic shadows, Golden hour, Tungsten warmth

Camera language: Follows, Tracks, Pans across, Circles around, Tilts upward, Pushes in / pulls back, Overhead view, Handheld, Over-the-shoulder, Wide establishing shot, Static frame

Pacing: Slow motion, Time-lapse, Rapid cuts, Lingering shot, Continuous shot, Freeze-frame, Fade-in/out, Seamless transition, Sudden stop

Audio: Coffeeshop noise, Wind and rain, Forest ambience with birds, Energetic announcer, Resonant voice with gravitas, Whisper, Mutter, Shout

---

## 8. Known Limits

- On-screen text: keep any text short and prominent; verify spelling and add critical logos/titles in post if precision matters.
- Complex physics: highly chaotic motion (crowds colliding, debris explosions, turbulent cloth) can still produce artifacts. Steer toward simpler, plausible motion.

---

## 9. Common Mistakes to Catch Before Outputting a Prompt

- Reads like a still photo → rewrite with active present-tense action.
- Multiple competing actions in one shot → cut to the single dominant event.
- Mixed/inconsistent lighting logic → pick one light source and stick to it.
- Multi-shot written as a numbered list → convert to prose with named transitions.
- Multi-shot cut with no re-established framing or audio continuity → add both.
- Image-to-Video prompt re-describing the source image instead of just the motion → strip appearance detail, keep motion/camera/audio only.
- FLF2V prompt missing a described transition between very different frames → add explicit connecting motion.
- Emotion described as a label instead of a physical cue → convert to physical cues.
- Padding a prompt to make the video "longer" → duration is automatic; describe the real action.

---

## 10. Output Format

When the user gives you an idea, respond with:
1. The finished prompt (mode-appropriate format), ready to paste into LTX-2.5.
2. If useful, one short line noting anything you assumed so they can correct it.

Don't restate this whole skill file back to the user — just produce the prompt.

How to use: Open claude.ai, paste this file as your first message, then follow up with: "Write an LTX-2.5 prompt for: [describe your scene in plain language]". Copy the output into your ComfyUI Gemma text encoder node.

Frequently asked questions

This almost always means one of the six core elements is missing, or the prompt is too vague for the shot scale you asked for. LTX-2.5's Gemma-based text encoder renders what you write literally, so a short or list-style prompt gives it too little to work with and it fills the gap on its own.

Single-Shot describes one continuous take with one camera setup. Multi-Shot describes 2–4 cuts inside a single generation, written as one chronological paragraph with named transitions between each cut. Using Single-Shot phrasing for a multi-cut idea is the most common cause of broken-looking cuts.

The underlying six-element approach still applies, but LTX-2.5 adds native multi-shot generation and automatic duration prediction, so older prompts padded with filler sentences to "fill time" will underperform. Trim to the real action and pick the correct mode before reusing an old prompt.

For Single-Shot, 4–8 sentences where every sentence adds a concrete visual or audio detail. Multi-Shot needs more, roughly one full beat per cut. Image-to-Video and FLF2V prompts should stay short and motion-focused — the source image already supplies the appearance.

Yes, through Image-to-Video. Upload the source image and write a short, motion-only prompt — camera movement, the subject's action, and audio. Re-describing what's already visible in the image tends to cause drift instead of helping.

FLF2V (First-Last-Frame-to-Video) generates the motion connecting a starting image and an ending image you provide — useful for a product spin, a transformation, or a controlled transition. The prompt only needs to describe the transition itself, not either frame.

What to do next

Run the diagnostic on your last broken clip, fix the one element it points to, and regenerate.

If the prompt checks out and the output is still wrong, the settings guide covers every node, VRAM requirement, and sampler value LTX-2.5 needs in ComfyUI.

Tutorial · Roadmap Level 2 · Video Generation series · ComfyUI interface overview · LTX-2.5 ComfyUI settings guide