Earngenix Logo
Skip to main content

Model Comparison · All Levels · Updated 2026

Best Local AI Video Model 2026: LTX-2.3 vs WAN 2.2 vs HunyuanVideo 1.5

Three open-weight video models now run locally in ComfyUI — and all three do both text-to-video and image-to-video. This comparison tests LTX-2.3, WAN 2.2, and HunyuanVideo 1.5 on VRAM, speed, realism, anime style, camera motion, physics, audio, and licensing, for both generation modes, so you download the right one first.

3

Models compared

9

Categories

RTX 4090

Tested on

All levels

Skill level

By Earngenix Team ·

⚡ Quick Answer

For most people running ComfyUI locally, LTX-2.3 is the best all-around local AI video model for both text-to-video and image-to-video — it's the fastest of the three, runs on the least VRAM (8GB+ with GGUF), and is the only model that generates synced audio with the video. If photorealistic humans or believable physics are your priority — for T2V or animating a reference photo — and you have 16GB+ VRAM, WAN 2.2 produces the strongest results. HunyuanVideo 1.5 is roughly a year old at this point, and with LTX-2.3 and WAN 2.2 both having moved on since, it no longer leads any category here — it's still usable, but there's no longer a clear reason to reach for it over the other two.

Picking the best local AI video model between LTX-2.3, WAN 2.2, and HunyuanVideo 1.5 comes down to three questions: how much VRAM do you have, what does your video need to look like, and do you need audio. All three do both text-to-video (T2V) — generating a clip from a prompt alone — and image-to-video (I2V) — animating a reference photo — so this comparison covers both modes rather than treating them as separate decisions. It runs all three through the same categories — VRAM, speed, realism, stylized output, camera motion, physics, audio, and licensing — so you pick the right one before downloading 20GB+ of the wrong model.

All three run natively in ComfyUI with official or community-maintained nodes. If you haven't installed ComfyUI yet, install ComfyUI first before downloading any of the models below.

What You Need Before You Start

Check your GPU's VRAM before picking a model — this decides which one is even usable on your hardware, before quality enters the conversation. VRAM needs are the same whether you're generating text-to-video or image-to-video.

VRAM tiers across all three models

8–12 GB VRAM: LTX-2.3 (GGUF quantized) is your only realistic option among these three at this tier, for T2V or I2V.
14–16 GB VRAM: All three become usable — LTX-2.3 FP8, WAN 2.2 14B FP8 with block swap, and HunyuanVideo 1.5 with model offloading enabled.
24 GB VRAM (RTX 4090 class): All three run comfortably at full precision without offloading tricks, in either mode.

Quantization, in plain terms, is a compressed version of a model file — GGUF and FP8 formats trade a small amount of quality for a much smaller file size and lower VRAM use. All three models publish quantized versions.

Warning: Don't download the full BF16/FP16 version of any of these three models "to be safe" if you're under 16GB VRAM. It will not load, and you'll have wasted 20–40GB of download and disk space. Check the VRAM table in each model's dedicated guide (linked further down) before downloading.

For a full GPU buying guide across all three models — which card to buy at each budget, running on 8GB VRAM, fixing "CUDA out of memory" errors, and AMD/Apple Silicon/cloud rental options — see the VRAM & Hardware Guide for AI Video Generation.

LTX-2.3 vs WAN 2.2 vs HunyuanVideo 1.5 at a Glance

A fast top-level scan before the category-by-category breakdown below. Note that all three support both text-to-video and image-to-video — this isn't a text-to-video-only comparison.

LTX-2.3WAN 2.2HunyuanVideo 1.5
DeveloperLightricksAlibabaTencent
Parameters22B5B / 14B8.3B
Text-to-video (T2V)✅ Yes✅ Yes✅ Yes
Image-to-video (I2V)✅ Yes✅ Yes✅ Yes
Min VRAM8GB (GGUF)12GB (block swap)14GB (offloaded)
Comfortable VRAM16GB+ (FP8)24GB (14B, no offload)24GB
Native audio✅ Yes❌ No❌ No
Speed (relative)FastestSlow (fast w/ Lightning LoRA)Middle
Best forSpeed, low VRAM, audioPhotorealism & natural physics (T2V + I2V)Legacy option — no longer leads a category
LicenseOpen weights, commercial use permittedApache 2.0Tencent Community License (restricted)

Which Model Is Fastest?

LTX-2.3 is roughly 2–3x faster than WAN 2.2 or HunyuanVideo 1.5 at comparable settings, in both T2V and I2V. On an RTX 4090 with the FP8 distilled model at 8 steps, LTX-2.3 generates a 5-second clip in 20–40 seconds.

HunyuanVideo 1.5 sits in the middle — its smaller 8.3B parameter count (versus the older 13B HunyuanVideo) is specifically what made single-GPU generation practical, but it's still slower than LTX-2.3's distilled pipeline.

WAN 2.2 is the slowest of the three at standard settings — a full 14B render with 20+ steps can take 40–50 minutes for a single 5-second clip, for either mode. With the community Lightning LoRA — a distilled step-reduction LoRA that cuts sampling to 4 steps per expert model — that drops to roughly 1–3 minutes, closing most of the speed gap.

Tip: If you're iterating on a prompt and don't need final quality yet, generate drafts at lower resolution or with a distilled model — a smaller, faster version of the same model trained to approximate the full model's output in fewer steps — regardless of which of the three you're using, or which mode you're generating in. Switch to full precision only for your final render.

See Each Model's Own Text-to-Video Showcase

Before the matched-prompt breakdown below, here's each model's own T2V demo clip from its dedicated setup guide — these use different prompts per model, so treat them as "what this model looks like at its best," not a direct comparison.

LTX-2.3

Lighthouse on an Atlantic Cliff (LTX-2.3's own T2V demo)

LTX-2.3 · BF16 Distilled · 2560×1408

WAN 2.2

A Girl Walking on the Street (Wan 2.2's own T2V demo)

Wan 2.2 · Lightning 4-step · 832×480

HunyuanVideo 1.5

Jungle Temple Explorer (HunyuanVideo 1.5's own T2V demo)

HunyuanVideo 1.5 · 640×480

Image-to-Video Compared: Animating a Reference Photo

This is the category most text-to-video-only comparisons skip — but all three models here ship a dedicated I2V pipeline, and for a lot of real workflows (bringing a product photo or character portrait to life) it matters more than T2V does. Each model below animates its own reference photo from its dedicated setup guide — not the same source image across all three — so read this as "what each model's I2V looks like at its best," the same honest framing as the T2V demos above.

LTX-2.3 reference input photo

LTX-2.3 — input photo

WAN 2.2 reference input photo

WAN 2.2 — input photo

HunyuanVideo 1.5 reference input photo

HunyuanVideo 1.5 — input photo

LTX-2.3

Beauty Tutorial Portrait Animation (LTX-2.3's own I2V demo)

LTX-2.3 · FP8 Distilled · denoise 0.85

WAN 2.2

Girl Walking, Animated from a Still Photo (Wan 2.2's own I2V demo)

Wan 2.2 · Lightning 4-step · 832×480

HunyuanVideo 1.5

Neon Rain Walk, Animated from a Still Photo (HunyuanVideo 1.5's own I2V demo)

HunyuanVideo 1.5 · 480×640

On quality: WAN 2.2 generally holds up best for photorealistic I2V — skin texture and hair stay consistent frame to frame, particularly on close-up portraits. LTX-2.3 is close behind and is the only one of the three that can add synced audio to an animated image in the same generation pass. HunyuanVideo 1.5 does well with natural-feeling motion once the reference image and prompt are aligned, though its CLIP Vision step is an easy one to misconfigure — see the dedicated guide's troubleshooting section if your output ignores your input image.

Warning: The three clips above are each model's own showcase output on its own reference photo — not the same image run through all three models. A real side-by-side (same photo, same prompt, all three models) is planned for the "Realistic People" section below once those clips are generated.

Which Model Has the Best Video Quality?

"Best quality" isn't one number — it splits by what you're actually generating. Each section below uses the exact same prompt run through all three models, in text-to-video mode unless noted.

Realistic People and Photorealism

WAN 2.2 is generally considered the strongest for photorealistic human subjects — skin texture, hair, and lighting hold up best in close-up shots, in both T2V and I2V. HunyuanVideo 1.5 is close behind, with community testing praising its texture consistency, particularly in fine-tuned variants. LTX-2.3 is competitive at medium and wide shots, but its VAE architecture shows more softness in extreme close-ups than the other two.

LTX-2.3

LTX-2.3

Realistic — same prompt

WAN 2.2

WAN 2.2

Realistic — same prompt

HunyuanVideo 1.5

HunyuanVideo 1.5

Realistic — same prompt

same prompt — all three models
A cinematic medium shot of a woman in her late 20s sitting at a cozy café beside a rain-streaked window, gently holding a warm cup of coffee with both hands. Raindrops slowly slide down the glass while soft afternoon light filters through the clouds. She slowly looks up toward the window, a subtle smile appears on her face, then she lowers her gaze back to the cup. Steam rises naturally from the coffee as she softly blinks. The camera performs a slow dolly-in with subtle handheld motion, creating an intimate, peaceful atmosphere. Natural movement, realistic facial expressions, gentle hair motion, cinematic lighting

Anime and Stylized Video

LTX-2.3 tends to produce the strongest results for stylized, animated, and non-photorealistic content — motion graphics, illustration styles, and artistic camera work. If your channel or client work leans stylized rather than photoreal, this is where LTX-2.3's speed advantage compounds: more iterations in the same amount of time.

LTX-2.3

LTX-2.3

Anime/Stylized — same prompt

WAN 2.2

WAN 2.2

Anime/Stylized — same prompt

HunyuanVideo 1.5

HunyuanVideo 1.5

Anime/Stylized — same prompt

same prompt — all three models
An anime girl stands alone on a rooftop at sunset, her silver hair and jacket flowing gently in the evening wind. She takes a few slow steps toward the edge, pauses, and watches the sun disappear below the skyline. Birds fly across the distant sky while clouds slowly drift overhead. The camera starts behind her, slowly circles around to the front, and finishes with a close-up of her peaceful smile. Smooth cinematic camera movement, expressive animation, natural hair and clothing motion, warm golden evening light. 
Warning: Comparison clips above are placeholders — swap in real matching-prompt generations from all three models before publishing.

Camera Motion and Cinematography

LTX-2.3 follows detailed camera instructions — dolly, pan, push-in, lens type — closely, and specifying exact cinematography terms in the prompt produces noticeably different results. WAN 2.2 also handles camera movement well and holds background elements steady during complex foreground motion. HunyuanVideo 1.5 is generally solid but can show slight warping on very long, complex camera moves.

LTX-2.3

LTX-2.3

Camera motion — same prompt

WAN 2.2

WAN 2.2

Camera motion — same prompt

HunyuanVideo 1.5

HunyuanVideo 1.5

Camera motion — same prompt

same prompt — all three models
The camera glides slowly through an ancient stone corridor toward a weathered wooden door. Warm candlelight dances across the rough stone walls while dust drifts gently through the air. As the camera reaches the doorway, it smoothly arcs to the right, revealing a wooden table holding a softly flickering candle beside an old parchment. The scene remains calm and continuous with slow cinematic movement, subtle environmental animation, warm ambient lighting, and a mysterious atmosphere.
Warning: Comparison clips above are placeholders — swap in real matching-prompt generations from all three models before publishing.

Natural Motion and Physics

WAN 2.2 currently leads here. Fluid dynamics — water, smoke, fire — cloth simulation, and physical object interactions look more physically grounded than in LTX-2.3 or HunyuanVideo 1.5 at equivalent settings, in both T2V and I2V. If your scene depends on believable physical motion rather than a static or slow-moving subject, WAN 2.2 is worth the extra generation time (or the Lightning LoRA if you want that quality without the wait).

HunyuanVideo 1.5 was the physics leader in this comparison when it first released, but AI video moves fast — LTX-2.3 and WAN 2.2 have both had major updates since, and at roughly a year old, HunyuanVideo 1.5 has been passed on this front too. It's not a bad model, it's just no longer the newest one in the room.

Native Audio

LTX-2.3 is the only model of the three that generates audio synced to the video in the same pass — not added afterward as a separate step, for text-to-video or image-to-video. Neither WAN 2.2 nor HunyuanVideo 1.5 generates audio natively; if you need sound with either of those, you're adding it in a separate pipeline step outside ComfyUI's video generation.

No comparison clips needed here — this category has one answer, not three. LTX-2.3's own demos above already include synced audio generated in the same pass as the video, for both T2V and I2V.

Licensing — What You Can Actually Use Commercially

This matters more than most comparisons mention, and it applies the same whether you're generating T2V or I2V output. WAN 2.2 is released under the Apache 2.0 license — the most permissive of the three, with no meaningful commercial restrictions.

HunyuanVideo 1.5 uses the Tencent Hunyuan Community License — not Apache 2.0 or MIT. It allows personal, research, and most commercial use, but it explicitly does not apply in the EU, UK, or South Korea, and any product built on it that reaches over 100 million monthly active users needs a separate license from Tencent.

LTX-2.3's weights are published on HuggingFace under an open license permitting personal and commercial use. Terms can update, so check the current license file on the HuggingFace repo before publishing commercially — this applies to all three models, not just LTX-2.3.

Winner by Category

CategoryWinner
Fastest generationLTX-2.3
Lowest VRAM requirementLTX-2.3
Photorealistic humans (T2V)WAN 2.2
Image-to-video qualityWAN 2.2
Anime / stylized outputLTX-2.3
Camera motion accuracyLTX-2.3
Natural physics / motionWAN 2.2
Native audioLTX-2.3
Most permissive licenseWAN 2.2

So Which One Should You Actually Use?

Pick LTX-2.3 if…

You want one model that covers the most ground — fastest generation, lowest VRAM floor (8GB with GGUF), strong camera-motion accuracy, and the only one of the three with native audio for either T2V or I2V. This is the practical default for most local ComfyUI users.

Pick WAN 2.2 if…

Photorealistic human subjects or believable physics are the priority — flowing water, smoke, cloth — for text-to-video or animating a reference photo, and you have at least 16GB VRAM. The Apache 2.0 license also makes it the safest legal choice for unrestricted commercial use.

Pick HunyuanVideo 1.5 if…

You already have a workflow or LoRAs built around it, or LTX-2.3 and WAN 2.2 genuinely don't fit your hardware. It's a solid model, but at roughly a year old it no longer leads any category against the two newer models above — treat it as a fallback rather than a first choice in 2026.

Nothing stops you from running all three — the models don't conflict, and many workflows keep two installed: one for speed (LTX-2.3) and one for a specific quality need (WAN 2.2 or HunyuanVideo 1.5), regardless of whether you're mostly doing T2V, I2V, or both.

Full Setup Guides for Each Model

Each of these has a complete, tested ComfyUI setup guide covering both text-to-video and image-to-video, with exact model downloads, folder structure, and settings.

Troubleshooting: Common Setup Issues Across All Three

"This node type does not exist" (red nodes on any of the three workflows)

You're missing a custom node the workflow depends on — this happens most often with HunyuanVideo 1.5's CLIP Vision node (needed for I2V specifically), LTX-2.3's GGUF workflow, or WAN 2.2's WanVideoWrapper nodes.

  1. Click Manager in the top menu.
  2. Click Install Missing Custom Nodes.
  3. Restart ComfyUI completely, not just a browser refresh.

Black frames or no output on any model

Almost always a missing or misplaced text encoder file. Each of the three needs its text encoder(s) in models/text_encoders/ exactly as named in its dedicated guide above — double-check the filename, not just the folder.

Image-to-video output ignores your reference photo

This is specific to I2V and usually means a wiring problem, not a bad image. On HunyuanVideo 1.5, confirm CLIPVisionLoader and CLIPVisionEncode are actually connected into the HunyuanVideo15ImageToVideo node — if that link is missing, ComfyUI still generates a video, it just won't use your photo. On LTX-2.3 and WAN 2.2, confirm you loaded the I2V-specific model file, not the T2V one, into the checkpoint/model loader.

"CUDA out of memory" during generation

Your VRAM doesn't match the model variant you downloaded.

  1. Confirm you downloaded the quantized (FP8 or GGUF) version, not full BF16/FP16, if you're under 16GB.
  2. Lower resolution or frame count first before switching models entirely.
  3. Close other GPU-heavy applications and retry.

For anything not covered here, see the full ComfyUI troubleshooting guide.

Frequently Asked Questions

LTX-2.3 is the strongest all-around pick for most ComfyUI users — it's the fastest, runs on the lowest VRAM, and is the only one with native audio. WAN 2.2 wins on photorealistic quality and on natural physics/motion, for both T2V and I2V. HunyuanVideo 1.5 is roughly a year old now and no longer leads any category against the other two.

Yes. All three ship dedicated image-to-video pipelines alongside text-to-video — this comparison deliberately covers both, since a text-to-video-only comparison would miss half of what each model actually does.

Yes. Each model uses its own set of files in separate ComfyUI model folders, so installing all three doesn't cause conflicts — you're limited only by disk space, not compatibility.

LTX-2.3, using the GGUF quantized version, runs on 8GB VRAM. WAN 2.2 is also usable around 12GB with block swap enabled, though HunyuanVideo 1.5 needs roughly 14GB as a floor.

No. Neither generates audio natively, for text-to-video or image-to-video. LTX-2.3 is the only one of the three that produces synced audio and video in the same generation pass.

Mostly, but with conditions. The Tencent Hunyuan Community License permits commercial use, except in the EU, UK, and South Korea, and products exceeding 100 million monthly active users need a separate license from Tencent.

LTX-2.3 generally produces the strongest results for stylized and non-photorealistic content, and its speed advantage means more iterations to refine a style in the same amount of time.

What to Do Next

Start with LTX-2.3 — then branch out to WAN 2.2 or HunyuanVideo 1.5

Start with LTX-2.3 if you're not sure which to try first — it has the lowest VRAM floor, so it's the most likely to actually run on your hardware today, and it covers both T2V and I2V from one workflow. Get one clip generating successfully, then come back and try WAN 2.2 or HunyuanVideo 1.5 for the specific quality your project needs.

Published: 2026-07-17 · Last updated: 2026-07-18 · Models compared: LTX-2.3 (Lightricks), WAN 2.2 (Alibaba), HunyuanVideo 1.5 (Tencent) · Both text-to-video and image-to-video covered for each model

Discussion

Join the discussion

Sign in to leave a comment or reply

💬

No comments yet

Be the first to share your thoughts!