New: create a free account and get 14 days ad-free.Sign up free
EarnGeniX
Skip to main content

Workflow Blog · Beginner–Intermediate · Updated September 2026

MiniMax H3 FastVideo in ComfyUI: Run the 4-Step Model Today

A distilled MiniMax H3 checkpoint that runs in 4–8 steps instead of 20–50 — exact files, real settings, and tested speed on an RTX 4090.

Free

Cost

≈49 GB

Min Disk

Beginner–Int.

Skill level

4–8

Steps

By Earngenix Team · · Tested on ComfyUI 0.30.0+, RTX 4090

⚡ Quick Answer

MiniMax H3 FastVideo is a 4-step distilled version of MiniMax H3, built by the FastVideo research team (hao-ai-lab) and converted for ComfyUI by Kijai. It swaps H3's normal 20–50 sampling steps for 4–8, runs on a standard UNETLoader plus a SageAttention patch node, and turns a several-minute generation into roughly two and a half to three minutes on consumer hardware. This guide covers the exact files, the settings that actually matter, and real tested speed on an RTX 4090.

If you've already run base MiniMax H3 in ComfyUI, you know the wait: a single clip can take a long time even with SageAttention turned on. MiniMax H3 FastVideo is a separate, distilled checkpoint built specifically to fix that, and it landed in the ComfyUI ecosystem only in the last few weeks. This guide walks through downloading the right files, wiring the workflow, and understanding the handful of settings that control speed and quality.

Hardware and versions used for this guide: ComfyUI 0.30.0+, tested on an RTX 4090 (24GB VRAM).

What to Expect: Example Outputs

Before you download anything, here's what the FastVideo workflow actually produces — both clips below were generated as image-to-video, using the same 4-file setup covered further down this page.

FastVideo example

Image-to-video output from the FastVideo 4-step checkpoint.

FastVideo example

A second clip generated with the same settings, different prompt and source image.
Tip: Both clips above were generated at 8 sampling steps, which is the workflow's default — not the minimum of 4. The settings section further down explains exactly why 8 is a better starting point than 4.

What Is MiniMax H3 FastVideo, and Why Is It Faster?

MiniMax H3 normally generates video through a diffusion model (this is the part of the AI that turns random noise into a finished video, a little at a time, over many repeated passes called steps) — and it typically needs somewhere between 20 and 50 of those steps to produce a clean result. Each step means running the full model again, so more steps directly means more waiting.

FastVideois a separate checkpoint (this means a different saved version of the model's weights, not a plugin or add-on) built by the FastVideo research team at hao-ai-lab. They used a technique called DMD2 distillation to train a version of H3 that produces a comparable result in just 4 steps. Kijai — a well-known ComfyUI node developer — then converted that checkpoint into a format ComfyUI can load natively, which is the file this guide uses.

The workflow in this guide loads that checkpoint through a plain UNETLoader node, the same node type used for any other diffusion model in ComfyUI. It also uses a SageAttentionpatch node (this speeds up the model's attention calculations, a core part of how it processes each step) instead of FastVideo's more specialized Video Sparse Attention system — that specialized system still needs two ComfyUI pull requests to be merged before it works out of the box, so SageAttention is the practical path today.

FastVideo vs. MiniMax H3 Turbo LoRA — What's the Difference?

You may have also come across the MiniMax H3 Turbo LoRA, and it's easy to assume they're the same thing. A LoRA (short for Low-Rank Adaptation) is a small patch file you load on top of an existing model to change how it behaves — the Turbo LoRA, built by a different creator called larryvrh, is applied on top of the standard base MiniMax H3 checkpoint. FastVideo, by contrast, is a complete, standalone checkpoint with the distillation baked directly into the weights — there's no separate base model to load alongside it, and no LoRA loader node involved. If you want the LoRA route instead, we cover that setup in our MiniMax H3 Turbo LoRA workflow guide.

What You Need Before You Start

Unlike base MiniMax H3, FastVideo only comes in one size — there's no bf16, int8, or pruned choice to make here, since the distillation itself is the compression. You need exactly four files, each going into a different ComfyUI models folder:

FileGoes InSizeWhat It Is
minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot.safetensorsdiffusion_models/≈21 GBThe FastVideo diffusion model itself — this is the file that replaces H3’s normal 20–50 step sampling with 4–8 steps.
qwen3vl_32b_minimax_h3_int8_convrot.safetensorstext_encoders/≈27.1 GBThe text encoder — this is the part of the model that reads your written prompt and turns it into something the diffusion model can understand.
minimax_h3_video_vae_fp16.safetensorsvae/≈500 MBThe video VAE (Variational Auto-Encoder) — decodes the model’s internal representation into the actual video frames you see.
minimax_h3_audio_vae_fp32.safetensorsvae/≈190 MBThe audio VAE — decodes the model’s internal representation into the stereo sound that plays alongside the video.
Tip: These four buttons link directly to the files on Hugging Face — the diffusion model comes from Kijai's ComfyUI conversion, and the text encoder plus both VAE files come from Comfy-Org's official MiniMax H3 repository, since FastVideo reuses H3's standard text encoder and VAEs.
Warning: Skip the audio VAE and the workflow will still run — your video just won't have any sound. This is one of the most common setup mistakes with every MiniMax H3 workflow, not just this one.

Where Everything Goes

ComfyUI/ ├── models/ │ ├── diffusion_models/ │ │ └── minimax_h3_fastvideo_vsa_datafree_1300step_4step_int8_convrot.safetensors │ ├── text_encoders/ │ │ └── qwen3vl_32b_minimax_h3_int8_convrot.safetensors │ └── vae/ │ ├── minimax_h3_video_vae_fp16.safetensors │ └── minimax_h3_audio_vae_fp32.safetensors

You also need ComfyUI 0.30.0 or later for native MiniMax H3 node support, and the SageAttention custom node installed, since the workflow's attention patch has nothing to connect to without it. If you haven't installed it yet, follow our SageAttention install guide first — it takes a few minutes and this workflow won't run without it.

Tip: For a full breakdown of how much VRAM each MiniMax H3 variant needs and which quantization level fits your GPU, see our VRAM guide for AI video in ComfyUI. FastVideo is one of the lighter options on that list.

How to Set Up the MiniMax H3 FastVideo Workflow

This is an image-to-video workflow — you provide a starting image and a prompt describing what happens next, and FastVideo animates it. The graph looks busy the first time you open it, but you only need to touch four things: the loader nodes, the SageAttention node, your prompt, and Queue Prompt.

(minimax_h3_fastvideo_i2v.json)
  1. Confirm your four loader nodes are correct. Before touching anything else, check that UNETLoader, CLIPLoader, and both VAELoader nodes each have the matching file selected from the table above — not blank, and not a file from a different MiniMax H3 workflow. If a file doesn't show up in a dropdown, it's sitting in the wrong folder.
  2. Install and connect SageAttention. This workflow includes a PathchSageAttentionKJ node sitting between the model loader and the sampler. We need this connected because it's what lets the model run its attention calculations faster — without it installed, this node will error out the moment you try to generate. Install it through ComfyUI Manager first if you haven't already.
  3. Load your starting image. Connect an image to the MiniMax H3 Image To Video node's image input. This is the frame FastVideo will animate from.
  4. Write a structured prompt. MiniMax H3 responds best to prompts broken into shots with timestamps, describing what the camera sees and what changes. We cover the exact format in the prompt section just below — write your prompt directly into the MiniMax H3 Image To Video node's text field.
  5. Set your duration and resolution. Still on the same node, set how many seconds of video you want and pick a resolution. As covered further down, your requested duration gets rounded up to the model's fixed frame grid, so don't worry if the output runs a little longer than what you typed.
  6. Queue the prompt. Click the orange Queue Prompt button in the top-right of the screen. A progress bar appears underneath it. On an RTX 4090, a 5-second clip takes roughly 150 to 180 seconds — see the speed section below for the full breakdown.
Full MiniMax H3 FastVideo node graph in ComfyUI🔍 Click to zoom
The full FastVideo workflow — the loader nodes and SageAttention patch feed into the sampler on the left.
MiniMax H3 Image To Video node with a structured shot-by-shot prompt filled in🔍 Click to zoom
A structured, shot-by-shot prompt in the MiniMax H3 Image To Video node.

For the full breakdown of MiniMax H3's prompt structure — the shot format, camera direction, and worked examples — see our MiniMax H3 prompt guide. The same structure applies whether you're on the base model or FastVideo.

Finished MiniMax H3 FastVideo output shown in the ComfyUI output panel🔍 Click to zoom
A finished clip in the ComfyUI output panel after queuing.

What Do the FastVideo Settings Actually Control?

This workflow has five settings worth understanding, in the order you're most likely to need to touch them.

How Many Steps Should You Use?

The Scheduler node's steps value controls how many times the model refines the video before finishing. Every single step re-runs the full model, so step count is directly proportional to how long you wait. The workflow ships with 8 stepsas its default, which noticeably improves motion quality over the model's minimum of 4 steps for a modest time cost. Going up to 12–14 steps is fine if you want to double-check framing and motion before committing to a longer render, but pushing much past that stops helping — FastVideo was distilled specifically to work well in a small number of steps, and extra steps beyond roughly 14–20 mostly just add wait time.

Why Doesn't My Video Match the Duration I Set?

MiniMax H3 generates video in fixed frame blocks at 24 frames per second, so typing "2.0 seconds" doesn't always give you exactly 2.0 seconds back — the workflow rounds your request up to the nearest length the model can actually produce:

You SetFramesReal Length
1.0 s391.62 s
2.0 s562.33 s
2.5 s or 3.0 s733.04 s
3.5 s903.75 s
4.0 s1074.46 s
4.5 s or 5.0 s1245.17 s
Warning: Don't type a duration between two of these values expecting something in between — the workflow always rounds up, never down. Setting 2.5 seconds gives you the same 3.04-second result as setting 3.0 seconds.

What Resolution Should You Pick?

Resolution barely affects how much VRAM you use, since the model's weights take up far more memory than the video frames being generated — what it does change is how long generation takes. FastVideo was trained on a 768-pixel short edge, capped at 768×1344, so going below 768 on the short edge costs real quality for only a modest speed gain:

Aspect RatioMegapixelsResult
1:10.56768 × 768
4:30.751024 × 768
3:40.75768 × 1024
16:90.981344 × 768 (default)
9:160.98768 × 1344

1344×768 at 5 seconds is the heaviest combination this checkpoint officially supports — around 37,000 tokens passing through the model at once. It still fits in 8GB of VRAM, but it's the first place you'd see an out-of-memory error if something else is using your GPU at the same time.

Should You Touch Sigma Shift or Text Encoder Placement?

The workflow includes a MiniMax H3 Sigma Shiftnode, which controls how the model paces its denoising schedule across video and audio separately. The values shipped with the workflow match the model's own built-in defaults, so this node is effectively a no-op unless you go looking for it — it's there so you have the option to lower the video shift value for tighter, less drifty motion if you want to experiment. Leave it alone for your first few generations.

The CLIPLoadernode also has a device setting that can be switched to run the text encoder on your CPU instead of your GPU. Only do this if you're running out of memory during the text-encoding step — it permanently pins around 15.7GB in your system RAM and runs a 32-billion-parameter model on your CPU, which is slower than letting ComfyUI stream it through your GPU as needed.

How Fast Is MiniMax H3 FastVideo on an RTX 4090?

These numbers come from our own testing on an RTX 4090 (24GB VRAM), using the default 8-step setting and the SageAttention patch node connected and working.

A 5-second clip at the default 1344×768 resolution took roughly 150 to 180 seconds to generate on our RTX 4090 — under three minutes, compared to several minutes for the same length and resolution on base MiniMax H3 without FastVideo.

Both example clips in the "What to Expect" section above were generated at these exact settings: 8 steps, 1344×768, 5 seconds. If your own generation times are noticeably slower, check three things first: whether SageAttention is actually connected and installed (not bypassed), whether another application is using your GPU at the same time, and whether you've bumped the resolution or duration above the defaults shown in the settings section.

Troubleshooting MiniMax H3 FastVideo Errors

"This node type does not exist" (MiniMaxH3ImageToVideo)

What causes it: Your ComfyUI installation is older than version 0.30.0, so it doesn't recognize MiniMax H3's native nodes yet.

How to fix it: Open ComfyUI Manager (the puzzle-piece icon in the top toolbar) and click Update ComfyUI. Restart ComfyUI completely afterward — node definitions only load on startup.

SageAttention Node Errors on Generation

What causes it: The PathchSageAttentionKJ node is connected in the workflow, but the SageAttention custom node itself isn't installed on your system.

  1. Install SageAttention through ComfyUI Manager, following our install guide.
  2. If you want to test the workflow before installing it, right-click the PathchSageAttentionKJ node and choose Bypass (or select it and press Ctrl+B). Generation will still work, just noticeably slower.

Generated Video Has No Sound

What causes it: The audio VAE (minimax_h3_audio_vae_fp32.safetensors) isn't loaded, or it's connected to the wrong VAELoader node.

How to fix it: Check both VAELoader nodes — one should point to the video VAE, the other to the audio VAE. Confirm a VAEDecodeAudio node is present and connected through to SaveVideo.

Frequently Asked Questions

No. FastVideo is a fully distilled standalone checkpoint built by the FastVideo team (hao-ai-lab), converted for ComfyUI by Kijai. The Turbo LoRA is a separate community project by larryvrh that patches the base MiniMax H3 model at load time. They are different files from different teams, and they are not interchangeable.

The FastVideo checkpoint fits in 8GB of VRAM for most settings, since the weights themselves dominate memory use rather than the resolution or duration you pick. The heaviest setting the model officially supports — 1344×768 at 5 seconds — is the point where you'd first risk running out of memory on an 8GB card.

Yes, within a range. The workflow ships with 8 steps as its default, which gives noticeably cleaner motion than 4 steps for a modest time cost. Going much above 12 to 14 steps stops helping quality much, since the model was distilled specifically to work in a small number of steps.

This guide covers image-to-video, which is the mode the current FastVideo checkpoint is built around. Text-to-video generation is supported by the underlying model, but first-and-last-frame control and full reference-to-video are not yet available in this distilled version.

The original weights are published on Hugging Face under FastVideo's FastH3 collection. The ComfyUI-ready version used in this guide is Kijai's converted int8 file, linked directly in the What You Need section above.

No. FastVideo is a research project from hao-ai-lab, a video-generation acceleration lab, built independently on top of the MiniMax H3 base model that MiniMax itself released. Kijai then converted FastVideo's checkpoint into a format ComfyUI can load natively.

What to Do Next

Run your first FastVideo generation at the default settings.

Download the workflow above, keep the defaults — 8 steps, 1344×768, 5 seconds — for your first clip, and compare your own generation time against the numbers in this guide before changing anything.

Published: 2026-09-04 · Last updated: 2026-09-04· Workflow structure verified against Kijai's ComfyUI-native MiniMax H3 FastVideo checkpoint.

Discussion

Join the discussion

Sign in to leave a comment or reply

💬

No comments yet

Be the first to share your thoughts!