New: create a free account and get 14 days ad-free.Sign up free
EarnGeniX
Skip to main content

Workflow · ComfyUI · Updated July 2026

Wan Video Upscale in ComfyUI: Full Workflow Guide

A two-stage pipeline that redraws your video with Wan's diffusion model and a FusionX LoRA stack, then upscales and smooths it with RealESRGAN and RIFE — full node setup, settings, and fixes for the errors that stall most runs.

12 GB

Min. VRAM

RTX 4090

Tested on

Intermediate

Skill level

6

LoRAs stacked

By Earngenix Team · · Tested on RTX 4090

⚡ Quick Answer

To upscale a video with Wan in ComfyUI, you run it through two stages: first Wan 2.1's 14B model re-renders the video at low resolution with a stack of FusionX LoRAs to fix detail and motion, then a RealESRGAN model and RIFE frame interpolation scale it up to full resolution and smooth, high frame-rate output. Minimum 12 GB VRAM, with Block Swap for 8 GB cards.

Most "upscale" workflows just resize a video and sharpen the pixels. This one is different — Wan's diffusion model actually redraws the video at a low resolution first, fixing blur, compression artifacts, and warped motion, before a classic upscaler blows it up to full size. That's why it's called an enhance-and-upscale pipeline, not a plain upscale.

This guide walks through the exact wan video upscale comfyui workflow: which nodes to connect, which LoRAs to stack, and the settings that keep it from running out of memory.

How Does Wan Video Upscale Work in ComfyUI?

The pipeline has two separate stages that most tutorials online mix up:

  1. Enhance pass (Wan diffusion). Your source video is loaded, resized down to a small working resolution (this workflow uses 512×896), and fed into WanVideoSampler with a low denoise value. The Wan model doesn't generate a new video from scratch — it redraws the existing frames, using the LoRA stack to sharpen detail and stabilize motion.
  2. Upscale pass (RealESRGAN + RIFE). The enhanced frames go through ImageUpscaleWithModel using RealESRGAN_x2 to double the resolution, then ImageScale stretches it to your final target (2880×2160 in this build), and RIFE VFI doubles the frame rate (24fps → 48fps) so motion looks smooth instead of choppy.
Tip: The WanVideoSampler's denoise value controls how much the video changes. 0.3–0.5 keeps the original style and just cleans it up. 0.7–0.8 adds noticeable creative changes. Above 0.85 with no prompt, Wan starts generating unpredictable new content instead of enhancing your footage — stay under 0.5 if you just want a cleaner, sharper version of your original video.

What You Need Before You Start

Grab the workflow JSON first so you can follow along with the real node graph while you download the models below.

Download Workflow JSON
Tested on
RTX 4090 (24 GB)ComfyUI portable, Windows
Minimum VRAM
12 GB8 GB possible with Block Swap

Custom Nodes

Models

FileFolderSource
Wan2_1-T2V-14B_fp8_e4m3fn.safetensorsmodels/diffusion_modelsDownload ↗
Wan2_1_VAE_bf16.safetensorsmodels/vaeDownload ↗
umt5-xxl-enc-bf16.safetensorsmodels/text_encodersDownload ↗
RealESRGAN_x2.pthmodels/upscale_modelsDownload ↗
rife49.pthmodels/rifeDownload ↗

LoRA Stack

All six go in ComfyUI/models/loras:

LoRAFileStrengthSource
MPS RewardsWan2.1-Fun-14B-InP-MPS.safetensors1.0HuggingFace ↗
AccVidWan21_AccVid_T2V_14B_lora_rank32_fp16.safetensors1.0HuggingFace ↗
MoviiGenWan21_T2V_14B_MoviiGen_lora_rank32_fp16.safetensors1.0HuggingFace ↗
LightX2V FastWan21_T2V_14B_lightx2v_cfg_step_distill_lora_rank32.safetensors1.0HuggingFace ↗
Realism BoostWan14B_RealismBoost.safetensors0.5HuggingFace ↗
Detail EnhancerDetailEnhancerV1.safetensors0.5HuggingFace ↗
Warning: Six LoRAs stacked at once is heavy. If you're on 12 GB VRAM or less, drop MoviiGen and AccVid first — Realism Boost and Detail Enhancer matter most for upscale quality.

Before / After: What the Full Pipeline Produces

Here's the same clip run through the full pipeline above — Wan enhance pass, RealESRGAN upscale, RIFE interpolation — so you can see what to expect before running it yourself.

🎬 Before / After — Same Clip, Full Pipeline

Split-screen: source clip on the left, output after the Wan enhance pass, RealESRGAN upscale, and RIFE frame interpolation on the right.

◀ Before (source)After (upscaled) ▶

🎬 Before / After — Same Clip, Full Pipeline

Split-screen: source clip on the left, output after the Wan enhance pass, RealESRGAN upscale, and RIFE frame interpolation on the right.

◀ Before (source)After (upscaled) ▶
// Fix this part

Step-by-Step Setup

1

Load your source video

Add a VHS_LoadVideo node (from Video Helper Suite). Click choose video to upload and select your file. Leave frame_load_cap at 0 to process the whole clip, or set a number to test on a short clip first.

VHS_LoadVideo node with a video loaded and the preview thumbnail visible🔍 Click to zoom
VHS_LoadVideo with a clip loaded and preview visible.
2

Resize the video down for the enhance pass

Connect the loaded video to an ImageResizeKJv2 node (from KJNodes). Set width and height to a size your GPU can handle — this workflow uses 512x896. Set the resize method to nearest-exact and crop mode to center.

This step exists because the Wan diffusion pass is expensive. Running it at full resolution would need far more VRAM than most GPUs have — you upscale to full size after the enhance pass, not before.

3

Load the Wan model, VAE, and text encoder

Add three nodes: WanVideoModelLoader (select Wan2_1-T2V-14B_fp8_e4m3fn.safetensors, base_precision fp16_fast, attention_mode sdpa), WanVideoVAELoader (Wan2_1_VAE_bf16.safetensors), and LoadWanVideoT5TextEncoder (umt5-xxl-enc-bf16.safetensors).

Tip: If your GPU doesn't support Torch Compile, keep attention_mode set to sdpa and bypass the WanVideoTorchCompileSettings node entirely. This workflow ships with it disabled by default for exactly this reason.
4

Build the LoRA stack

Add six WanVideoLoraSelect nodes and chain them together (each one's prev_lora output feeds the next one's prev_lora input). Set each LoRA and strength as listed in the table above. Connect the final chained output into the WanVideoModelLoader's LoRA input.

Tip: The LightX2V Fast LoRA lets you drop your sampler steps as low as 2 if you raise its strength to 1.0 — useful for quick test renders before committing to a full pass.
Six WanVideoLoraSelect nodes chained together feeding into the model loader🔍 Click to zoom
The full six-LoRA chain feeding into WanVideoModelLoader.
5

Set up VRAM management

Add a WanVideoBlockSwap node and connect it to the model loader. Start with it bypassed. If you hit an out-of-memory error in step 7, enable it and raise the swap block count gradually (up to 40) until generation completes without crashing.

Warning: Block Swap trades speed for memory. Only enable it if you're actually running out of VRAM — it will noticeably slow down every render.
6

Configure the sampler

Add a WanVideoSampler node. Set steps to 4, cfg to 1.0, denoise to 0.5 (for cleanup without changing the video's content), and scheduler to flowmatch_causvid.

Leave the prompt empty in WanVideoTextEncode if you just want cleanup with no style changes. If your video is longer than 81 frames, add a WanVideoContextOptions node set to uniform_standard with context frames at 81 — this splits long videos into chunks so you don't run out of memory.

Warning: If you get an error mentioning "FlowMatch," switch the scheduler from flowmatch_causvid to unipc (or dpm++_sde / beta) instead.
7

Run the enhance pass and check the output

Click Queue Prompt. Generation time depends on your video length and GPU — expect several minutes for a short clip on a 24 GB card. Once done, a WanVideoDecode node converts the result back to viewable frames, and a VHS_VideoCombine node saves it as an MP4.

The enhanced output next to the original in a side-by-side compare🔍 Click to zoom
The ImageConcatMulti side-by-side compare of original vs. enhanced pass.
8

Upscale and interpolate

Feed the enhanced frames into UpscaleModelLoader (set to RealESRGAN_x2.pth) → ImageUpscaleWithModel → ImageScale (set to your target resolution, e.g. 2880x2160, method lanczos) → RIFE VFI (rife49.pth, multiplier 2 to double your frame rate).

9

Export the final video

Connect the RIFE output to a final VHS_VideoCombine node. Set frame_rate to double your source (e.g. 24 → 48 if you used a multiplier of 2 in RIFE). Click Queue Prompt to render the finished, upscaled video.

Troubleshooting

"CUDA out of memory" during the enhance pass

Your working resolution or LoRA stack is too heavy for your VRAM.

  1. Lower the resolution in ImageResizeKJv2 (try 384x672 instead of 512x896).
  2. Enable WanVideoBlockSwap and raise the swap block count gradually.
  3. Drop MoviiGen and AccVid from the LoRA chain — they add load without changing upscale quality much.

Error message contains "FlowMatch"

Your ComfyUI-WanVideoWrapper version doesn't support the flowmatch_causvid scheduler as configured.

  1. Open WanVideoSampler.
  2. Change scheduler from flowmatch_causvid to unipc.
  3. Re-run. dpm++_sde or beta also work if unipc gives soft results.

Output video looks smeared or overly changed from the original

Your denoise value is too high for a cleanup pass.

  1. Open WanVideoSampler and lower denoise to between 0.3 and 0.5.
  2. Make sure WanVideoTextEncode has an empty prompt — a prompt at high denoise pushes Wan toward generating new content instead of preserving your original.

Frequently Asked Questions

Yes — swap the checkpoint in WanVideoModelLoader for a Wan 2.2 model and check that your LoRAs are built for the same version, since Wan 2.1 and 2.2 LoRAs are not always cross-compatible. The node structure and settings stay the same.

The Wan diffusion pass is what fixes blur and compression artifacts, and running diffusion at full resolution needs far more VRAM than most consumer GPUs have. Processing small, then upscaling with RealESRGAN afterward, gets you the quality benefit without the memory cost.

On an RTX 4090, a short clip at 512x896 with 4 sampler steps takes a few minutes for the enhance pass, plus additional time for the RealESRGAN upscale and RIFE interpolation. Longer videos or lower-end GPUs will take considerably longer — use WanVideoContextOptions to chunk anything over 81 frames.

No. Realism Boost and Detail Enhancer have the biggest visible impact on upscale quality. MPS Rewards, AccVid, MoviiGen, and LightX2V mainly affect motion quality and render speed — drop them first if you're short on VRAM.

Yes — this pipeline works the same way on Wan-, LTX-, or other AI-generated clips. It's especially useful for cleaning up low-resolution AI video output before final export.

A plain ESRGAN upscale only enlarges existing pixels — it can't fix blur or motion artifacts already baked into the video. This workflow's Wan diffusion pass redraws the frames first, so the ESRGAN upscale that follows has cleaner source material to work with.

What to Do Next

Download the workflow and run it on a short test clip first.

Run it once with the default settings (denoise 0.5, no prompt) before processing a full video — once you confirm it runs on your GPU, adjust the resolution and LoRA strengths to match your hardware.

Published: 2026-07-30 · Last updated: 2026-07-30 · Tested on RTX 4090

Discussion

Join the discussion

Sign in to leave a comment or reply

💬

No comments yet

Be the first to share your thoughts!