New: create a free account and get 14 days ad-free.Sign up free
EarnGeniX
Skip to main content

Tutorial · Updated September 2026

ComfyUI Low VRAM: The Setting Nobody Tells Beginners About

It's your system RAM, not your GPU. Real settings for 4GB, 6GB, 8GB and 12GB cards — and why the flag every guide tells you to add doesn't do anything anymore.

Free

Cost

4 GB

Min VRAM

Beginner

Skill level

~10 min

Fix time

🗺️ Part of the Free ComfyUI RoadmapLevel 1: Beginner FoundationsBeginnerStep: Troubleshooting Common Issues
View Full Roadmap

By Earngenix Team · · Tested on ComfyUI [0.26.0], [confirm GPU + RAM]

⚡ Quick Answer

If you have an NVIDIA GPU and a current copy of ComfyUI, the old --lowvram flag does nothing anymore — a newer system called Dynamic VRAM handles memory automatically, and it's on by default. This means your system RAM, not just your GPU's VRAM, decides how much you can actually generate. If you're on AMD, Mac, or running ComfyUI inside WSL, Dynamic VRAM isn't active by default, and the older rules and flags still apply to you.

If you searched for ComfyUI low VRAM, you've probably already found guides telling you to add --lowvram to your launch command. Here's the problem: on a current install with an NVIDIA card, that flag doesn't do anything. The official source code says so directly — its own help text reads, "Doesn't do anything if dynamic vram is enabled." And dynamic VRAM is enabled by default.

This guide explains what actually controls memory usage today, what you can realistically run on 4 GB, 6 GB, 8 GB, and 12 GB cards, and what to do when something still runs out of memory.

Why ComfyUI Low VRAM Advice From Last Year Is Probably Wrong Now

ComfyUI used to work like this: before running a workflow, it would guess how big your model was and how much VRAM (short for Video RAM — this is the memory built into your graphics card, separate from your computer's regular system memory) you had, then pick a loading strategy. If the guess was wrong, you got an out-of-memory crash. The --lowvram flag existed to force a more careful, cautious loading strategy for people with small GPUs.

Sometime around early-to-mid 2026, ComfyUI replaced that system with something called Dynamic VRAM. Instead of guessing upfront, it watches your GPU's memory in real time while the workflow runs, and moves pieces of the model between your GPU, your regular system RAM, and even your hard drive as needed — automatically, without you setting anything.

On a current install with an NVIDIA graphics card, Dynamic VRAM turns on by itself. Once it's on, the old --lowvram flag genuinely has no effect, because the whole point of that flag was to trigger the old cautious-loading behavior, and that old system isn't running anymore.

Warning: This only applies with all three of the following true: (1) you have an NVIDIA graphics card, not AMD or Intel; (2) you're running ComfyUI on Windows, Mac, or native Linux — not inside WSL; (3) your ComfyUI install is fairly recent, with a Python library called PyTorch at version 2.8 or newer. If any of those three aren't true for you, ComfyUI falls back to its older memory system, and --lowvram still works exactly like it always has.

What You Need Before You Start

Before you touch any settings, check two things: what GPU brand you have, and roughly how much system RAM your computer has. Both affect which advice in this article applies to you.

  1. Check your GPU brand, VRAM, and system RAM. Open Task Manager (press Ctrl+Shift+Esc), click the Performance tab, then check GPU for your dedicated GPU memory (this is your VRAM) and Memory for your total system RAM. If you want the full click-by-click walkthrough with screenshots, our best GPU for ComfyUI guide covers it in detail, including the Mac version of this check.
  2. Update ComfyUI to a current version. Dynamic VRAM only exists in newer builds. If you haven't updated in a while, follow the steps in how to update ComfyUI before continuing — running an old version means none of the Dynamic VRAM settings below will even be present.
Task Manager Performance tab showing GPU dedicated memory next to system RAM🔍 Click to zoom
Checking your GPU's VRAM and your total system RAM in Task Manager.

You'll also see two file types mentioned in this guide: the checkpoint (this is the main AI model file — it controls what your generated images look like, usually saved as a safetensors file, a safer, faster-loading format than the older .ckpt format) and GGUF files (a compressed version of a model that trades a small amount of quality for a much smaller file size — these end in .gguf).

How Much VRAM (and RAM) Do You Actually Need? Picking Your Tier

VRAM size is still the single biggest factor in what you can run, even with Dynamic VRAM helping in the background. Find your card below and use the model suggested for that tier.

VRAMModel FileNotes
4 GBv1-5-pruned-emaonly-fp16.safetensorsSD 1.5 only, 512×512. VRAM is still the hard limit at this size.
6 GBz_image_turbo-Q4_K_M.ggufNeeds the ComfyUI-GGUF custom node. Keep an eye on system RAM too.
8 GBz_image_turbo-Q8_0.ggufOr FP8 builds. SDXL runs comfortably here without quantization.
12 GBflux1-dev (fp8 / GGUF)Full SDXL, and 2–3 stacked LoRAs without trouble.

4 GB VRAM — What You Can Realistically Run

Four gigabytes is ComfyUI's practical floor. At this size, VRAM itself is still the hard limit that Dynamic VRAM cannot get around — there simply isn't enough room on the card for anything bigger, regardless of how much system RAM you have.

Model: Stable Diffusion 1.5, using the file v1-5-pruned-emaonly-fp16.safetensors. Resolution: stay at 512×512 — going higher will very likely cause an out-of-memory error on a 4 GB card. Place the file in your ComfyUI/models/checkpoints/ folder.

6 GB VRAM — What You Can Realistically Run

At 6 GB, you can run a modern, good-quality model, but you need the compressed GGUF version, not the full-size original.

Model: Z-Image Turbo, using the GGUF file z_image_turbo-Q4_K_M.gguf (about 4.98 GB). Note the underscores in the filename — some other sites write it with hyphens, which is incorrect.

Tip: You'll need the ComfyUI-GGUF custom node (made by developer city96) before this file will load. Without it, ComfyUI won't recognize the .gguf format. On 6 GB cards, running out of regular system RAM — not VRAM — is a common cause of crashes with GGUF models. Aim for at least 16 GB of system RAM as a baseline.

8 GB VRAM — What You Can Realistically Run

Eight gigabytes is a comfortable, popular tier. You can run SDXL at full quality, or step up to a higher-quality Z-Image Turbo file.

Model: Z-Image Turbo, using z_image_turbo-Q8_0.gguf (about 7.22 GB), or an FP8 version of the same model if you'd rather skip the GGUF node entirely. FP8 stores a model's numbers using less precision — it shrinks the file and speeds it up, with a small, usually unnoticeable, quality trade-off. SDXL checkpoints also run comfortably here without needing any compressed version at all.

12 GB VRAM — What You Can Realistically Run

At 12 GB, most current models run well, including Flux.1-dev in its FP8 or GGUF form, full SDXL, and stacking two or three LoRAs (small add-on files that change a model's style) on top of a base model without trouble.

Tip: None of these four tiers match your card, or you're outgrowing yours? The settings on this page get more out of the GPU you already have — they won't turn a 4 GB card into a 12 GB one. If you're ready to upgrade, our best GPU for ComfyUI in 2026 guide sorts cards by VRAM tier instead of speed, which is the number that actually decides what you can run.

The System RAM Setting That Actually Decides What Runs (Dynamic VRAM Explained)

Here's the part most guides skip. If Dynamic VRAM is active on your system (NVIDIA card, not WSL, current PyTorch — see the checklist above), it doesn't just manage your GPU's memory. It streams data between four places, from fastest to slowest: your VRAM, then a special "reserved" section of system RAM, then regular system RAM, then your hard drive.

What this means in plain terms: even if a model is technically too big for your VRAM, ComfyUI can often still run it by temporarily parking the parts it isn't using in your system RAM. The catch is that if your system RAM is also low, or your hard drive is a slow, older type, this streaming process either fails outright or becomes extremely slow.

This is why two people with the exact same graphics card can have completely different experiences — one has 64 GB of system RAM and everything runs smoothly, while the other has 16 GB and keeps hitting errors or waiting far longer than expected.

Settings Worth Knowing

These are all typed as extra text after main.py in your launch command:

  • --reserve-vram [number] — tells ComfyUI to always keep a certain amount of VRAM free for other programs, in gigabytes. Useful if you're also gaming or browsing while generating.
  • --disable-pinned-memory — occasionally fixes out-of-memory errors that don't make sense given how much RAM you have. This is a genuinely unusual-sounding fix, but multiple real user reports confirm turning this off stopped otherwise unexplainable crashes.
  • --fast-disk — if your models are stored on a fast NVMe solid-state drive (not a regular hard drive, and not a USB external drive), this can speed up the streaming process described above.
If you're on AMD, Mac, or WSL: none of this applies to you by default, because Dynamic VRAM isn't active. Instead, the traditional --lowvram flag (which forces text encoders to run on your CPU instead of your GPU, freeing up VRAM at the cost of some speed) is still your main tool, along with the tier guidance above.

How to Check Your Dynamic VRAM Settings and Launch ComfyUI Correctly

Follow these steps to find out whether Dynamic VRAM is active on your system, and how to add a setting if you need one.

  1. Find your ComfyUI launch file. If you installed the Desktop app, this is handled automatically and you'll need to add settings through the app's settings menu instead of a launch file. If you installed the Portable version, look for a file called run_nvidia_gpu.bat (Windows) in your main ComfyUI folder.
  2. Open the launch file in a plain text editor — right-click it and choose Edit, or open Notepad first and drag the file into it. Don't double-click it, or it will run instead of opening for editing.
  3. Look at the last line of the file. It will start with something like .\python_embeded\python.exe -s ComfyUI\main.py. This is the command that starts ComfyUI.
  4. Add your flag to the end of that line, with a space before it — for example, add --reserve-vram 1 if you want to reserve 1 GB of VRAM. Save the file when you're done.
  5. Launch ComfyUI normally by double-clicking the file you just edited. Watch the black terminal window that opens.
  6. Check for confirmation. You should see a line mentioning model loading and memory — this confirms Dynamic VRAM (or your added flag) is active. If you don't see any memory-related messages at all in the first few lines, something didn't apply correctly, and it's worth double-checking your edit for typos.
ComfyUI terminal window on startup showing a memory-related confirmation line🔍 Click to zoom
The console line that confirms your memory settings applied on startup.

Troubleshooting

Running into something not covered below? Our general ComfyUI troubleshooting guide covers the errors that show up across every workflow, not just this one.

"torch.OutOfMemoryError: Allocation on device"

What causes it: Your GPU ran out of VRAM partway through generating an image or video. This is the current, most common version of the out-of-memory error — an older ComfyUI version might instead show the longer message "Allocation on device 0 would exceed allowed memory. (out of memory)," which means the same thing.

  1. Lower your image resolution first — this has the single biggest impact on VRAM usage.
  2. Reduce your batch size (the number of images generated at once) to 1, if it isn't already.
  3. Switch to a smaller or more compressed version of your model — for example, move from an FP16 or FP8 checkpoint down to a GGUF version, using the tier guide above to pick the right one for your card.
  4. If you're on an NVIDIA card with Dynamic VRAM active, try adding --disable-pinned-memory to your launch command — this has fixed OOM errors for some users even when they had plenty of VRAM on paper.
Red torch.OutOfMemoryError: Allocation on device error box in ComfyUI🔍 Click to zoom
The current out-of-memory error box — start with the resolution and batch size fixes above.

--lowvram Isn't Doing Anything — Why, and What Actually Works Instead

What causes it: You're on an NVIDIA card with Dynamic VRAM active, which makes --lowvram a no-op by design, not a bug.

How to fix it: Stop relying on --lowvram and instead pick the right model size for your VRAM tier, using the guide above. If you're specifically trying to free up VRAM for other programs, use --reserve-vram instead. If you're seeing a completely different error — such as red boxes on nodes rather than an out-of-memory crash — that's usually a missing-nodes error instead, which has a different fix.

Generation Finishes but Takes 5–10x Longer Than It Should

What causes it: Under Dynamic VRAM, when a model is too big for your VRAM and your system RAM is also tight, ComfyUI keeps moving data back and forth between VRAM, RAM, and your hard drive during every single step of generation — sometimes called "spilling." It works, but it's slow.

  1. Check your available system RAM using the Task Manager steps from earlier in this article. If it's close to full, close other programs before generating.
  2. If your models are stored on a regular hard drive rather than an SSD, moving them to an SSD (ideally NVMe) will noticeably help.
  3. Try adding --fast-disk to your launch command if you're on a fast NVMe drive.
  4. As a last resort, drop down one VRAM tier in the model guide above — a smaller, well-fitted model will always be faster than a large one that's constantly spilling to RAM.

Once your settings match your hardware, a good next check is whether all your model files are actually sitting in the right folders — our checkpoint models guide covers exactly where each file type belongs.

Frequently Asked Questions

Yes, but with real limits. Stick to Stable Diffusion 1.5 at 512×512 resolution using the v1-5-pruned-emaonly-fp16.safetensors checkpoint. Newer, larger models like SDXL or Flux generally won't fit, even with Dynamic VRAM's help, because 4 GB is small enough that VRAM itself is still the hard limit.

Only if you're on AMD, Mac, or running ComfyUI inside WSL, or if you're on an older ComfyUI version without Dynamic VRAM. If you have an NVIDIA card on a current install, the flag is ignored, and picking the right model size for your VRAM tier matters more than any launch flag.

Sixteen gigabytes is a reasonable minimum for image generation on smaller VRAM tiers like 6 GB. If you plan to run larger models where Dynamic VRAM has to spill data from VRAM into system RAM, 32 GB is safer, and 64 GB removes most RAM-related slowdowns entirely.

The most likely reason is system RAM. Two people can have the exact same graphics card, but if one has significantly less system RAM, Dynamic VRAM has to spill data to a slower location more often, which can make generation several times slower even though the GPU itself is identical.

No. Dynamic VRAM is currently NVIDIA-only, and it also doesn't activate inside WSL. On AMD, Mac, or WSL, ComfyUI uses its older memory system, where your VRAM size remains the main hard limit and the --lowvram flag still functions the way older guides describe.

As a starting point: Q4_K_M for 6 GB cards, and Q8_0 for 8 GB cards, both using Z-Image Turbo as an example. Generally, a higher number after the Q means better quality but a larger file — pick the highest number that comfortably fits your VRAM tier.

What to Do Next

Explore what's new in ComfyUI for 2026.

Some of what changed there affects how Dynamic VRAM behaves, so it's worth a read once your settings are sorted. For a structured path through the rest of ComfyUI's beginner fundamentals, see the full roadmap.

What to Read Next

If your settings are sorted, the two guides worth reading next are how checkpoint models work, since picking the right file format matters just as much as your VRAM tier, and our missing nodes guide if a red error box shows up that isn't about memory. If you're finding your card just isn't enough no matter what you change, our GPU buying guide breaks down exactly how much VRAM to buy for the models you want to run.

— Written as a personal recommendation from the Earngenix team.

Published: 2026-09-29 · Last updated: 2026-09-29 · Settings verified against the official ComfyUI startup-flags documentation.

Discussion

Join the discussion

Sign in to leave a comment or reply

💬

No comments yet

Be the first to share your thoughts!