New: create a free account and get 14 days ad-free.Sign up free
EarnGeniX
Skip to main content

Workflow Blog · All Levels · Updated July 2026

How to Make a Full AI Music Video in ComfyUI (Free, Local)

Song first, then a consistent AI character, then lip-synced video with LTX-2.3, then a final edit — the exact pipeline, tested on real hardware, no subscriptions required.

$0

Cost

RTX 4090

GPU tested

40s

Max clip length

All levels

Skill level

By Earngenix Team · · Tested on ComfyUI, RTX 4090

⚡ Quick Answer

To make a full AI music video in ComfyUI, generate your song first (Suno or the free ACE-Step 1.5), then generate a consistent AI character with Krea2 based on that song, then feed the character images and the audio into the LTX-2.3 workflow to produce lip-synced video clips, and stitch the clips together in a video editor. On an RTX 4090, each LTX-2.3 clip can run up to 40 seconds.

Making an AI music video isn't one generation — it's four separate steps chained together: song, character, synced video, edit. Get the order wrong (starting with the character before the song exists) and your character's mood, outfit, and performance style won't match the track. Here's the exact AI music video ComfyUI pipeline, tested end to end.

The full AI music video made with this exact pipeline.

What You Need Before You Start

Tested on
RTX 409024 GB VRAM
Minimum VRAM (LTX-2.3)
8 GBFP8 / GGUF builds available
Max clip length (RTX 4090)
40 secper LTX-2.3 generation

🅛🅣🅧 Gemma-3 Model Loader (pick one, based on your VRAM)

LTXV Audio Text Encoder Loader

Download the LTX-2.3 Audio-Sync Workflow (JSON)

Other tools used across this pipeline: Krea2 for character image generation (free/local via ComfyUI), ACE-Step 1.5 (free, local) or Suno (paid, cloud) for the song, Claude or ChatGPT for writing prompts, and any video editor to combine clips — this guide uses DaVinci Resolve (free).

1. Generate the Song First

Start with the song, not the character. The song's mood, tempo, and lyrics decide what your character looks like, wears, and how she performs — building the character first means guessing, and you'll usually have to redo it once the track is done.

  1. Generate your track with Suno, or with the free local alternative, ACE-Step 1.5, running inside ComfyUI.
  2. Export the final audio file once you're happy with the take.
ACE-Step 1.5 node graph in ComfyUI for free local song generation🔍 Click to zoom
Screenshot: ACE-Step 1.5 node graph in ComfyUI.

🎧 Listen to the Generated Song

Full setup for the free option: ACE-Step 1.5: the best free Suno AI alternative. If you want background on how Suno itself works before deciding which to use, see What Is Suno AI? A Complete Guide.

Tip: Keep a copy of the lyrics as plain text. You'll reuse specific lines in Step 3 when writing prompts for the singing performance shots.

2. Build a Consistent AI Character for the Song

With the song done, build the character who's going to perform it. This is two passes in Krea2: one to generate the character's look, and a second using Krea2's character consistency feature to produce multiple images of the same face for different shots in the video.

  1. Generate your initial character image (a singer) in Krea2, based on the mood of the song.
  2. Run that image through Krea2 character consistency to generate additional shots — different angles, outfits, or stage settings — while keeping the same face.
First Krea2-generated AI singer character image🔍 Click to zoom
Screenshot: the first Krea2-generated singer image.

Character Consistency Outputs (Same Face, Different Shots)

Character consistency output 1, same singer face
Character consistency output 2, same singer face
Character consistency output 3, same singer face
Warning: Skipping character consistency and generating each shot from scratch will give you a different-looking singer in almost every clip. LTX-2.3 will still lip-sync a mismatched face just fine — it won't look like the same person from clip to clip.

Full setup guides: Krea2 in ComfyUI: full workflow · Krea2 character consistency in ComfyUI.

3. Write Your Prompts with Claude or ChatGPT

Both the character image prompts in Step 2 and the video-motion prompts in Step 4 came from Claude and ChatGPT rather than being written from scratch by hand.

  • For text-to-image prompts: describe the singer's appearance, outfit, pose, and lighting, and ask the model to turn that into a single detailed Krea2 prompt.
  • For image-to-video prompts: describe the performance you want (camera movement, expression, energy level matching a specific point in the song) and ask for a prompt LTX-2.3 can use alongside the audio.
Tip: Paste a line or two of the actual lyrics into your prompt request so the AI-written prompt matches what the character is "singing" in that clip, not a generic performance description.

4. Turn the Character Into a Lip-Synced Video with LTX-2.3

LTX-2.3 in ComfyUI is not a single-purpose "lip sync" workflow — it's one model that covers four different generation modes in the same node setup:

  • Text-to-video — generate a clip from a prompt alone, no input image
  • Image-to-video — animate a single still image
  • First-and-last-frame — give it a starting and ending image and it fills in the motion between them
  • Image + audio-to-video — the mode used here, which takes a character image and an audio file and generates lip-synced motion matched to the audio
Pick the mode based on what you need for that specific clip. For a singing performance shot, use image + audio-to-video with one of your Step 2 character images and the song from Step 1.
LTX-2.3 audio-sync workflow loaded in ComfyUI🔍 Click to zoom
Screenshot: the LTX-2.3 workflow loaded in ComfyUI, ready to run.
  1. Load a character image with LoadImage, and load the song (or a trimmed segment of it) with LoadAudio.
  2. Connect both to the LTX-2.3 audio-video generator node, along with your motion prompt from Step 3.
  3. Set your resolution and frame rate, then run the generation. A progress bar appears; generation time depends on clip length and your GPU.
  4. The output is an MP4 with the character's mouth and expression moving in time with the audio.

Output Examples — Lip Sync in Action

Warning: On an RTX 4090, the maximum reliable generation length per clip is 40 seconds. Longer requests either fail or degrade in quality — generate in 40-second (or shorter) segments and stitch them in your editor instead.

Broader LTX-2.3 setup reference (all four modes): LTX-2.3 in ComfyUI: full workflow guide. Official node repo: ComfyUI-LTXVideo on GitHub.

5. Edit the Clips Into a Full Music Video

LTX-2.3 outputs individual clips, not a finished video — you still need an editor to arrange them in order, trim transitions, and sync clip boundaries to the song structure.

  1. Import all generated clips and the full song into your editor. This guide uses DaVinci Resolve, which is free.
  2. Line up each clip against the section of the song it was generated for (verse, chorus, bridge).
  3. Trim and crossfade between clips at natural cut points in the music.
  4. Export the final video.
Clips and song lined up on the timeline in DaVinci Resolve🔍 Click to zoom
Screenshot: clips lined up against the song structure in DaVinci Resolve.

Troubleshooting

"This node type does not exist" (red node in the LTX-2.3 workflow)

The workflow uses a custom node you haven't installed.

  1. Open Manager in the top menu.
  2. Click Install Missing Custom Nodes.
  3. Restart ComfyUI completely.

The character's mouth barely moves, or doesn't sync to the audio

This usually means the audio wasn't connected to the LTX-2.3 audio-video node, or a plain image-to-video mode was used instead of image + audio-to-video.

  1. Confirm LoadAudio is connected directly into the LTX-2.3 generator node, not just present on the canvas.
  2. Confirm you're running the audio-sync mode of the workflow, not the plain image-to-video mode.

Generation fails or crashes on longer clips

LTX-2.3 clips longer than your GPU can handle (around 40 seconds on an RTX 4090) will fail or run out of memory.

  1. Shorten the clip length and generate in segments.
  2. Stitch the segments together in your video editor instead of generating one long clip.

Frequently Asked Questions

Yes. The song sets the mood, tempo, and lyrics that the character's look and performance are built around. Generating the character first usually means redoing it once the track is finished.

Yes. ACE-Step 1.5 is a free, local text-to-music model that runs inside ComfyUI and works as a Suno alternative for this pipeline.

No. LTX-2.3 is a single workflow that also handles text-to-video, image-to-video, and first-and-last-frame generation. Lip-synced audio-to-video is one mode among several — use whichever mode fits the clip you're making.

On an RTX 4090, the maximum reliable length per clip is 40 seconds. Generate longer videos as multiple shorter clips and combine them in your editor.

Yes — the AI models just speed up prompt writing. Any well-detailed prompt describing the character's appearance or the performance and camera movement works the same way.

The full pipeline was tested on an RTX 4090 (24 GB VRAM). LTX-2.3 has lower-VRAM FP8 and GGUF builds available for smaller GPUs, though clip length and generation speed will be more limited.

What to Do Next

Download the workflow and run it with one existing character image.

Drag the audio-sync workflow JSON into ComfyUI and run it once with one of your character images and a short clip of your song to confirm everything connects correctly, then move on to generating your full set of clips.

Published: 2026-07-27 · Last updated: 2026-07-27 · Tested on ComfyUI, RTX 4090 (24 GB VRAM)

Discussion

Join the discussion

Sign in to leave a comment or reply

💬

No comments yet

Be the first to share your thoughts!