Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Turn prompts into clips with synced stereo audio in ComfyUI — the comfyui minimax h3 workflow runs open weights locally, up to 2K at 24fps. No API key needed.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Gemini Omni
Gemini Omni Video Generator

Nano Banana2
Best Image Generator
Why Creators Run comfyui minimax h3 Inside ComfyUI
Packed into ComfyUI as downloadable weights, the comfyui minimax h3 pipeline is built on MiniMax's omni-modal engine, so a single context can hold text, pictures, footage, and sound at once. Every clip arrives with its own stereo track — speech, effects, and score produced alongside the picture rather than dubbed in later. Results reach roughly 15 seconds at up to 2K and 24fps, and each setting stays editable through the node graph.
- Sound Built Into Every ClipSpeech, effects, and score are produced in the same pass as the picture, so each MP4 leaves the comfyui minimax h3 graph already mixed in stereo.
- Weights You Actually OwnBecause the comfyui minimax h3 model lives on your own machine, resolution, clip length, and sampling settings are yours to tune — nothing is metered or capped.
- Mix Any Reference TypeFeed prompts, stills, clips, and audio samples into a single run, and these nodes hold a face, a look, a camera path, or a voice steady across the output.
Three Steps to Your First comfyui minimax h3 Render
From install to export: three quick moves get the comfyui minimax h3 graph producing video and sound on your own hardware.
Capabilities Inside the comfyui minimax h3 Graph
Ready-made templates, multimodal inputs, built-in stereo tracks, reference locking, and an optional Sage Attention boost — together the comfyui minimax h3 nodes cover the whole pipeline from prompt to finished file on your desktop.
Three Templates Out of the Box
Text, image, and reference driven examples all ship inside the comfyui minimax h3 library, so every generation mode has a working starting graph.
One Shared Context
Words, pictures, footage, and sound are read side by side by the model, letting any mix of reference types shape the very same clip.
Lock What Matters
Hold a face, a look, a movement, a camera path, or a voice steady using source material — the comfyui minimax h3 R2V node accepts as many as 9 stills, 3 clips, and 3 audio files.
Legible Text and Logos
Lettering and logo marks come out crisp, and the comfyui minimax h3 model follows plain-language instructions about how references relate to one another.
Optional Speed Boost
Slot a Patch Sage Attention KJ node into the comfyui minimax h3 graph and renders finish in about half the time, with barely any drop in fidelity.
Precise Size and Length Controls
Set an aspect ratio and megapixel target, and the comfyui minimax h3 Resolution Selector works out width and height on the model's 32-step grid, with duration ticking in 17-frame blocks at 24fps.
comfyui minimax h3: Questions Answered
Straight answers about installing, running, and tuning the MiniMax H3 model inside ComfyUI.
What exactly is the comfyui minimax h3 workflow?
It is the official ComfyUI integration of MiniMax H3, an omni-modal generation model that MiniMax published as downloadable weights. Feeding it text, pictures, footage, or audio lets it produce a clip and its stereo soundtrack together in one pass.
How good is the output?
Clips from the comfyui minimax h3 graph run up to about 15 seconds at 2K and 24fps. The native canvas keeps a 768px short edge, tops out at 768x1344, and rounds dimensions to a multiple of 32.
Which generation modes ship with it?
Three examples are bundled in the comfyui minimax h3 library: text-to-video, image-to-video with optional first and last frame control, and reference-to-video for holding a character, style, motion, camera move, or voice in place.
Does the model produce sound too?
It does. Speech, effects, and music are all generated by the comfyui minimax h3 model alongside the visuals, then delivered already synchronized inside one MP4 file.
What do I need to get started?
Install ComfyUI 0.30.0 or newer, go to Template Library > Video, pick one of the comfyui minimax h3 examples, and accept the download prompt that pulls weights from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to make renders faster?
Yes. Add the SageAttention and KJNodes extensions, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in your graph — render times typically drop by around half.
Your Next Clip Starts in the comfyui minimax h3 Graph
Keep the comfyui minimax h3 model on your own machine, tweak every parameter, and let text, image, or reference prompts carry you from idea to finished video with sound — no subscription in the way.
