Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Bring MiniMax H3 into ComfyUI to build open-weight clips with synchronized stereo audio — feed text, imagery, or reference footage and render up to 2K at 24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.
What Makes the MiniMax H3 ComfyUI Workflow Stand Out
Integrating MiniMax H3 into ComfyUI grants you an omni-modal generation engine with fully open weights. The model interprets text, stills, motion, and sound together in a single context, synthesizing video with synchronized stereo audio — speech, effects, and music all produced in the same pass. You can render clips up to 2K at 24fps for roughly 15 seconds while tuning every node parameter by hand.
- Built-In Stereo SoundSpeech, sound effects, and musical score are rendered together with the frames in a single MP4, perfectly aligned by the MiniMax H3 integration inside ComfyUI.
- Full Parameter FreedomExecute MiniMax H3 on your own hardware and tune resolution, clip length, and every diffusion setting — with no external API quotas or request limits.
- Rich Reference InputsFeed text, stills, video, and audio cues into one generation run — the MiniMax H3 ComfyUI nodes lock down a character's look, style, motion, camera path, or voice.
Getting Started with MiniMax H3 in ComfyUI
Follow three straightforward steps to begin generating open-weight videos with synced audio through the MiniMax H3 node set in ComfyUI.
Key Capabilities of the MiniMax H3 ComfyUI Pipeline
From bundled ComfyUI templates to open-weight multimodal generation with native stereo sound, reference-based conditioning, and optional Sage Attention acceleration — the MiniMax H3 integration adds a complete offline video studio to ComfyUI.
Bundled Template Library
The MiniMax H3 pack for ComfyUI ships with text-to-video, image-to-video, and reference-to-video node examples, each matched to a distinct production workflow.
Unified Multi-Modal Context
The MiniMax H3 engine interprets text, pictures, motion, and sound in one combined context, letting you mix all reference types inside a single synthesis pass.
Reference-Guided Synthesis
Lock in a character's identity, artistic style, motion pattern, camera angle, or voice from source materials — up to 9 images, 3 videos, and 3 audio clips via the MiniMax H3 R2V node.
Clear Text and Brand Rendering
On-screen words and brand logos stay crisp with the MiniMax H3 model, and its instruction following describes reference relationships in natural language.
Sage Attention Acceleration
Add a Patch Sage Attention KJ node to the MiniMax H3 ComfyUI workflow and nearly double rendering speed with negligible quality impact.
Smart Resolution & Duration Grid
The Resolution Selector derives width and height from aspect ratio plus megapixel target, snapped to the model's 32-multiple grid and 17-frame block duration at 24fps.
MiniMax H3 ComfyUI — FAQ
Straight answers to common questions about running the MiniMax H3 model inside ComfyUI.
What exactly is the MiniMax H3 workflow in ComfyUI?
It is ComfyUI's official integration of MiniMax H3, an open-weight omni-modal generation model from MiniMax. The pipeline creates video with native stereo audio from text, stills, footage, and sound references in one forward pass.
What output quality can I achieve?
The MiniMax H3 ComfyUI workflow renders clips up to 2K resolution at 24fps for about 15 seconds. Its native canvas uses a 768px short edge, capped at 768x1344 pixels and rounded to multiples of 32.
Which generation modes are included?
Three templates are packaged: text-to-video (T2V), image-to-video (I2V) with optional first/last-frame conditioning, and reference-to-video (R2V) that fixes character, style, motion, camera, or voice.
Does it produce audio alongside the video?
Yes — MiniMax H3 inside ComfyUI generates native stereo audio including dialogue, effects, and music, all modeled together with the visuals in one pass and wrapped into a single synced MP4.
How do I get everything up and running?
Install ComfyUI 0.30.0 or later, open Template Library > Video, select a MiniMax H3 template, and follow the on-screen instructions to fetch models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Is there a way to speed up generation?
Yes — set up SageAttention and the KJNodes custom nodes, then place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the MiniMax H3 pipeline to roughly double generation speed.
Start Creating Videos with MiniMax H3 in ComfyUI Today
Run MiniMax H3 locally from ComfyUI with open weights and stereo audio at full parameter control — T2V, I2V, and R2V node templates are just a click away.
