Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Use the comfyui minimax h3 workflow in ComfyUI to generate open-weight videos from text, images, or references, with native stereo audio at up to 2K 24fps.
All Tools
Discover our comprehensive AI-powered animation toolkit

Suno AI Music
Create Professional Music with AI

Seedance2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1

Kling3.0
Next-Gen AI Video Generator

Nano Pro
Advanced AI Image Generator
Sora2 Video Generator
Advanced AI Video Generator for High-Quality Videos
Kling Motion Control
Turn reference images into amazing motion videos in minutes

Moive Maker
Turn Ideas into Stunning Movies with AI
Reasons to Run the comfyui minimax h3 Pipeline in ComfyUI
The comfyui minimax h3 workflow brings MiniMax H3, an open-weight omni-modal model, directly into ComfyUI. It processes text, images, video, and audio in one unified context, producing clips with native stereo audio—dialogue, effects, and music generated together. Output can reach 2K at 24fps for up to 15 seconds, and every node parameter stays adjustable.
- Built-In Synchronized AudioVoice, sound effects, and music render alongside the visuals in a single MP4, perfectly time-aligned through the comfyui minimax h3 workflow.
- Full Local ControlWith the comfyui minimax h3 nodes, you can run the model on your own hardware and tweak resolution, duration, and diffusion parameters without API limits.
- Diverse Reference SupportFeed in text, images, video, and audio simultaneously to lock down a character's look, style, motion, camera angles, or voice—all handled by the comfyui minimax h3 node setup.
Three Steps to Generate with the comfyui minimax h3 Workflow
Follow this quick guide to produce open-weight videos with native audio using the comfyui minimax h3 workflow in ComfyUI.
Core Capabilities of the comfyui minimax h3 Node Workflow
The comfyui minimax h3 setup includes three native ComfyUI templates, open-weight multimodal generation, built-in stereo audio, reference-guided control, and optional Sage Attention acceleration—everything you need for local video production.
Three Ready-Made Templates
The comfyui minimax h3 template pack covers text-to-video, image-to-video, and reference-to-video, with each example demonstrating a different generation mode.
One Model, Many Input Types
The comfyui minimax h3 model understands text, images, video, and audio as a single shared context, letting you blend multiple reference types in a single run.
Strong Reference Following
Fix a character's identity, style, movement, camera motion, or voice by providing up to 9 images, 3 videos, and 3 audio clips via the comfyui minimax h3 R2V node.
Sharp Text and Brand Output
Spelled words and logos appear cleanly with the comfyui minimax h3 model, while natural-language instructions define how references relate to the final video.
Faster Generation with Sage Attention
Add the Patch Sage Attention KJ node to the comfyui minimax h3 pipeline and roughly double your generation speed with minimal quality loss.
Flexible Resolution and Timing
The resolution selector in the comfyui minimax h3 workflow calculates width and height from aspect ratio and megapixels, snapping to the model's 32-multiple grid and 17-frame-per-block duration at 24fps.
Common Questions About the comfyui minimax h3 Workflow
Find quick answers on setting up and running the MiniMax H3 model inside ComfyUI with audio and video generation.
How does the comfyui minimax h3 workflow work?
It uses ComfyUI's native integration of MiniMax H3, an open-weight omni-modal model. The comfyui minimax h3 workflow generates video with synced stereo audio from text, images, video, and audio references in one forward pass.
What video quality can I expect?
You can get up to 2K resolution at 24fps for around 15 seconds. The native canvas uses a 768px short edge, capped at 768x1344 pixels and rounded to a multiple of 32.
What generation modes are included?
The comfyui minimax h3 templates include text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks character, style, motion, camera, or voice.
Does it really produce audio?
Yes. The comfyui minimax h3 model outputs native stereo audio—voice, effects, and music—modeled together with the visuals and saved in a single MP4 file.
How do I start generating?
Update ComfyUI to 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 workflow, and follow the prompt to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Can generation be sped up?
Yes—install SageAttention and the KJNodes custom nodes, then add a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to double the speed.
Start Creating with the comfyui minimax h3 Node Pipeline Today
Generate open-weight, audio-synced videos locally in ComfyUI using the comfyui minimax h3 workflow. Supports text-to-video, image-to-video, and reference-to-video modes.
