Start the comfyui minimax h3 Workflow
Describe your scene or attach a reference, then run the comfyui minimax h3 workflow for clips with synced audio.
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Use the comfyui minimax h3 workflow in ComfyUI to generate open-weight videos from text, images, or references, with native stereo audio at up to 2K 24fps.

All Tools

Discover our comprehensive AI-powered animation toolkit

Reasons to Run the comfyui minimax h3 Pipeline in ComfyUI

The comfyui minimax h3 workflow brings MiniMax H3, an open-weight omni-modal model, directly into ComfyUI. It processes text, images, video, and audio in one unified context, producing clips with native stereo audio—dialogue, effects, and music generated together. Output can reach 2K at 24fps for up to 15 seconds, and every node parameter stays adjustable.

  • Built-In Synchronized Audio
    Voice, sound effects, and music render alongside the visuals in a single MP4, perfectly time-aligned through the comfyui minimax h3 workflow.
  • Full Local Control
    With the comfyui minimax h3 nodes, you can run the model on your own hardware and tweak resolution, duration, and diffusion parameters without API limits.
  • Diverse Reference Support
    Feed in text, images, video, and audio simultaneously to lock down a character's look, style, motion, camera angles, or voice—all handled by the comfyui minimax h3 node setup.

Three Steps to Generate with the comfyui minimax h3 Workflow

Follow this quick guide to produce open-weight videos with native audio using the comfyui minimax h3 workflow in ComfyUI.

Core Capabilities of the comfyui minimax h3 Node Workflow

The comfyui minimax h3 setup includes three native ComfyUI templates, open-weight multimodal generation, built-in stereo audio, reference-guided control, and optional Sage Attention acceleration—everything you need for local video production.

Three Ready-Made Templates

The comfyui minimax h3 template pack covers text-to-video, image-to-video, and reference-to-video, with each example demonstrating a different generation mode.

One Model, Many Input Types

The comfyui minimax h3 model understands text, images, video, and audio as a single shared context, letting you blend multiple reference types in a single run.

Strong Reference Following

Fix a character's identity, style, movement, camera motion, or voice by providing up to 9 images, 3 videos, and 3 audio clips via the comfyui minimax h3 R2V node.

Sharp Text and Brand Output

Spelled words and logos appear cleanly with the comfyui minimax h3 model, while natural-language instructions define how references relate to the final video.

Faster Generation with Sage Attention

Add the Patch Sage Attention KJ node to the comfyui minimax h3 pipeline and roughly double your generation speed with minimal quality loss.

Flexible Resolution and Timing

The resolution selector in the comfyui minimax h3 workflow calculates width and height from aspect ratio and megapixels, snapping to the model's 32-multiple grid and 17-frame-per-block duration at 24fps.

FAQ

Common Questions About the comfyui minimax h3 Workflow

Find quick answers on setting up and running the MiniMax H3 model inside ComfyUI with audio and video generation.

1

How does the comfyui minimax h3 workflow work?

It uses ComfyUI's native integration of MiniMax H3, an open-weight omni-modal model. The comfyui minimax h3 workflow generates video with synced stereo audio from text, images, video, and audio references in one forward pass.

2

What video quality can I expect?

You can get up to 2K resolution at 24fps for around 15 seconds. The native canvas uses a 768px short edge, capped at 768x1344 pixels and rounded to a multiple of 32.

3

What generation modes are included?

The comfyui minimax h3 templates include text-to-video (T2V), image-to-video (I2V) with optional first/last-frame control, and reference-to-video (R2V) that locks character, style, motion, camera, or voice.

4

Does it really produce audio?

Yes. The comfyui minimax h3 model outputs native stereo audio—voice, effects, and music—modeled together with the visuals and saved in a single MP4 file.

5

How do I start generating?

Update ComfyUI to 0.30.0 or later, open Template Library > Video, choose a comfyui minimax h3 workflow, and follow the prompt to download models from the Hugging Face Comfy-Org/MiniMax-H3 repository.

6

Can generation be sped up?

Yes—install SageAttention and the KJNodes custom nodes, then add a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow to double the speed.

Start Creating with the comfyui minimax h3 Node Pipeline Today

Generate open-weight, audio-synced videos locally in ComfyUI using the comfyui minimax h3 workflow. Supports text-to-video, image-to-video, and reference-to-video modes.