Check the license before the download
MiniMax released H3 under its own Community License, not Apache or MIT. The license grants local-use rights only in its "Applicable Territory" and excludes the United States, European Union, United Kingdom, and South Korea. Canada is not on that exclusion list, but readers elsewhere should check the current text before downloading or running the weights. Companies whose products generate more than US$20 million in annual revenue also need separate written authorization.
That caveat belongs first because ComfyUI makes the technical install easier than the legal one. A download button is not a universal permission slip.
What you need
Use ComfyUI 0.30.0 or later for the base H3 templates. The minimum sensible local package for text-to-video and first- or last-frame image-to-video is four files, about 42.5GB in total.
The 21GB minimax_h3_fl2va_pruned_int8_convrot.safetensors file goes in ComfyUI/models/diffusion_models/. The 15.7GB qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors text encoder goes in ComfyUI/models/text_encoders/. Put both minimax_h3_video_vae_fp16.safetensors, about 5.21GB, and minimax_h3_audio_vae_fp32.safetensors, about 605MB, in ComfyUI/models/vae/.
For reference-to-video, use the 21GB minimax_h3_ref2va_pruned_int8_convrot.safetensors diffusion file instead. The text encoder and both VAEs stay the same. Start with the pruned INT8 set; heavier checkpoints consume more storage and memory without making the first setup more informative.
Let the template do the wiring
Update ComfyUI, restart it, then open Template Library, choose Video, and select a MiniMax H3 workflow. The base templates cover Text to Video, Image to Video, and Reference to Video. Opening one should prompt you to download anything missing. ComfyUI's native support means the basic path does not need a third-party custom-node bundle.
If you prefer manual downloads, place every file in the folder above and restart ComfyUI so the loaders rescan the model directories. A missing audio VAE is the usual reason a run produces picture without sound; a missing or misplaced text encoder prevents the prompt stage from loading.
Make the first run deliberately boring
Choose Text to Video, leave the template's default resolution and sampler settings alone, and render four or five seconds before trying a 15-second clip. Keep batch size at one. The first run is the slowest because ComfyUI has to load the 32-billion-parameter text encoder, encode the prompt, then swap to the diffusion model and VAEs.
Use a prompt that asks for one shot rather than a sequence of edits: subject, action, setting, camera movement, light, and the sound that should exist in the scene. H3 generates 32kHz stereo audio in the same job as the video, so dialogue, ambience, or effects should be part of the direction rather than an afterthought.
Once that works, move to Image to Video for first- and last-frame control, or Reference to Video when you need character, object, motion, style, video, or audio references. Change one variable at a time. When a 42.5GB stack fails, random node surgery is slower than confirming the four files, their folders, and a known-good template.
The hardware answer is: it depends on your patience
Full-precision H3 is a data-centre-sized model. ComfyUI's quantized package cuts the working set from roughly 123.6GB to about 42.5GB, then relies on dynamic offloading to move pieces between VRAM and system RAM. ComfyUI says that can run on a GPU such as the RTX 3060. It does not publish an official per-card timing or VRAM matrix.
That distinction matters. "Runs" on 12GB does not mean all weights fit in 12GB or that generation will be quick. The 21GB diffusion file alone is larger than the card. Community guidance treats 24GB as the practical tier for the pruned INT8 path, while 12–16GB cards trade money for waiting. Budget generous system RAM and free disk space for the model files, caches, and rendered video. Do not promise yourself a generation time until you test your own card.
The 2K claim needs an asterisk
The local base model produces clips with a 768px short edge, commonly around 1344×768 for 16:9, at 24 frames per second. H3-Regenerate-2K is a separate hosted stage, not part of the open local checkpoint. A local ComfyUI render can be upscaled afterwards with another model, but that is your added pipeline, not H3 running 2K offline.
The useful result is still substantial: up to 15 seconds of local video with synchronized stereo audio, generated in one pass. Just describe it accurately. Local H3 is a 768p video-and-audio model with an optional hosted route to 2K.
The setup worth keeping
Keep one untouched official template as your baseline. Duplicate it before adding acceleration nodes, experimental quantizations, ControlNet, or custom reference logic. If a later workflow breaks, the baseline tells you whether the problem is H3, the files, or your additions.
The recommended first-day setup is therefore plain: current ComfyUI, the official pruned INT8 four-file set, a built-in template, a five-second test, and no optimization until the baseline completes with audio. It is not the flashiest graph. It is the fastest way to find out whether H3 is actually usable on your machine.
Sources
- [1] ComfyUI documentation, "MiniMax H3 Video Generation Guide"Read source
- [2] ComfyUI blog, "MiniMax H3 Day-0 Support in ComfyUI"Read source
- [3] MiniMax H3 Community License Agreement (Aug 2, 2026)Read source
- [4] Local AI Master, "MiniMax H3 Local Setup: Run the Open Video Model in ComfyUI"Read source