What Is MiniMax H3?
MiniMax H3 is MiniMax's open-weight, omni-modal video model — one model that takes text, images, video, and audio together and returns video with synchronized native audio in a single pass.
- Native Audio + VideoScore, dialogue, foley, and room tone generated together with the picture and timed to the cut — no separate audio pass, no manual sync.
- Multimodal ReferencesUp to 9 images, 3 video clips, and 3 audio tracks in one generation. Carry a subject, a camera move, and a voice forward together.
- Open-Weight Model, HostedMiniMax published the weights alongside the hosted API. VidCella runs it pay-per-generation — no local GPU, no setup, no download.
- 480p or 768p, 5-15 SecondsTwo output resolutions and seven aspect ratios, from 21:9 ultrawide through 9:21 vertical.
How to Use MiniMax H3 on VidCella
Four steps from prompt to finished clip:
MiniMax H3 Features on VidCella
Full MiniMax H3 capability surface, pay-per-generation:
Native Stereo Audio
Music, dialogue, foley, and ambience generated with the video rather than added after it.
Multimodal References
9 images + 3 videos + 3 audio tracks per generation, all in one call.
Text, Image & Reference to Video
Three modes on one model — prompt only, first/last frame control, or full multimodal reference.
480p & 768p · 5-15 Seconds
Seven aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and 9:21.
Open Weights, No GPU Needed
The weights are public, but you don't need them — generate from any browser, pay per generation.
Room to Create
VidCella runs H3 with lighter filtering than most hosted platforms. Handle real-person likeness, trademarks, and applicable law responsibly.
FAQs about MiniMax H3
Common questions about MiniMax H3 on VidCella
What is MiniMax H3?
MiniMax's open-weight, omni-modal video model, released August 2026. It takes text, image, video, and audio references and returns video with native synchronized stereo audio in a single pass.
What makes MiniMax H3 different from other video models?
It generates audio and video together rather than adding sound afterward. It also accepts up to 9 images, 3 videos, and 3 audio tracks as references in one call, carrying subject, motion, and voice forward together.
Do I need a powerful GPU to use MiniMax H3?
No — running the open weights locally takes a 24GB+ VRAM card to be comfortable. VidCella runs H3 hosted, so you generate from any browser with no GPU, download, or setup.
Can I use MiniMax H3 if I'm in the US, EU, UK, or South Korea?
Yes — VidCella runs H3 as a hosted service, and the regional carve-outs in MiniMax's open-weight community licence govern self-hosting the downloaded weights rather than hosted use. If you do plan to deploy the weights yourself, review the licence terms for your region first.
What resolutions, durations, and aspect ratios are supported?
480p or 768p output, 5 to 15 seconds, and seven aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and 9:21. Image-to-video follows your first frame's ratio, and reference videos are supported at 480p output only.
How much does MiniMax H3 cost on VidCella?
Text-to-video and image-to-video run 12 credits/s at 480p and 24 credits/s at 768p; reference-to-video is 15 and 30 credits/s respectively. Reference inputs add 6 credits per image, 6 per audio track, and 15 credits per second of reference video — pay-per-generation, no subscription.
