MiniMax H3: Open-Weight AI Video with Native Audio

MiniMax's open-weights omni-modal model — text, images, video, and audio in, video with native stereo audio out. 480p or 768p, 5-15 seconds, from 12 credits/s.

Loading...

What Is MiniMax H3?

MiniMax H3 is MiniMax's open-weight, omni-modal video model — one model that takes text, images, video, and audio together and returns video with synchronized native audio in a single pass.

  • Native Audio + Video
    Score, dialogue, foley, and room tone generated together with the picture and timed to the cut — no separate audio pass, no manual sync.
  • Multimodal References
    Up to 9 images, 3 video clips, and 3 audio tracks in one generation. Carry a subject, a camera move, and a voice forward together.
  • Open-Weight Model, Hosted
    MiniMax published the weights alongside the hosted API. VidCella runs it pay-per-generation — no local GPU, no setup, no download.
  • 480p or 768p, 5-15 Seconds
    Two output resolutions and seven aspect ratios, from 21:9 ultrawide through 9:21 vertical.

How to Use MiniMax H3 on VidCella

Four steps from prompt to finished clip:

MiniMax H3 Features on VidCella

Full MiniMax H3 capability surface, pay-per-generation:

Native Stereo Audio

Music, dialogue, foley, and ambience generated with the video rather than added after it.

Multimodal References

9 images + 3 videos + 3 audio tracks per generation, all in one call.

Text, Image & Reference to Video

Three modes on one model — prompt only, first/last frame control, or full multimodal reference.

480p & 768p · 5-15 Seconds

Seven aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and 9:21.

Open Weights, No GPU Needed

The weights are public, but you don't need them — generate from any browser, pay per generation.

Room to Create

VidCella runs H3 with lighter filtering than most hosted platforms. Handle real-person likeness, trademarks, and applicable law responsibly.

FAQ

FAQs about MiniMax H3

Common questions about MiniMax H3 on VidCella

1

What is MiniMax H3?

MiniMax's open-weight, omni-modal video model, released August 2026. It takes text, image, video, and audio references and returns video with native synchronized stereo audio in a single pass.

2

What makes MiniMax H3 different from other video models?

It generates audio and video together rather than adding sound afterward. It also accepts up to 9 images, 3 videos, and 3 audio tracks as references in one call, carrying subject, motion, and voice forward together.

3

Do I need a powerful GPU to use MiniMax H3?

No — running the open weights locally takes a 24GB+ VRAM card to be comfortable. VidCella runs H3 hosted, so you generate from any browser with no GPU, download, or setup.

4

Can I use MiniMax H3 if I'm in the US, EU, UK, or South Korea?

Yes — VidCella runs H3 as a hosted service, and the regional carve-outs in MiniMax's open-weight community licence govern self-hosting the downloaded weights rather than hosted use. If you do plan to deploy the weights yourself, review the licence terms for your region first.

5

What resolutions, durations, and aspect ratios are supported?

480p or 768p output, 5 to 15 seconds, and seven aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4, 21:9, and 9:21. Image-to-video follows your first frame's ratio, and reference videos are supported at 480p output only.

6

How much does MiniMax H3 cost on VidCella?

Text-to-video and image-to-video run 12 credits/s at 480p and 24 credits/s at 768p; reference-to-video is 15 and 30 credits/s respectively. Reference inputs add 6 credits per image, 6 per audio track, and 15 credits per second of reference video — pay-per-generation, no subscription.