MiniMax H3 — AI Video Generation with Native Audio

Turn text, images, video and audio references into up to 15 seconds of 2K footage with synchronized stereo sound. MiniMax H3 follows complex instructions and keeps your product, logo and typography consistent — built for ads, e-commerce and brand work.

MiniMaxH3 example 1
MiniMaxH3 example 2
MiniMaxH3 example 3

What is MiniMax H3?

MiniMax H3 is the omni-modal video generation model released by MiniMax on July 31, 2026, with open weights following on August 3. It is the third generation in the Hailuo line after Hailuo 01 and Hailuo 02 — the community also calls it Hailuo 3.0. Instead of treating text, images, video and audio as separate inputs, H3 reads them in one multimodal context and generates up to 15 seconds of video at up to 2K, with native stereo sound produced in the same pass.

H3 does not try to out-render cinematic models frame by frame. Its bet is different: understand complex instructions, keep every element of your source material under control, and deliver picture and sound together. That makes it a natural fit for commercial packaging work — ads, e-commerce videos, brand assets, UI and game promos. Creators consistently highlight its text and logo stability: turning a poster into a motion ad or animating typography at a level competing models struggle to match.

On the Artificial Analysis leaderboard, MiniMax H3 ranks #1 in video editing, #2 in text-to-video and #3 in image-to-video — the first openly downloadable model to top a category there. The 768p H3-Base checkpoint is open source with Day-0 support across Hugging Face, ComfyUI, vLLM and major hardware vendors, while 2K output runs through the hosted H3-Regenerate-2K pipeline.

How to Create Videos with MiniMax H3

From a blank prompt or a folder of brand assets to a finished clip with sound — in four steps.

1

Add your references (or start from text)

Upload up to 9 images, 3 video clips and 3 audio tracks — 12 files max. Product shots, brand color cards, reference footage and background music can all steer the result. No references? Plain text-to-video works too.

2

Write the prompt

Describe the scene, camera moves and dialogue. MiniMax H3 reliably handles dialogue in 11 languages, including English, Chinese, Japanese, Korean, French, German, Italian, Spanish, Portuguese, Russian and Arabic.

3

Pick duration and frame

Choose 4 to 15 seconds at 24 FPS, in aspect ratios from 21:9 cinematic to 9:16 vertical — 16:9, 4:3, 1:1 and 3:4 included.

4

Generate, then upscale to 2K

Preview at 768p, then let H3 regenerate the result at 2K using the full original context — real regeneration that recovers small text and fine detail, not a blur-and-sharpen upscale.

Why Creators Choose MiniMax H3

Four capabilities that set MiniMax H3 apart from other AI video generators.

🎨

Omni-modal references

Mix images, video clips and audio in one request. Real-world tests show multi-reference runs — product photo plus color card plus reference footage plus soundtrack — hold consistency noticeably better than single-reference generations.

Native stereo audio

Sound is not an afterthought: 32kHz stereo audio is generated in the same pass as the video, synchronized from the start. No separate dubbing or alignment step.

Text & logo stability

H3's standout strength. Poster-to-motion ads, animated logos and on-screen typography stay legible and on-brand — the kind of visual packaging work reviewers say competing models can't yet match.

📱

True 2K regeneration

Instead of a super-resolution filter, H3 feeds the 768p result back through the model with the original context and regenerates at 2K — recovering details that upscalers can only guess at.

Subscription Plans

Save more with our subscription plans. Cancel or change anytime.

Free

$0USD /mo
  • 3.5 credits
  • Up to 1 image at 1K resolution
  • 1 parallel task
  • All-in-one multi-model support
  • Text & Image to Image
Flash Sale 50% Off

Pro

$14.50$29USD /mo

Billed annually

  • 800 credits/month
  • Up to 266 images/month at 1K
  • 3 parallel tasks
  • All-in-one multi-model support
  • Text & Image to Image
  • Priority generation queue
  • No watermark
  • Commercial license
  • Early access to new models

Lite

$10$15USD /mo

Billed annually

  • 300 credits/month
  • Up to 100 images/month at 1K
  • 2 parallel tasks
  • All-in-one multi-model support
  • Text & Image to Image
  • Priority generation queue
  • No watermark
  • Commercial license
  • Early access to new models

One-time Credit Packs

Top up when you need extra generations, without starting a subscription.

Small Pack

$25USD one-time
  • 300 credits included
  • Valid for 365 days
  • Includes all Lite subscription features except credits

Medium Pack

$55USD one-time
  • 800 credits included
  • Valid for 365 days
  • Includes all Lite subscription features except credits

Large Pack

$99USD one-time
  • 1800 credits included
  • Valid for 365 days
  • Includes all Lite subscription features except credits

Ready to create with MiniMax H3?

Generate ad-ready video with synchronized sound, right in your browser. Free starter quota, no setup.

Start creating now

MiniMax H3 — Frequently Asked Questions

Specs, licensing, pricing and honest limitations — everything you need to know before you build with MiniMax H3.

What is MiniMax H3?

MiniMax H3 is an omni-modal video generation model from MiniMax, released on July 31, 2026, with weights opened on August 3. It is the third generation of the Hailuo family (often called Hailuo 3.0) and generates video with native stereo audio from text, images, video and audio references.

How is MiniMax H3 different from Hailuo 02?

H3 abandons the Hailuo-02 architecture in favor of a 33B dense single-stream omni transformer that treats task generalization as the core design goal. Attention and FFN layers contain no modality-specific structures, and a three-dimensional MM-RoPE encodes time and space positions.

What are the output specs?

4–15 seconds at 24 FPS. Default output is 768p (short side); 2K is produced by the H3-Regenerate-2K pass. Aspect ratios include 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, and audio is 32kHz stereo generated together with the picture.

Which dialogue languages are supported?

Eleven languages are stably supported: Chinese, English, Japanese, Korean, French, German, Italian, Spanish, Portuguese, Russian and Arabic.

Is MiniMax H3 open source?

Partially. The 768p H3-Base model is open source, and community builds run it on cards as small as an RTX 3060 (ComfyUI's optimized variant needs about 42.5GB total memory, down 66% from full precision). The H3-Context-IR preprocessing system and the H3-Regenerate-2K pipeline stay hosted — so local deployments top out at 768p, and 2K requires the official API.

What are the license restrictions?

The community license excludes the US, EU, UK and South Korea from running the weights locally or using their outputs (the hosted API remains available, and exemptions can be requested). Companies above $20M annual revenue need a separate commercial license, commercial use requires displaying 'MiniMax H3' in your UI, and using H3 or its outputs to improve other AI models is prohibited everywhere.

How much does MiniMax H3 cost?

Official API pricing is ¥0.8 per second for 2K and ¥0.5 per second for 768p — which MiniMax says is under a third of mainstream 2K pricing and about half of typical 720p pricing.

What are its limitations?

Generation is slow: a 5-second 480p clip takes roughly 5–10 minutes on consumer hardware, multi-reference runs cost double, and the INT8 quantized build loses noticeable detail. Teams chasing cinematic frame quality should look elsewhere — MiniMax itself notes that multimodal context understanding and fine detail still have room to improve.