MiniMax H3 launched on July 31, 2026 as a multimodal video model for generation and editing. It accepts text, images, video, and audio in one context, produces up to 15 seconds of 2K video with native stereo sound, and is already available through fal's hosted API.
MiniMax also said it plans to publish the model weights “in the coming days.” That is an open-weight announcement, not proof that a downloadable checkpoint is available today. As of August 1, the official MiniMax Hugging Face organization does not list an H3 repository.
MiniMax H3 Quick Facts
| Item | Verified status on August 1, 2026 |
|---|---|
| Release announcement | July 31, 2026 |
| Other names | MiniMax H3; some provider material also uses Hailuo 3.0 |
| Inputs | Text, images, video, and audio in one reference context |
| Output | 2K video at 24 FPS with native stereo audio |
| Duration | 5–15 seconds on fal's current endpoints |
| Reference limits | Up to 9 images, 3 videos, and 3 audio clips; 12 files total |
| Generation modes | Text-to-video, first/last-frame, multimodal reference-to-video, and editing |
| Hosted API | Available through fal at launch |
| Hosted price | $0.26 per generated second at 2K on fal |
| Open weights | Announced, but an official public checkpoint was not verified for this update |
What Is MiniMax H3?
MiniMax H3 is a general-purpose audiovisual generation model. Instead of treating text-to-video, motion transfer, voice reference, and local editing as unrelated tools, it reads multiple media types as a shared context and returns a synchronized video and audio result.
That design matters when a shot needs more than a prompt. A creator can use one image for character identity, another for art direction, a video for movement, an audio clip for voice or rhythm, and text for camera and editing instructions. H3 is designed to resolve those references together.
MiniMax positions the model for advertising, branding, ecommerce, product design, UI motion, games, and other commercial creative work. Those use cases match its emphasis on readable text, product preservation, motion transfer, and targeted editing. Narrative production is still possible, but teams should test character continuity and causal action across several shots rather than infer episode-level reliability from launch demos.
MiniMax H3 Specs and Reference Limits
fal's launch-partner documentation exposes three H3 endpoint families:
- Text to video: a prompt generates a 5-to-15-second 2K clip.
- Image to video: one image becomes the opening frame, with an optional final frame for transition control.
- Reference to video: up to 9 images, 3 video clips, and 3 audio tracks can guide identity, style, motion, camera, or sound.
The current reference endpoint caps the complete upload set at 12 files. Reference video and audio are each limited to 15 seconds in total, and prompts can contain up to 7,000 characters. Text and reference generation support common landscape, square, portrait, and ultrawide aspect ratios; first/last-frame generation follows the input image.
These are hosted-provider limits. A future official MiniMax API or local release can expose a different schema, so record the provider and exact endpoint with every output.
Native Audio and Editing
Every documented H3 generation includes stereo audio. The launch material describes dialogue, music, ambience, and sound effects timed to the picture, plus voice transfer from an audio reference.
H3 also promotes instruction-based editing: replacing a product, rewriting signage, changing dialogue, relighting a scene, or adding and removing objects while preserving untargeted content. This is more production-relevant than regenerating an otherwise approved clip to fix one detail.
Treat both capabilities as testable claims. A useful acceptance check measures lip synchronization, voice identity, unwanted music, text stability, whether the edit leaks outside the selected subject, and whether a second edit preserves the first.
MiniMax H3 Demo and Hands-On Video
MiniMax and fal also publish selected examples on the MiniMax H3 launch page. Those examples show typography, interface motion, character work, motion reference, compositing, and video editing. Because they are launch selections, use them to understand supported workflows—not to estimate first-try success rate.
MiniMax H3 API Access
fal exposes H3 through these current model paths:
minimax/h3/text-to-videominimax/h3/image-to-videominimax/h3/reference-to-video
The endpoints are asynchronous hosted inference routes available through JavaScript, Python, or REST. They return a generated video URL after the queued request completes. Store the request ID, full input manifest, endpoint path, price at submission time, and output URL in the project record.
Do not assume that Hailuo-3.0, minimax/h3, and a future downloadable checkpoint are interchangeable. Provider aliases can wrap different defaults, safety layers, codecs, or optimizations even when they refer to the same model family.
MiniMax H3 Pricing
fal currently charges $0.26 per generated second at 2K for text-to-video and image-to-video. At that rate:
| Duration | Base generation price |
|---|---|
| 5 seconds | $1.30 |
| 10 seconds | $2.60 |
| 15 seconds | $3.90 |
For reference-to-video, the generated output is also $0.26 per second. Audio references and the first five images are free on the current fal price sheet; each additional image costs $0.08, and reference video costs $0.26 per input second at 2K.
These figures are a dated provider snapshot, not a permanent MiniMax list price. Production budgets should include failed jobs, retries, reference-video charges, storage, review time, and editing. Compare total cost per approved second rather than the base price of one request.
Is MiniMax H3 Open Source?
“Open source” is too broad for the current status. MiniMax announced that it intends to release H3's model weights in the days after launch, subject to laws and regulations. The announcement emphasizes hardware compatibility and community integration.
Before calling H3 locally available, verify four separate artifacts:
- an official MiniMax repository;
- downloadable model weights with hashes;
- inference code and a supported runtime;
- a license that permits the intended research or commercial use.
None of those should be inferred from a hosted API. This article will be updated when the official checkpoint and license can be reviewed directly.
Who Should Test MiniMax H3?
H3 belongs on the shortlist when a workflow needs 2K output, native audio, first/last-frame control, motion reference, several reference modalities, targeted editing, or accurate on-screen graphics.
For motion comics and short drama, run a sequence benchmark with at least:
- two dialogue close-ups using the same character and voice;
- a full-body action shot with hand-to-object contact;
- two-character staging with a camera move;
- a vertical shot that preserves costume and background;
- one local edit to a prop or sign;
- one shot that must match an earlier approved frame.
AniKuku keeps scripts, shots, character assets, scene references, prompts, and output versions outside the provider. That makes it possible to test H3 on selected shots without rebuilding the episode around one API.
FAQ
Is MiniMax H3 available now?
Yes through fal's hosted API as of August 1, 2026. Availability through MiniMax's own global platform or other partners may differ.
Is MiniMax H3 the same as Hailuo 3.0?
MiniMax uses H3 in its announcement, while some provider and API material uses Hailuo 3.0 for the same new model family. Preserve the exact provider model ID because names alone do not guarantee identical settings.
Does MiniMax H3 generate audio?
Yes. The current launch and API materials describe native stereo audio for every generation, including dialogue, music, effects, and ambience.
Can MiniMax H3 run locally?
MiniMax announced an open-weight release, but an official public checkpoint was not verified on August 1. Wait for the repository, weights, runtime instructions, hardware requirements, and license before planning local deployment.
Is MiniMax H3 better than Seedance?
H3 is easier to test today because a documented hosted endpoint is live. Seedance 2.0 remains highly competitive, while Seedance 2.5 is official but still marked coming soon. The best choice depends on reference control, continuity, audio, retry rate, latency, and cost. See the Seedance vs MiniMax H3 comparison.