fal Launches MiniMax H3 as an Official API Partner, Bringing the Open-Weight Multimodal Video Model to Developers

Developers can now run MiniMax’s next-generation model on fal through three dedicated endpoints—text-to-video, image-to-video, and reference-to-video—with native stereo audio and output up to 1440p (2K).

fal, the generative media platform for developers, today announced that it is an official API launch partner for MiniMax H3, MiniMax’s next-generation open-weight multimodal video model. H3 is available on fal now through three dedicated endpoints—text-to-video, image-to-video, and reference-to-video—so teams can integrate it into production applications with a single API call.

H3 moves beyond single-purpose tools toward general multimodal intelligence: it interprets text, images, video, and audio within one unified context, reading creative intent across every input and producing coherent, complete audio-visual content in a single end-to-end pass. On fal’s high-performance inference platform, developers get that capability with the speed, scalability, and reliability needed to ship it to users.

“Open source means control, customizability, and freedom. That matters to a lot of our customers—being able to train on their own images, video, and audio together unlocks real personalization. It’s a big deal for brands, film, and plenty of other industries.” — Gökay Aydoğan, Creative Engineer, fal

What developers get with MiniMax H3 on fal

Native multimodal understanding and generation. MiniMax H3 accepts combined input across text, images, audio, and video, interprets the creative intent behind them, and generates coherent audio-visual content in one integrated pass—rather than stitching together the output of separate systems.

Precise multimodal editing and control. Replace, remove, or add characters and objects; swap backgrounds, relight scenes, and adjust visual effects; and edit dialogue, transfer vocal identity, or clone a voice. Localized edits apply where intended while untouched areas stay stable, with strong instruction-following throughout.

Production-ready across use cases. Built for real commercial work in film and entertainment, advertising and branding, e-commerce, and gaming—handling dynamic typography, VFX, product showcases, UI/UX motion design, game visuals, and stylized content.

Fal is consistently mentioned as one of the best places to run your AI infrastructure and APIs. The fact that they’re able to leverage this new state-of-the-art model on their platform is a game changer. They run at some of the best inference speed in the industry.

Key specifications

Output duration

5–15 seconds

Resolution

1440p (2K) mode; 768p mode coming soon, upscalable to 1440p

Frame rate

24 FPS

Audio

Native stereo audio included in every generation

Aspect ratios

21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus an Auto mode in reference-to-video

Input modes

Text-to-video, image-to-video (first / last frame), and reference-to-video (up to 9 images, 3 video clips, and paired audio; up to 12 files total)

Prompt length

Up to 7,000 characters

Pricing and the full set of API parameters for each endpoint are listed on the model pages on fal. For API workflows, URL-based media input is recommended.

From brand films to game concepts

Teams building on fal can use H3 across cinematic brand films and trailers; creative shorts, VFX packages, and motion-design social assets; AI-native storytelling such as vertical dramas, motion comics, and character performances with AI dubbing; product and e-commerce marketing; digital experiences and game concepts such as UI walkthroughs and interaction demos; and stylized animation spanning game cinematics, anime promos, and IP content—rendering text, subtitles, brand assets, UI, and game visuals while preserving character identity, camera language, and edit rhythm.

Availability on fal

MiniMax H3 is available on fal today through three dedicated endpoints:

Image → Video fal.ai/models/minimax/hailuo-03/image-to-video

Text → Video fal.ai/models/minimax/hailuo-03/text-to-video

Reference → Video fal.ai/models/minimax/hailuo-03/reference-to-video

About fal

fal is the generative media platform for developers, providing fast API access to the world’s best AI image, video, audio, and 3D models. Founded in 2021 and headquartered in San Francisco, fal runs a high-performance inference platform on serverless GPUs, hosting more than 1,000 production-ready models and serving over 2.5 million developers and companies including Amazon MGM Studios, Canva, and Adobe. Learn more at fal.ai.

About MiniMax

MiniMax is an artificial-intelligence company founded in 2021 and headquartered in Shanghai, China, developing foundation models across text, audio, image, and video. Its Hailuo video models and the MiniMax Open Platform make its multimodal technology available to creators and developers worldwide.

Media Contact
Company Name: Features & Labels, Inc. (fal.ai)
Contact Person: Media Relations
Email: Send Email
Country: United States
Website: https://fal.ai/minimax-h3