Back to home

MiniMax H3 Max by fal: What It Is, Costs, and Limits

MiniMax H3 Max is fal Research’s post-trained version of the open-weight MiniMax H3 video model. fal serves it through text-to-video and image-to-video APIs at 480p or 768p, with native audio and 5–15-second outputs. It prioritizes faster hosted generation; standard H3 remains the route for 2K, video editing, or local ComfyUI work.

Written and fact-checked by: TF6coolAccess: hosted API and web toolOutput: 480p or 768p · 5–15 seconds · 24 fpsUpdated: September 4, 2026

Quick answer

Use H3 Max when you want a fast, hosted 480p or 768p text-to-video or image-to-video job with synchronized audio. Use standard MiniMax H3 when the job needs 2K output, video editing, downloadable weights, or a local ComfyUI setup. Reference-to-video reached the Max family as a preview on August 31, 2026, and a cheaper H3 Max Turbo preview followed on September 2. The name “Max” is a speed-and-adherence branch, not a superset of every H3 capability.

This guide is for developers and video creators deciding whether to call fal’s hosted H3 Max endpoints or invest in the broader standard H3 path. It covers product and API selection, not prompt craft or a hands-on quality test.

480p / 768p — The two output tiers listed by the H3 Max API.

5–15 sec — Select a duration from five through fifteen seconds.

Native audio — Video and synchronized audio are generated together.

Official fal H3 Max demo. Streamed from fal’s public media host; this site did not create or rehost the clip. Treat it as a selected product example, not a reproducible benchmark. If playback fails, view the official examples on fal.

What is a post-trained hosted video model?

A post-trained hosted video model starts with an existing base model, receives additional training or tuning for a narrower goal, and runs on the provider’s infrastructure instead of your machine. You access the result through a web tool or API; the provider controls the served weights, runtime, pricing, and operational limits.

This distinction matters because the base model’s license and the hosted product’s terms are separate questions. “Built from open weights” does not automatically mean the post-trained branch is downloadable, and a hosted speed claim does not predict performance on a local GPU.

What is MiniMax H3 Max by fal?

H3 Max is that pattern applied to MiniMax H3. MiniMax supplied the open-weight base model; fal says fal Research post-trained it for stronger instruction following and faster generation, then co-optimized it with fal’s inference stack. The resulting H3 Max product is accessed through fal’s hosted tool and APIs.

MiniMax supplies the base

MiniMax created and released MiniMax H3 as an open-weight video model. This site’s local workflow guides cover that standard family and its T2V, I2V, and R2V files.

fal supplies the Max branch and service

fal Research created the H3 Max branch and fal provides the checked access paths. Its launch report attributes the speed and prompt-adherence changes to post-training plus inference co-optimization.

Practical consequence: use fal’s current service terms and endpoint specification for H3 Max. Do not import assumptions from the standard H3 weight license merely because H3 is the base model. The local-weight license map is a separate path.

How the H3 Max stack works

The product can be understood as four layers. Only the last layer is the interface you call; the middle two explain why H3 Max should not be treated as either a MiniMax rename or a downloadable local checkpoint.

  1. MiniMax H3 — Open-weight base model
  2. fal Research — Post-training for adherence and speed
  3. fal inference — Co-optimized hosted runtime
  4. Tool or API — T2V, I2V, reference-to-video, Director, Turbo

Conceptual stack. Based on fal’s product and launch descriptions. This site did not observe the training pipeline or reproduce fal’s backend optimization.

H3 Max advantages

The value proposition is narrow: fast managed generation with a small API surface. These are reasons to test H3 Max, not guarantees that it will beat standard H3 for every prompt or production constraint.

Fast reported backend inference

fal reports under three seconds of inference for one 5-second 768p example. That is a useful latency target for a hosted test, but queueing, upload, prompt expansion, transfer, and download still affect wall time.

Turbo measurements: September 3–4, 2026

The timings below were collected on one fal account in one region, around 21:00 on September 3 and 01:00 on September 4, 2026 (US Pacific). Requests used queue.fal.run, a 1.5-second polling interval and concurrency 3. Prompt expansion was balanced except in the row marked disabled. Wall time runs from sending the submit request to receiving the result JSON; it excludes downloading the video. Inference is fal's timings.inference field, and RTF is wall time divided by clip duration.

Endpoint and outputWall p50 / p95Inference p50RTF p50 / p95Jobs
Turbo 480p · 10 s · balanced4.3 s / 12.9 s1.10 s0.43 / 1.2924
H3 Max 480p · 10 s · balanced6.0 s / 14.6 s1.75 s0.60 / 1.4629
Turbo 768p · 10 s · balanced8.9 s / 18.9 s4.66 s0.89 / 1.8930
H3 Max reference-to-video 480p · 5 s · balanced6.9 s / not measured2.11 s1.4 / not measured3
Turbo 480p · 10 s · disabled2.5 s / 4.2 s1.07 s0.25 / 0.4230

The polling interval adds up to 1.5 seconds of quantisation error to queue and overhead readings. Rejected 403 submissions were excluded. These are one account's observations, with no claim of statistical significance or a service-level guarantee.

The September 3 price-card readings recorded Turbo 480p at $0.00625 → $0.025 per output second: the discount scheduled to end September 7 followed by the listed standard rate. For a 10-second clip that is $0.0625 → $0.25. Turbo 768p was $0.01 → $0.04 per second, or $0.10 → $0.40 for 10 seconds. The disabled-expansion row used the same 480p rate. These are historical card readings and arithmetic, not a current quote.

The lower disabled-mode latency came with a visible change in these tests: the clips looked like CG toy renders in a continuous take, with no automatic soundtrack or shot planning from the prompt rewriter.

Native synchronized audio

The checked H3 Max endpoints generate video and audio together, avoiding a separate sound-generation step when a single short clip is the desired deliverable.

One small API surface

Text-to-video, image-to-video, and the H3 Max Turbo variants share the same duration, resolution, seed, and prompt-expansion controls, so one integration covers the family. That keeps the first integration smaller than a multi-workflow local graph.

Simple job-size math

Standard rates are stated per output second, so a 5-, 10-, or 15-second request can be estimated before submission. Retries and future price changes remain outside that basic calculation.

H3 Max limits and reasons to choose another path

H3 Max exchanges breadth and local control for a focused hosted path. Check these limits before treating the Max name as an all-purpose upgrade.

Resolution stops at 768p

The checked API offers 480p and 768p. A delivery requirement for 2K points to a standard H3 endpoint that explicitly lists that tier.

Editing and 2K stay on standard H3

By September 3, 2026 the H3 Max family on fal covered text-to-video, image-to-video, a reference-to-video preview, Director sessions, and the H3 Max Turbo preview. Video editing and the 2K tier still belong to the standard H3 family.

The checked route is hosted

fal’s materials did not link H3 Max weights. Local reproducibility, owned-hardware measurements, or an inspectable ComfyUI graph therefore require standard H3.

Terms and live facts can move

Pricing, free allowances, rankings, and service behavior can change. This guide dates those claims and separates fal’s reports from independent snapshots.

H3 Max vs standard MiniMax H3

The useful difference is not “new versus old.” It is a narrower hosted branch optimized for speed versus a broader model family with more tasks, higher output tiers, and a local-weight path.

Decision pointH3 Max on falStandard MiniMax H3
AccessHosted fal tool and APIHosted endpoints plus downloadable open weights for local workflows
Video tasksText-to-video, image-to-video, reference-to-video (preview since August 31, 2026), Director sessions, and H3 Max Turbo T2V/I2V (preview since September 2)T2V, I2V, reference-to-video, and video editing across the wider product family
Resolution480p or 768pfal lists options up to 2K, depending on the endpoint
Duration5–15 secondsVaries by standard H3 endpoint or local workflow
AudioNative synchronized audioNative audio support; exact controls depend on the chosen path
I2V framingStart image and optional end image; output follows the image aspect ratioBroader I2V and reference workflows, with settings determined by endpoint or graph
Local weightsNo public H3 Max weight link appeared in the fal materials checked on August 31Open-weight download and local ComfyUI route, subject to the H3 Community License
Best fitLow-latency hosted generation at 480p or 768p2K, editing, local control, or a reproducible owned-hardware workflow

Product capabilities were reconciled against the fal H3 Max overview and the checked endpoint pages. “No public H3 Max weight link” describes those checked materials; it is not a claim that a release can never happen.

MiniMax H3 Max pricing

fal’s published standard rate is $0.05/sec at 480p and $0.08/sec at 768p. The H3 Max Turbo preview endpoints list half of that ($0.025 and $0.04), and the reference-to-video preview lists $0.08/sec at either resolution. The totals below are direct multiplication from the card; the batch this site ran on September 3, 2026 was billed at exactly the card rate.

Clip length480p at $0.05/sec768p at $0.08/sec
5 seconds$0.25$0.40
10 seconds$0.50$0.80
15 seconds$0.75$1.20

Launch discount: 75% off until September 7, 2026

Read on September 3, 2026, the endpoint card showed $0.0125/sec at 480p and $0.02/sec at 768p, marked 75% off with the discount ending September 7, 2026; the Turbo card showed $0.00625 and $0.01 on the same terms. On August 31 the same card had shown $0.025 and $0.04 with a September 1 end date, so the window has already moved once. Standard rates apply after September 7. This guide keeps the standard rates as the evergreen basis; see the source reconciliation, then confirm the price on the live endpoint before a large batch.

Standard totals calculated from the fal H3 Max endpoint price card and checked August 31 and re-read September 3, 2026. They do not include assumptions about retries, storage, egress, taxes, or future account pricing.

Who should — and should not — use H3 Max?

Choose from the output you must deliver and the runtime you want to own. The model name is less important than resolution, task type, access method, and how much of the workflow you need to inspect.

Use H3 Max for hosted T2V or I2V

Your output target is 480p or 768p, native audio matters, and reducing setup or turnaround matters more than owning the runtime. Start with balanced prompt expansion before paying the latency cost of quality mode.

Use standard H3 for 2K

H3 Max stops at 768p in the checked API. If delivery resolution is the hard requirement, choose a standard H3 endpoint that explicitly lists 2K rather than assuming the Max suffix includes every higher tier.

Use standard H3 for editing; R2V has a Max preview

Video editing belongs to the wider standard H3 family. Reference-to-video exists on H3 Max as a preview endpoint since August 31, 2026, priced at $0.08/sec with no launch discount. For local reference workflows, start from the pinned workflow map.

Use local H3 for control and evidence

Downloadable weights and ComfyUI let you pin files, inspect every node, retain outputs, and measure your own hardware. The raw GPU ledger shows that cost without comparing consumer hardware to fal’s different backend workload.

How to get started with MiniMax H3 Max

Validate the fit in four steps before building a larger integration. The examples below use a 5-second 768p request and balanced prompt expansion so the cost and latency choices are visible rather than hidden in defaults.

1 · Test one clip

Use fal’s official tool for a short prompt. Check the live free-tier label and judge motion, audio, and prompt following against your own content.

2 · Choose the endpoint

Use text-to-video when the prompt is the only creative input. Use image-to-video when a start frame must anchor composition or identity.

3 · Set the job size

Choose 480p or 768p and a duration from 5 to 15 seconds. Multiply duration by the current per-second rate before batching requests.

4 · Call it server-side

Install fal’s client, keep FAL_KEY in a server environment variable, submit the task-specific endpoint, and retain the returned video URL.

JavaScript examples

Both examples use fal.subscribe for a minimal queue-aware call. Production code should also decide how to handle uploads, retries, timeouts, logs, and webhook verification.

Text to video

import { fal } from "@fal-ai/client";

fal.config({ credentials: process.env.FAL_KEY });

const result = await fal.subscribe("minimax/h3-max/text-to-video", {
  input: {
    prompt: "A paper kite rises above a windy coastal cliff",
    duration: 5,
    resolution: "768P",
    aspect_ratio: "16:9",
    prompt_expansion_mode: "balanced"
  },
  logs: true
});

console.log(result.data.video.url);

Image to video

const result = await fal.subscribe("minimax/h3-max/image-to-video", {
  input: {
    prompt: "The camera arcs left as the fabric moves in the wind",
    image_url: "https://your-cdn.example/start-frame.jpg",
    duration: 5,
    resolution: "768P",
    prompt_expansion_mode: "balanced"
  },
  logs: true
});

console.log(result.data.video.url);

Do not ship FAL_KEY to the browser

The direct credential belongs in a server-side environment variable or a server proxy that creates restricted requests. A key embedded in client JavaScript can be copied and used against your balance. fal’s client also supports queue status and webhooks for jobs that should not hold a request open.

Endpointsminimax/h3-max/text-to-video · minimax/h3-max/image-to-video
Resolution480P or 768P; 768p is the documented default
Duration5 through 15 seconds
T2V aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, or 9:16
I2V inputsRequired start image URL; optional end image URL; output aspect ratio follows the input image
Prompt expansionbalanced adds roughly one second by fal’s description; quality may add up to about 30 seconds

Check the live text-to-video API reference and image-to-video API reference before pinning an SDK schema. The examples above are intentionally minimal and do not handle uploads, retries, or webhook verification.

Where H3 Max ranked when checked

Two independent leaderboards placed H3 Max first in their image-to-video views on August 31, 2026. That supports the model’s I2V quality position at that moment. It does not establish a permanent rank, a universal “best video model” claim, or this site’s own benchmark.

Leaderboard snapshotH3 Max readbackScope and boundary
Design Arena#1 · Elo 1349 · 3,574 battles · 2,416 wins / 1,158 losses · 67.6% win rate · fal · 6.1sImage-to-video leaderboard snapshot. Pairwise votes and ranks can move as new battles arrive.
Artificial Analysis#1 · Elo 1203 · 95% interval −9 / +9 · 5,400 samplesImage-to-video with audio leaderboard. This is not the text-to-video table or an all-model composite.

Read the current Design Arena image-to-video leaderboard and Artificial Analysis image-to-video leaderboard with its With Audio tab selected. fal’s own landing page showed older snapshot values when checked; this guide uses the live independent tables above and dates the readback.

Four details fal’s pages do not state consistently

These differences were visible across fal’s product page, numbered walkthrough, FAQ, launch blog, and endpoint card on August 31, 2026, and re-checked on September 3. They are reported as conflicts rather than silently “fixed” into one unsupported answer.

QuestionWhat the checked sources saidHow this guide handles it
Unsigned free resolutionThe product headline and FAQ said five free 5-second 768p generations per day without sign-up. The numbered step on the same experience said five 5-second 480p clips. fal’s August 31 comparison article added a second allowance: five clips a day on the tool page plus five more in the signed-in sandbox.The count and duration agree; the resolution does not. Verify the live tool before treating either tier as an entitlement.
Launch discount windowThe launch blog said the first week. The landing FAQ said the first 14 days. The endpoint price card said September 1 on August 31, then 75% off until September 7 when re-read on September 3. Vercel’s AI Gateway lists its own 50% window to September 13, and the Director endpoint a separate discount to September 14.Use standard rates for evergreen costs. Date every promo readback, treat the endpoint card’s current deadline as the operative one, and advise a live price check.
“Under three seconds”fal’s launch material described a 5-second 768p generation in under three seconds. Its response example exposed about 2.5 seconds in timings.inference.Call this backend inference time. Do not promise equivalent click-to-download latency after queueing, upload, prompt expansion, transfer, and download.
Weights and commercial usefal marketed H3 Max as commercially usable through its service, but the checked H3 Max materials did not link a public weight download.Treat H3 Max as the checked hosted product and review current fal terms. Keep the standard H3 Community License and its territory rules separate.

Sources checked: the fal H3 Max product page, fal tool, launch blog, and the linked API endpoint cards. This site did not reproduce fal’s backend timing or test the free allowance. Independent guide; not affiliated with MiniMax or fal.

MiniMax H3 Max questions

What is MiniMax H3 Max?

MiniMax H3 Max is a fal Research post-training of the open-weight MiniMax H3 video model, served on fal as text-to-video and image-to-video endpoints. It targets faster hosted generation and prompt adherence at 480p or 768p, with native synchronized audio and selectable 5-to-15-second output lengths.

Is MiniMax H3 Max open source or open weight?

The underlying MiniMax H3 model is available as open weights, but the fal materials checked on August 31, 2026 did not link downloadable H3 Max weights. fal presents H3 Max as a hosted tool and API. That access model should not be confused with the separate community license for standard H3 weights.

How much does MiniMax H3 Max cost on fal?

fal lists standard rates of $0.05 per second for 480p and $0.08 per second for 768p. That makes a 5-second clip $0.25 or $0.40 before any future pricing change. A 75% launch discount ($0.0125 and $0.02 per second) was on the endpoint card until September 7, 2026, and the batch this site ran was billed at exactly the card rate.

How fast is MiniMax H3 Max?

fal reports under three seconds of inference for a 5-second 768p example, and its sample response showed about 2.5 seconds in the inference timing field. That is backend inference time, not guaranteed end-to-end latency; queueing, upload, prompt expansion, network transfer, and download time can all add delay.

What is the difference between H3 Max and standard MiniMax H3?

H3 Max is the faster hosted fal option for 480p or 768p text-to-video and image-to-video, with a reference-to-video preview since August 31, 2026 and a cheaper H3 Max Turbo preview since September 2. Standard MiniMax H3 is the broader family: fal lists up to 2K and video editing, while the open weights support local workflows. Choose from required capabilities, not the word Max.

Can I use MiniMax H3 Max for free or for commercial work?

fal advertises limited free generations and describes H3 Max output as available for commercial use, but its free-tier page conflicts on whether unsigned clips are 480p or 768p. Check the live tool and current fal terms before relying on either allowance. Standard H3 weights have a separate community license and territory rules.

Your MiniMax H3 Max decision

Use H3 Max when a managed 480p or 768p T2V/I2V API solves the real job. If you need 2K, video editing, downloadable weights, or a locally inspectable run, continue with standard MiniMax H3 instead.