
MiniMax H3 Max Spicy
MiniMax H3 Max Spicy is a flagship uncensored video generation model powered by the high-tier MiniMax H3 Max architecture. Operating as a Spicy uncensored model, it completely bypasses standard safety guardrails to unlock maximum motion amplitude, intense physical interactions, and complex cinematic action sequences without content restrictions. Building on the enhanced rendering capacity and higher detail resolution of the Max tier, MiniMax H3 Max Spicy maintains exceptional frame consistency, realistic lighting, and photorealistic textures throughout dramatic camera movements. It is the premier tool for unrestricted dynamic video production, high-impact VFX pre-visualization, commercial action direction, and boundary-pushing creative AI visual assets.
Read Me
MiniMax H3 Max Video Spicy API
MiniMax H3 Max is a video generation model post-trained by fal Research on MiniMax H3's open weights, specialized for inference speed and prompt adherence, and introduced by fal on August 27, 2026. It retains the core multimodal capabilities of H3 while dramatically compressing generation time — fal's own measurement completes a 5-second 768P clip in roughly 3 seconds — positioning it for rapid drafts and batch short-video production. Each run generates 4–15 seconds with native 768P output and synchronized stereo audio, supporting text-to-video, first-frame-driven, and first/last-frame-constrained generation modes.
The model is offered through the iCreat platform as an API using a two-step asynchronous task flow: submit a task to obtain a task_id, then poll the result endpoint until a terminal status. Resolution is fixed at 768P, billed by output video duration at $0.08 per second, with reference media input not billed; six aspect ratios from 21:9 to 9:16 are supported, and text-to-video requests must specify a concrete ratio.
Model Positioning
MiniMax H3 Max targets video production where speed comes first. Unlike the standard H3 — which covers 2K HD output and broader multimodal referencing — it trades a focused capability set (native 768P, no 2K path, first/last-frame control) for second-level response per generation through post-trained inference optimization. Within the MiniMax H3 family it complements the standard version: the standard H3 serves wider multimodal conditioning and HD delivery, while H3 Max serves fast-draft, batch, and high-retry short-content pipelines. Typical needs include rapid ad-creative validation, batch social media video production, and storyboard sample generation.
Core Capabilities
Ultra-Fast Generation
Inference is optimized through dedicated post-training; speed improves dramatically over the original H3 at comparable quality. A 5-second 768P clip completes in roughly 3 seconds, cutting per-iteration wait time for large-scale material production.
Strong Prompt Adherence
Post-training focuses on prompt adherence and visual aesthetics. Camera positions, character actions, and object constraints are followed more reliably, reducing broken frames and off-track motion, so complex shot instructions land more stably.
Native Audio-Video Integration
Picture and stereo sound are generated jointly inside the model. No separate TTS or dubbing tool is needed — videos ship with scene-matched background audio, ambience, and character voices.
Multiple Input Modes
Three generation modes are supported: text-to-video, first-frame-driven generation, and first/last-frame-constrained generation. First/last-frame control anchors both the opening and closing images for controlled shot transitions.
Pricing
| Resolution | Unit Price (USD/second) | 5-Second Cost |
|---|---|---|
| 768p | $0.08 | $0.40 |
Total Cost = Unit Price × Output Video Duration. Only the output video duration is billed; reference media input is not charged separately. The costUSD field in the response returns the actual cost once the task succeeds.
Application Scenarios
- Rapid ad-creative validation: produce and compare multiple creative directions in parallel
- Batch social media video production: teams generate preview material at scale, then refine the picks
- Fast AI short-drama drafting: quickly produce storyboard samples to validate shots, characters, and plot
- E-commerce material previews: fast generation of product motion clips and promo samples
- Creative prototyping: turn written ideas into video sketches quickly to test prompt feasibility
Model Comparison
Same-Series Comparison
| Model | Output Resolution | Duration | Audio | Positioning |
|---|---|---|---|---|
| MiniMax H3 Max (this model) | 768P | 4–15 seconds | Native stereo | Fast-inference edition for rapid drafting |
| MiniMax H3 (standard) | 768P / 2K | 4–15 seconds | Native stereo | Full multimodal system for HD delivery |
| MiniMax H3 Max Image-to-Video | 768P | 4–15 seconds | Native stereo | First-frame-driven Max variant |
Cross-Model Comparison
| Model | Developer | Output Resolution | Duration | Audio |
|---|---|---|---|---|
| MiniMax H3 Max | MiniMax / fal | 768P | 4–15 seconds | Native stereo |
| Wan 2.6 Reference-to-Video | Alibaba Cloud | 720P / 1080P | 4–15 seconds | Native joint audio-video generation |
| Seedance 2.0 | ByteDance | Up to 2K | 4–15 seconds | Audio generation |
| Seedance 2.0 Mini | ByteDance | 720P / 1080P | 4–15 seconds | Audio generation |
Same-series data comes from MiniMax / fal official releases and channel documentation; this model's specifications follow the iCreat channel documentation.
Why Choose MiniMax H3 Max?
- Ultra-fast iteration: a 5-second clip in roughly 3 seconds makes retries nearly free — built for creative workflows that need constant adjustment
- Strong prompt adherence: shot and action instructions land reliably, cutting re-roll count
- Audio-visual in one pass: sound is generated with the picture, eliminating dubbing and sound-effect post-production
- Six aspect ratios covering mainstream landscape and portrait delivery; text-to-video must specify a concrete ratio, keeping output framing predictable
- Two-step async interface with clear status semantics;
status,costUSD, anderror_code/error_messagemake task progress, cost, and failure reasons transparent
Specifications
| Item | Description |
|---|---|
| Base URL | https://api.icreat.ai |
| Submit endpoint | POST /v1/task/submit/minimax/h3-max-global |
| Query result endpoint | POST /v1/task/result |
| Model ID | minimax/h3-max-global |
model parameter |
minimax/h3-max-global |
| Authentication | Authorization: Bearer header |
| Call pattern | Two-step asynchronous task (submit → poll) |
content |
Required, multimodal array: exactly one non-empty text item, may include first/last-frame images |
content[].type |
text / image_url |
content[].text |
Required when type is text; prompt limit 7,000 characters |
content[].image_url.url |
Required for image items; image URL accessible by the server |
content[].role |
Required for image items: first_frame / last_frame |
resolution |
Required, 768P |
duration |
Required, integer, 4–15 |
ratio |
Required, adaptive / 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16; text-to-video cannot use adaptive |
aigc_watermark |
Optional, default false |
| Task status | SUBMITTED / IN_PROGRESS / SUCCEEDED / FAILED |
| Output resource | type is video, includes url |
| Cost field | costUSD, returned only on SUCCEEDED |
| Failure fields | error_code / error_message, returned only on FAILED |
Architecture
MiniMax H3 Max is a specialized post-training of MiniMax H3's open weights. The original H3 system consists of three parts: H3-Context-IR parses free-form multimodal instructions into the representation the model consumes; H3-Base (a ~33B-parameter single-stream H3-Omni-Transformer that jointly predicts video and audio latents) performs core 768P audio-video generation; and H3-Regenerate-2K handles high-resolution regeneration. Building on the open H3 weights, fal applied post-training with its in-house reinforcement learning framework against real generation workloads and human preference data, co-designing the serving engine so that prompt adherence and audio-visual quality are retained while inference throughput improves dramatically. H3 Max weights are not open; it is cloud-served only. The 2K HD path remains in the standard H3 service.
Notes
- Submit and query must be chained with the same
task_id - While
statusisSUBMITTEDorIN_PROGRESS, the task is still processing; keep polling, with a suggested interval of 2–5 seconds FAILEDis terminal; readerror_codeanderror_messageto diagnose, then resubmitcontentmust contain exactly one non-emptytextitem; image items must not include thetextfieldlast_framerequiresfirst_frameto be provided as well; at most one of each- The
ratioof a text-to-video request cannot beadaptive; a concrete ratio must be specified - Image URLs must be non-empty and accessible by the server
aigc_watermarkdefaults to off; the exact watermark appearance depends on the provider route usedcostUSDis returned only when the task succeeds; failed tasks do not produce a cost field
FAQ
How do I get started with the MiniMax H3 Max API?
Register on the iCreat platform and obtain an API Key from the console, send a generation request to the submit endpoint, then poll the result endpoint with the returned task_id. All requests are authenticated with the Authorization: Bearer header. Keep your API Key safe and never expose it in client code or public repositories.
How is the cost calculated?
Total Cost = Unit Price × Output Video Duration. 768P is $0.08 per second; only the output duration is billed and reference media input is not charged separately. For example, a 5-second video costs $0.40. The costUSD field in the response gives the actual cost once the task succeeds.
What generation modes are supported?
Three modes: text-to-video (prompt only; ratio must be a concrete value), first-frame-driven (a first_frame image anchors the opening frame), and first/last-frame-constrained (both first_frame and last_frame control the transition between opening and closing frames). Both text and first-frame modes let you describe camera movement and scene sound in the prompt.
How do I choose an aspect ratio?
ratio supports six concrete values — 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 — plus adaptive. Text-to-video requests cannot use adaptive and must specify a concrete ratio; with first/last-frame images you may use adaptive, and the model follows the input image's framing.
How do I query the task result?
Send a POST request to https://api.icreat.ai/v1/task/result with the task_id in the body. A status of SUBMITTED or IN_PROGRESS means processing, with result as [] — keep polling. When status is SUCCEEDED, result returns the video resource array; read url to retrieve the video.
What if the task fails?
FAILED is a terminal status, and the response carries error_code and error_message explaining the failure. Check in order: whether content contains exactly one non-empty text item, whether last_frame is paired with first_frame, whether a text-to-video request mistakenly used adaptive, and whether duration is within 4–15. Fix any issue and resubmit. The cost field is returned only when the task succeeds.



