MiniMax H3 MAX

minimax/h3-max
OfficialImage-to-VideoText-to-Video

MiniMax H3 MAX is an economical, fast video generation model from MiniMax (the team behind the Hailuo video engine). It supports text-to-video, first-frame-to-video, and first-and-last-frame-to-video generation with direct 768P output. Dual first/last frame control precisely locks the starting and ending visuals, with 4–15 second clips across aspect ratios from 21:9 to 9:16. Billing is fully transparent — from $0.05/second at 480P and $0.08/second at 768P, charged only for the output seconds actually generated, with input text and frame images completely free — making it a cost-effective choice for e-commerce showcases, social media videos, and bulk creative production.

Resolution Unit Price (USD/second) 5-Second Cost
480p $0.05 $0.25
768p $0.08 $0.40

Read Me

MiniMax H3 MAX API

MiniMax H3 MAX is an economical and fast video generation model developed by MiniMax (Shanghai, Hailuo video engine team) and launched on the iCreat platform in September 2026. It supports text-to-video, first-frame-to-video, and first-and-last-frame-to-video generation with direct 768P output.

On the iCreat platform, the model is accessed via an asynchronous two-step API workflow (submit task -> poll query status). It features a unified endpoint for all generation modes, supports 480P and 768P resolutions, and is billed per second of generated video at $0.05/second for 480P and $0.08/second for 768P.

Model Positioning

MiniMax H3 MAX is positioned as a low-cost, fast-generation video model within the MiniMax H3 series. Compared to the full-modal flagship MiniMax H3 standard edition, it simplifies multimodal references and audio generation to focus on core image-to-video and text-to-video capabilities, offering a more economical option for high-throughput video generation tasks.

Core Capabilities

Multi-Mode Video Generation

Supports text-to-video, first-frame-to-video, and first-and-last-frame-to-video generation. Users can control the video's starting and ending visual states by providing reference images, enabling precise control over narrative transitions and dynamic effects.

Direct High-Resolution Output

Capable of directly outputting 768P resolution video. It also supports a 480P option for scenarios requiring faster generation or lower costs, providing flexibility for different production needs.

Flexible Aspect Ratios

Supports multiple aspect ratios including adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Image-to-video tasks can use the adaptive ratio to automatically match the aspect ratio of the input first frame.

Economical Per-Second Billing

Billed strictly based on the actual generated video duration in seconds, with no minimum duration threshold. Input text and reference images are not charged, ensuring cost-effectiveness for short video generation.

Pricing

Resolution Unit Price (USD/Second) 5-Second Cost
480P $0.05 $0.25
768P $0.08 $0.40

Note: Total cost = Unit price x Output video duration. Billed based on actual generated duration with no minimum duration threshold. Input text and reference images are not charged.

Application Scenarios

  • Fast generation of short videos for social media platforms
  • Creating dynamic video content from static images with first-frame control
  • Generating transition effects and narrative shorts using first and last frame control
  • High-volume video generation tasks requiring cost optimization
  • Rapid prototyping of video concepts before using higher-resolution flagship models

Model Comparison

Comparison Table 1: MiniMax H3 Series

Feature MiniMax H3 MAX MiniMax H3 (Standard)
Positioning Low-cost fast version Full-modal flagship
Endpoint Unified endpoint minimax/h3-max Separate text-to-video / image-to-video endpoints
Resolutions 480P / 768P 768P / 2K
Duration 4-15 seconds 4-15 seconds
480P Price $0.05/s (5s $0.25) -
768P Price $0.08/s (5s $0.40) $0.10/s (5s $0.50)
2K Price - $0.14/s (5s $0.70)
Audio No native audio 24 FPS, 32 kHz native stereo
Multimodal References First/last frame images Up to 9 images + 3 videos + 3 audio
Open Weights Not open-sourced 33B weights open-sourced

Comparison Table 2: Cross-Model Comparison

Feature MiniMax H3 MAX Hailuo 2.3 Seedance 2.0 Economy
Developer MiniMax MiniMax ByteDance
Resolutions 480P / 768P 768P / 1080P 480p / 720p
Duration per Request 4-15 seconds 6-10 seconds 4-15 seconds or -1
Frame Control First + last frame First frame only Text/Image/Video/Audio inputs
Native Audio No native audio No native audio Native audio-visual sync
Billing Per second Per generation (from 2 RMB each, CN list price) Per second (output seconds only)
480P/720p Price $0.05/s (5s $0.25) - $0.077/s (5s $0.3850)
768P/720p Price $0.08/s (5s $0.40) - $0.164/s (5s $0.8200)

Why Choose MiniMax H3 MAX?

  • Cost-effective generation with per-second billing starting at $0.05/second for 480P
  • Unified endpoint simplifies integration for text-to-video and image-to-video workflows
  • Supports both first-frame and first-and-last-frame control for enhanced narrative direction
  • Direct 768P output meets standard high-definition requirements
  • Flexible duration options from 4 to 15 seconds accommodate various short video needs
  • Economical positioning ideal for high-volume generation tasks without paying for unused modalities

Specifications

Field Value
Model Name MiniMax H3 MAX
Developer MiniMax
Model ID minimax/h3-max
model Parameter MiniMax-H3-Max
Release Date September 2026
Model Type Video Generation Model
Endpoint https://api.icreat.ai/v1/task/submit/minimax/h3-max
Query Endpoint https://api.icreat.ai/v1/task/result
Authentication Authorization: Bearer ${ICREAT_API_KEY}
Input Modes Text, First Frame Image, Last Frame Image
Supported Resolutions 480P, 768P
Aspect Ratios adaptive, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Duration per Request 4-15 seconds
Prompt Limit 7000 characters
Frame Limits Max 1 first frame, max 1 last frame
Billing Unit Per second of generated video
API Mode Asynchronous two-step (Submit -> Poll Query Status)

Architecture

MiniMax H3 MAX operates on an asynchronous two-step API architecture. Users submit a video generation task to the submission endpoint, including the prompt, reference images, resolution, duration, and ratio. The API returns a task ID immediately. Users must then poll the query status endpoint at 2-5 second intervals using the task ID. The task status transitions through SUBMITTED, IN_PROGRESS, SUCCEEDED, or FAILED. Upon success, the response includes the status, costUSD (as a string), and the result array containing the video URL.

Notes

  • The duration parameter must be an integer between 4 and 15; values outside this range will return an error.
  • The content array must contain exactly one non-empty text item; the prompt can be up to 7000 characters.
  • A maximum of one first frame and one last frame image is allowed. If last_frame is used, first_frame must also be provided.
  • For pure text-to-video, the ratio must be a specific aspect ratio (adaptive is not allowed). For image-to-video, adaptive can be used to match the input image.
  • The aigc_watermark parameter defaults to false; the actual watermark effect depends on the specific route.
  • The callback_url is a protocol-reserved field; currently, active polling is required to determine task status.
  • The query endpoint is https://api.icreat.ai/v1/task/result. FAILED is a terminal state; read error_code and error_message to troubleshoot before resubmitting.
  • The costUSD field is a string type and is only returned when the task succeeds.

Frequently Asked Questions

How do I call the MiniMax H3 MAX API?

To call the API, send a POST request to the submission endpoint https://api.icreat.ai/v1/task/submit/minimax/h3-max with your API key in the Authorization header. The request body must include the model name, content array, resolution, duration, and ratio. After receiving the task ID, poll the query endpoint https://api.icreat.ai/v1/task/result every 2-5 seconds until the status is SUCCEEDED or FAILED.

How is the model billed?

The model is billed per second based on the actual generated video duration. The rate is $0.05 per second for 480P resolution and $0.08 per second for 768P resolution. There is no minimum duration threshold, and input text or reference images are not charged.

How do I use the first and last frame features?

To use frame control, add image objects to the content array. Each image object must have type set to image_url, an image_url object containing the URL, and a role field set to either first_frame or last_frame. You can provide at most one of each. If you provide a last_frame, you must also provide a first_frame.

What are the ratio restrictions for pure text-to-video?

When generating video using only text (without any reference images), you must specify a concrete aspect ratio such as 16:9 or 9:16. The adaptive ratio option is only available for image-to-video tasks, where it automatically matches the aspect ratio of the provided first frame image.

What is the difference between MiniMax H3 MAX and the H3 standard version? MiniMax H3 MAX is an economical, fast version focused on core text and image-to-video generation up to 768P. The H3 standard version is a full-modal flagship supporting 2K resolution, native stereo audio, and extensive multimodal references (images, videos, and audio), but it comes at a higher price per second.

How do I troubleshoot a FAILED task?

If a task returns a FAILED status, it is a terminal state. You should check the error_code and error_message fields in the response to identify the issue. Common causes include invalid parameters, durations outside the 4-15 second range, or inaccessible image URLs. After correcting the issue, submit a new task.