Output duration, frame rate, and sound

H3 outputs 4 to 15 seconds at 24 FPS. Every result includes native stereo sound. The official manual lists a maximum prompt length of 7,000 characters.

Aspect ratios and resolution

Text-to-Video and Reference Generation accept 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Reference Generation can also choose an adaptive ratio. First/Last-Frame generation follows the supplied image ratio. The manual documents 768p output and a higher-resolution 1440p workflow; check the current API page before budgeting because product labels may differ by entry point.

First and last frame inputs

Supply zero, one, or two images. Each image dimension must be between 256 and 5,760 pixels and the ratio must remain between 5:2 and 2:5. Zero images becomes Text-to-Video.

Reference Generation inputs

Use up to 9 images, 3 videos, and 3 audio clips, with no more than 12 files total. Each video or audio clip must be 2 to 15 seconds, and total reference-video duration and total reference-audio duration are each capped at 15 seconds. Audio cannot be the only input.

Formats and request size

Video accepts H.264/AVC or H.265/HEVC with AAC or MP3 audio. Images accept JPG, JPEG, PNG, WebP, HEIC, and HEIF. Audio accepts WAV and MP3. Per-file limits are 50 MB for video, 30 MB for images, and 15 MB for audio. API request bodies are capped at 64 MB, so URL inputs are recommended.

Languages

The manual says H3 accepts multilingual prompts and output. Text-to-speech covers 11 languages precisely, including Chinese, English, Japanese, Korean, French, German, and Spanish, with broader exploratory support across more languages.