Output duration, frame rate, and sound
H3 outputs 4 to 15 seconds at 24 FPS. Every result includes native stereo sound. The official manual lists a maximum prompt length of 7,000 characters.
Aspect ratios and resolution
Text-to-Video and Reference Generation accept 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Reference Generation can also choose an adaptive ratio. First/Last-Frame generation follows the supplied image ratio. The manual documents 768p output and a higher-resolution 1440p workflow; check the current API page before budgeting because product labels may differ by entry point.
First and last frame inputs
Supply zero, one, or two images. Each image dimension must be between 256 and 5,760 pixels and the ratio must remain between 5:2 and 2:5. Zero images becomes Text-to-Video.
Reference Generation inputs
Use up to 9 images, 3 videos, and 3 audio clips, with no more than 12 files total. Each video or audio clip must be 2 to 15 seconds, and total reference-video duration and total reference-audio duration are each capped at 15 seconds. Audio cannot be the only input.
Formats and request size
Video accepts H.264/AVC or H.265/HEVC with AAC or MP3 audio. Images accept JPG, JPEG, PNG, WebP, HEIC, and HEIF. Audio accepts WAV and MP3. Per-file limits are 50 MB for video, 30 MB for images, and 15 MB for audio. API request bodies are capped at 64 MB, so URL inputs are recommended.
Languages
The manual says H3 accepts multilingual prompts and output. Text-to-speech covers 11 languages precisely, including Chinese, English, Japanese, Korean, French, German, and Spanish, with broader exploratory support across more languages.