← Back to all generators

bytedance/omni-human-1.5

A film-grade digital human model that generates realistic video from a single image, audio clip, and optional text prompt.

Capabilities

Reference ImagesSeed

Cost

Community model (estimated from hardware time)

Input Parameters

audiorequiredstring

Input audio file (MP3, WAV, etc.). Duration must be less than 35 seconds. If the audio exceeds 35 seconds, an error will be generated and the generation will fail.

imagerequiredstring

Input image containing a human subject, face or character.

fast_modeboolean

Enable fast mode to speed up generation by sacrificing some effects.

Default: false
promptstring

Optional prompt for precise control of the scene, movements, camera movements, etc. Supports Chinese, English, Japanese, Korean, Spanish, and Indonesian.

seedinteger

Random seed for reproducible generation.

Version: b0f93aebf8c3Updated: 8/1/202649.2K runs