← Back to all generators

inworld/realtime-tts-1.5-max

Highest-quality realtime text-to-speech with <200ms latency, emotion control, and 15-language support

Capabilities

No capability data available

Cost

Community model (estimated from hardware time)

Input Parameters

textrequiredstring

The text to convert to speech. Maximum 2,000 characters. Supports SSML break tags for pauses (e.g. `<break time="1s" />`), emotion markups (e.g. `[happy]`, `[sad]`), and non-verbal vocalizations (e.g. `[laugh]`, `[sigh]`).

audio_formatstring

Output audio format.

Default: "mp3"
mp3wavogg_opusflac
sample_rateinteger

Audio sample rate in Hz.

Default: 48000
8000160002205024000320004410048000
speaking_ratenumber

Speaking speed multiplier. Set to 0 for normal speed (1.0).

Default: 0min: 0, max: 1.5
temperaturenumber

Controls randomness when generating audio. Higher values produce more expressive results, lower values are more deterministic. Set to 0 to use the model default (1.1).

Default: 0min: 0, max: 2
text_normalizationstring

Controls whether numbers, dates, and abbreviations are expanded before synthesis. 'auto' lets the model decide, 'on' always normalizes, 'off' reads text as-is.

Default: "auto"
autoonoff
voice_idstring

The voice to use. Use a preset voice name (e.g. 'Ashley', 'Dennis', 'Alex') or a custom cloned voice ID.

Default: "Ashley"
Version: 4a2e51066a48Updated: 8/1/2026174.6K runs