inworld/realtime-tts-1.5-mini
Ultra-fast, cost-efficient realtime text-to-speech with ~120ms latency and 15-language support
Capabilities
Cost
Community model (estimated from hardware time)
Input Parameters
| Name | Type | Description | Default | Constraints |
|---|---|---|---|---|
text* | string | The text to convert to speech. Maximum 2,000 characters. Supports SSML break tags for pauses (e.g. `<break time="1s" />`), emotion markups (e.g. `[happy]`, `[sad]`), and non-verbal vocalizations (e.g. `[laugh]`, `[sigh]`). | — | — |
audio_format | string | Output audio format. | "mp3" | mp3wavogg_opusflac |
sample_rate | integer | Audio sample rate in Hz. | 48000 | 8000160002205024000320004410048000 |
speaking_rate | number | Speaking speed multiplier. Set to 0 for normal speed (1.0). | 0 | min: 0, max: 1.5 |
temperature | number | Controls randomness when generating audio. Higher values produce more expressive results, lower values are more deterministic. Set to 0 to use the model default (1.1). | 0 | min: 0, max: 2 |
text_normalization | string | Controls whether numbers, dates, and abbreviations are expanded before synthesis. 'auto' lets the model decide, 'on' always normalizes, 'off' reads text as-is. | "auto" | autoonoff |
voice_id | string | The voice to use. Use a preset voice name (e.g. 'Ashley', 'Dennis', 'Alex') or a custom cloned voice ID. | "Ashley" | — |
textrequiredstringThe text to convert to speech. Maximum 2,000 characters. Supports SSML break tags for pauses (e.g. `<break time="1s" />`), emotion markups (e.g. `[happy]`, `[sad]`), and non-verbal vocalizations (e.g. `[laugh]`, `[sigh]`).
audio_formatstringOutput audio format.
"mp3"sample_rateintegerAudio sample rate in Hz.
48000speaking_ratenumberSpeaking speed multiplier. Set to 0 for normal speed (1.0).
0min: 0, max: 1.5temperaturenumberControls randomness when generating audio. Higher values produce more expressive results, lower values are more deterministic. Set to 0 to use the model default (1.1).
0min: 0, max: 2text_normalizationstringControls whether numbers, dates, and abbreviations are expanded before synthesis. 'auto' lets the model decide, 'on' always normalizes, 'off' reads text as-is.
"auto"voice_idstringThe voice to use. Use a preset voice name (e.g. 'Ashley', 'Dennis', 'Alex') or a custom cloned voice ID.
"Ashley"787e87b8a178Updated: 8/1/2026129.3K runs
cinemasetfree