TTS is Microsoft's text to speech model. Microsoft Azure's text-to-speech service providing neural voice synthesis across a wide range of languages and voice styles.
Specifications
Canonical IDmicrosoft-tts
TypeText to Speech
StatusActive
CreatorMicrosoftMicrosoft
Providers
Input ModalitiesText
Output ModalitiesAudio

Capabilities

Input1/5
Text
Image·
Audio·
Video·
PDF·
Output1/5
Text·
Image·
Audio
Video·
Embedding·
Capabilities0/13
Reasoning·
Adaptive Reasoning·
Function Calling·
Parallel Function Calling·
Structured Outputs·
Native JSON Schema·
Web Search·
URL Context·
Computer Use·
Code Execution·
File Search·
Prompt Caching·
Assistant Prefill·

Pricing by Provider

US Dollar ($)
Per 1M tokens
ProviderStandard
Audio In
$ / 1K chars
Azure AI Foundry logo
Azure AI Foundry
azure/speech/azure-tts
$0.015

Cost Calculator

US Dollar ($)
Preset:

Versions

VersionReleasedContextInput / 1MOutput / 1MStatus
Inworld Realtime TTS 2Available
Step TTS 2Available
StyleTTS 2Available
TTS HD 2.5Available
TTS 1$15.00Available
TTS 1 HD$30.00Available
Inworld Realtime TTS 1.5 MaxAvailable
Inworld Realtime TTS 1.5 MiniAvailable
Inworld TTS 1 MaxAvailable
Inworld TTS 1.5 MaxAvailable
TTSCurrent

Model IDs

azure-neural
azure/speech/azure-tts
microsoft-tts