JBFqnCBsd6RMkjVDRZzb — that you select in the dashboard or pass in API requests. IntuneVoice maintains a library of 10,000+ voices. You can also clone a voice from an audio recording or generate one from a text description.Documentation
IntuneVoice Documentation
Explore our docs and guides to integrate IntuneVoice
How IntuneVoice works
IntuneVoice provides AI voice infrastructure: text-to-speech, speech-to-text, voice cloning, conversational agents, and generative audio. You can use it in four ways, suited to different audiences.
IntuneCreative is a no-code web application where creators, producers, and editors generate voiceovers, music, dubs, and studio projects directly in the browser.
IntuneAgents is the platform for designing and operating conversational voice agents, with a visual builder for non-technical users and full programmatic control for developers.
IntuneAPI exposes every capability as a REST interface with official Python and TypeScript SDKs, so developers can embed voice into their own applications and workflows.
IntuneVoice Reception is a ready-to-deploy AI phone receptionist for small and medium businesses that answers calls, books appointments, and manages day-to-day operations from a single dashboard.
Concepts
intune_v3 produces the most expressive output across 70+ languages. intune_flash_v2_5 targets real-time use at ~75ms latency. Each capability — speech-to-text, music, sound effects — has its own dedicated model.Choose your path
Meet the models
Intune v3
Our most emotionally rich, expressive speech synthesis model
- Dramatic delivery and performance
- 70+ languages supported
- 5,000 character limit
- Support for natural multi-speaker dialogue
Intune Multilingual v2
Lifelike, consistent quality speech synthesis model
- Natural-sounding output
- 29 languages supported
- 10,000 character limit
- Most stable on long-form generations
Intune Flash v2.5
Our fast, affordable speech synthesis model
- Ultra-low latency (~75ms†)
- 32 languages supported
- 40,000 character limit
- Faster model, 50% lower price per character for API generations
Scribe v2
State-of-the-art speech recognition model
- Accurate transcription in 90+ languages
- Keyterm prompting, up to 1000 terms
- Entity detection, up to 56
- Precise word-level timestamps
- Speaker diarization, up to 32 speakers
- Dynamic audio tagging
- Smart language detection
Scribe v2 Realtime
Real-time speech recognition model
- Accurate transcription in 90+ languages
- Real-time transcription
- Low latency (~150ms†)
- Precise word-level timestamps
† Excluding application & network latency
Browse by capability
Built withIntuneVoice