Voice & Audio
Text-to-speech
Section titled “Text-to-speech”| Tool id | Vendor | Credentials | Status | Divinci usage |
|---|---|---|---|---|
browser-localFree text-to-speech using your browser's built-in Web Speech API. Uses system voices, works offline. | Browser | Divinci-managed | available | — |
cartesia-sonicPremium text-to-speech with 40ms latency. Features emotion control, laughter tags, and instant voice cloning. | Cartesia | Divinci-managed | available | — |
cloudflare-aura-1High-quality context-aware text-to-speech powered by Deepgram on Cloudflare Workers AI. 12 voices, multiple audio formats. | Cloudflare | Divinci-managed | available | — |
cloudflare-aura-2Enhanced quality text-to-speech with improved pronunciation and natural prosody. 40 English and 10 Spanish voices. | Cloudflare | Divinci-managed | available | — |
cloudflare-melottsMulti-lingual text-to-speech by MyShell AI. Supports 16+ languages with natural-sounding voices. | Cloudflare | Divinci-managed | available | — |
deepgram-ttsDeepgram Aura TTS with 100+ voices across 7 languages, low latency, and pronunciation control. | Deepgram | BYOK | available | — |
elevenlabs-ttsHighly expressive, natural ElevenLabs voices with 29+ languages and instant voice cloning support. | ElevenLabs | BYOK | available | — |
openai-ttsOpenAI speech (gpt-4o-mini-tts) with steerable delivery and 10 expressive built-in voices. | OpenAI | BYOK | available | — |
vertex-ai-chirp3Google's highest-fidelity Chirp 3 HD voices. Eight expressive voices available across 30+ locales. | Divinci-managed | available | — | |
vertex-ai-neural2High-quality neural text-to-speech voices from Google Cloud. Natural-sounding with 1M free characters per month. | Divinci-managed | available | — | |
vertex-ai-standardBasic text-to-speech voices from Google Cloud. Affordable option with 4M free characters per month. | Divinci-managed | available | — | |
vertex-ai-studioPremium studio-grade text-to-speech voices from Google Cloud. Professional quality for production audio. | Divinci-managed | available | — |
Usage figures for this category are not collected yet — the weekly job populates them.
cloudflare-aura-2 is the default. browser-local uses the listener's own
device — no synthesis bill and no audio leaving the browser, at the cost of
whatever voices that device happens to have.
Transcription
Section titled “Transcription”| Tool id | Vendor | Credentials | Status | Divinci usage |
|---|---|---|---|---|
@cf/openai/whisper-large-v3-turboOpensource speech to text | OpenAI | Divinci-managed | available | — |
@openai/whisper-1Speech to text straight from the source | OpenAI | BYOK | available | — |
deepgram-nova-3Deepgram's Nova-3 speech-to-text — fast, accurate, multilingual | Deepgram | BYOK | available | — |
Usage figures for this category are not collected yet — the weekly job populates them.
Whisper on Workers AI is the default; Deepgram Nova 3 and OpenAI's hosted Whisper are BYOK alternatives. Speech-to-text has a fallback router, so a provider failing mid-job falls through rather than losing the job.
Speaker diarization
Section titled “Speaker diarization”| Tool id | Vendor | Credentials | Status | Divinci usage |
|---|---|---|---|---|
@divinci-ai/pyannote-segmentationIdentifies speaker during time durations | Pyannote | Divinci-managed | available | — |
@google-cloud/speechSpeech to text including speakers | Divinci-managed | available | — | |
@pyannote/pyannote-segmentationIdentifies speaker during time durations | Pyannote | BYOK | available | — |
Usage figures for this category are not collected yet — the weekly job populates them.
Diarization answers who spoke when, and runs alongside transcription rather than instead of it. Pyannote is available both on Divinci's infrastructure and on your own key.
Voice cloning & voiceprints
Section titled “Voice cloning & voiceprints”| Tool id | Vendor | Credentials | Status | Divinci usage |
|---|---|---|---|---|
@cartesia/voice-cloneCreates an instant voice clone from a short audio clip (minimum 5 seconds) | Cartesia | BYOK | available | — |
@google/chirp3-instant-custom-voiceCreates a Chirp 3 voice cloning key from ~10s of reference audio plus a voice-talent consent recording. Requires Google allow-list access. | Divinci-managed | available | — |
Usage figures for this category are not collected yet — the weekly job populates them.
| Tool id | Vendor | Credentials | Status | Divinci usage |
|---|---|---|---|---|
@pyannote/voiceprintCreates a voice embedding from an audio sample for speaker identification | Pyannote | BYOK | available | — |
Usage figures for this category are not collected yet — the weekly job populates them.
A voice clone synthesises new speech in a given voice. A voiceprint identifies a speaker across recordings. They are different capabilities with very different consent implications.
See also
Section titled “See also”- Voice, Phone & SMS
- Recall.ai connector — meeting capture with transcripts.