Skip to content

Voice & Audio

Copy page
Tool id Vendor Credentials Status Divinci usage
browser-local
Free text-to-speech using your browser's built-in Web Speech API. Uses system voices, works offline.
Browser Divinci-managed available
cartesia-sonic
Premium text-to-speech with 40ms latency. Features emotion control, laughter tags, and instant voice cloning.
Cartesia Divinci-managed available
cloudflare-aura-1
High-quality context-aware text-to-speech powered by Deepgram on Cloudflare Workers AI. 12 voices, multiple audio formats.
Cloudflare Divinci-managed available
cloudflare-aura-2
Enhanced quality text-to-speech with improved pronunciation and natural prosody. 40 English and 10 Spanish voices.
Cloudflare Divinci-managed available
cloudflare-melotts
Multi-lingual text-to-speech by MyShell AI. Supports 16+ languages with natural-sounding voices.
Cloudflare Divinci-managed available
deepgram-tts
Deepgram Aura TTS with 100+ voices across 7 languages, low latency, and pronunciation control.
Deepgram BYOK available
elevenlabs-tts
Highly expressive, natural ElevenLabs voices with 29+ languages and instant voice cloning support.
ElevenLabs BYOK available
openai-tts
OpenAI speech (gpt-4o-mini-tts) with steerable delivery and 10 expressive built-in voices.
OpenAI BYOK available
vertex-ai-chirp3
Google's highest-fidelity Chirp 3 HD voices. Eight expressive voices available across 30+ locales.
Google Divinci-managed available
vertex-ai-neural2
High-quality neural text-to-speech voices from Google Cloud. Natural-sounding with 1M free characters per month.
Google Divinci-managed available
vertex-ai-standard
Basic text-to-speech voices from Google Cloud. Affordable option with 4M free characters per month.
Google Divinci-managed available
vertex-ai-studio
Premium studio-grade text-to-speech voices from Google Cloud. Professional quality for production audio.
Google Divinci-managed available

Usage figures for this category are not collected yet — the weekly job populates them.

cloudflare-aura-2 is the default. browser-local uses the listener's own device — no synthesis bill and no audio leaving the browser, at the cost of whatever voices that device happens to have.

Tool id Vendor Credentials Status Divinci usage
@cf/openai/whisper-large-v3-turbo
Opensource speech to text
OpenAI Divinci-managed available
@openai/whisper-1
Speech to text straight from the source
OpenAI BYOK available
deepgram-nova-3
Deepgram's Nova-3 speech-to-text — fast, accurate, multilingual
Deepgram BYOK available

Usage figures for this category are not collected yet — the weekly job populates them.

Whisper on Workers AI is the default; Deepgram Nova 3 and OpenAI's hosted Whisper are BYOK alternatives. Speech-to-text has a fallback router, so a provider failing mid-job falls through rather than losing the job.

Tool id Vendor Credentials Status Divinci usage
@divinci-ai/pyannote-segmentation
Identifies speaker during time durations
Pyannote Divinci-managed available
@google-cloud/speech
Speech to text including speakers
Google Divinci-managed available
@pyannote/pyannote-segmentation
Identifies speaker during time durations
Pyannote BYOK available

Usage figures for this category are not collected yet — the weekly job populates them.

Diarization answers who spoke when, and runs alongside transcription rather than instead of it. Pyannote is available both on Divinci's infrastructure and on your own key.

Tool id Vendor Credentials Status Divinci usage
@cartesia/voice-clone
Creates an instant voice clone from a short audio clip (minimum 5 seconds)
Cartesia BYOK available
@google/chirp3-instant-custom-voice
Creates a Chirp 3 voice cloning key from ~10s of reference audio plus a voice-talent consent recording. Requires Google allow-list access.
Google Divinci-managed available

Usage figures for this category are not collected yet — the weekly job populates them.

Tool id Vendor Credentials Status Divinci usage
@pyannote/voiceprint
Creates a voice embedding from an audio sample for speaker identification
Pyannote BYOK available

Usage figures for this category are not collected yet — the weekly job populates them.

A voice clone synthesises new speech in a given voice. A voiceprint identifies a speaker across recordings. They are different capabilities with very different consent implications.