Voice systems combine transcription, synthesis, orchestration and telephony. Test them with realistic calls, accents, interruptions and failure conditions; a polished demo does not establish production reliability.
01 · Find the fit
Start with the job
Build inbound or outbound voice agents
Generate and localize spoken audio
Transcribe and analyze conversations
02 · Inspect the details
Compare the tradeoffs
Latency and interruption handling
Language, voice and telephony coverage
Consent, recording and escalation controls
03 · Verify before buying
Follow the evidence
Latency measured only in ideal conditions
Unclear consent or voice-cloning safeguards
Weak transfer and fallback behavior
Common evaluation signals
Voice AIFirst-party researchedhuman feedbackvoice AI evaluationemotional intelligencemultimodal research
Profiles in this category
Find your best fit.
Browse every catalogued voice ai profile, then use the evidence and freshness details to judge how much decision context is currently available.
Hume AI is the human evaluation layer for voice, speech, and conversational AI, offering multimodal emotional intelligence across 50+ languages. It provides evaluation APIs, the Kairos platform, Octave TTS and EVI speech-to-speech models, and SDKs for building empathic voice experiences.
Voice AIFirst-party researchedhuman feedbackvoice AI evaluation
PricingFree plan at $0/month with 10,000 TTS characters and 5 minutes of EVI usage · Starter plan at $3/month (first month 50% off) with 30,000 TTS characters · Creator plan at $7/month (regularly $14/month) with 140,000 TTS characters · Pro plan at $70/month with 1,000,000 TTS characters
AccessKairos Platform - builds evaluation suites, simulates agent-to-agent and human-to-agent co
Loudly is an AI music creation platform enabling creators, artists, and brands to generate royalty-free tracks, stems, samples, and remixes via text prompts or parameter inputs, with distribution to major streaming platforms and tiered licensing.
Voice AIFirst-party researchedAI music generationroyalty-free music
PricingFree tier available ("Start Free Now", "Sign up free") · Free plan is $0/month with no credit card required · Personal plan is $8/month billed annually ($96/year) · Pro plan is $24/month billed annually ($288/year)
Tool London, United Kingdom (HQ being expanded via UK Government partnership)
ElevenLabs is a Voice AI platform offering lifelike AI speech generation across 5,000+ voices and 70+ languages. It delivers ElevenCreative, ElevenAgents, and ElevenAPI products with Text to Speech, Speech to Text, voice cloning, music, and sound effects.
AssemblyAI provides AI models to transcribe and understand speech, with its flagship Universal-3.5 Pro handling real-time and pre-recorded audio across 99 languages. It offers Speech-to-Text, Speech Understanding, Guardrails, LLM Gateway, and Voice Agent APIs via cloud or self-hosted deployments.
PricingPay-as-you-go pricing by product, with a free API trial and custom plans available via sal · Pay-as-you-go pricing with no contracts, minimums, or monthly subscriptions · Failed transcripts are not charged · Free tier and credits available
AccessModels, APIs, and infrastructure in one place for building voice into any product on any s
Crowdin is an AI-powered localization platform launched in 2009, helping companies automate translation through 700+ integrations, 100+ file formats, and continuous localization for mobile, web, and software projects.
Pricing30-Day Enterprise Trial available · Free plan available · Pricing based on hosted words (source words multiplied by number of target languages) · 14-day free trial available
Lokalise is a continuous localization and translation management platform that integrates into development workflows to ship localized products faster. It offers AI-powered translation, automated workflows, and integrations with tools like Figma, GitHub, and Jira.
PricingExplorer plan is priced at $144/month. · Growth plan is priced at $375/month. · Advanced plan is priced at $999/month. · Enterprise plan is available via custom demo booking for organizations with bespoke securi
MemoQ is a translation management system offering AI-powered translation automation, computer-assisted translation, and terminology management for enterprises, language service providers, and professional translators.
Voice AIFirst-party researchedTranslation management system (TMS)Computer-assisted translation
PricingSee official pricing
AccessmemoQ on the web for online translation experience
Resemble AI is a generative AI security platform offering real-time, multimodal deepfake detection across text, audio, image, and video for enterprises, deployable on-prem or in the cloud.
Sonix is an AI transcription platform offering 99% accurate speech-to-text in 54+ languages, with translation to 55+ languages, automated captions, summaries, and sentiment analysis. It features a browser editor, RESTful API, and SOC 2 Type 2/HIPAA security since 2017.
Tool Stockholm, Sweden (Repslagargatan 17A, 118 46 Stockholm)
Soundtrap is a cloud-based digital audio workstation (DAW) for creating music and podcasts, featuring real-time collaboration, mixing tools, and a royalty-free sound library. It offers free and paid tiers, developer APIs, and supports external instruments.
Voice AIFirst-party researchedCloud-based DAWMusic production
PricingFree sign-up available · Free tier available at $0 with unlimited projects · Sound Starter plan at $9.59 USD per month ($115.08 billed yearly) · Music Production plan at $13.59 USD per month ($163.08 billed yearly)
AccessWeb-based online studio, no downloads required
Speechify is a voice AI platform offering text-to-speech, voice typing, AI podcasts, voice cloning, and dubbing across iOS, Android, web, Mac, Windows, and browser extensions, with 1000+ voices and 60+ languages on Premium.
Voice AIFirst-party researchedtext to speechvoice typing
PricingPremium plan is $29 per month · Free plan available with speeds up to 1.5x and 10 robotic sounding voices · Yearly billing saves 60% · Offers a free service plus paid services; free users are encouraged to upgrade
AccessiPhone & iPad, Android, Chrome Extension, Edge Extension, Web App, Mac App, Windows App
Splash is an AI music company (Popgun Labs Pty Ltd) offering an interactive online music creation platform with capabilities including text-to-singing, text-to-rap, generative music, voice transfer, composition, mastering, and collaborative remixing, accessible via Roblox, Amazon Alexa, Steam, YouTube, and SoundCloud.
Suno is an AI music generation platform by Suno, Inc. offering a web-based generative audio workstation (Suno Studio) and mobile apps. Users create songs from prompts, lyrics, or melodies with granular style controls, upload audio, and export up to 12 WAV stems for DAW workflows.
Voice AIFirst-party researchedAI music generationgenerative audio workstation
PricingFree tier: 10 free songs per day, free forever with no subscription needed · Pro plan: up to 500 custom songs per month with full commercial rights · Subscription offered monthly or yearly, with yearly saving 20% · Free Plan costs $0 per month
AccessAvailable on web (Suno Studio) and as mobile apps on iPhone and Android
Auphonic is an automatic audio post-production webservice offering AI-powered noise reduction, leveling, loudness normalization, speech-to-text, and multitrack processing for podcasts, educational content, video, and audiobooks.
Beatoven.ai, by Beatoven Private Limited (India), is an AI platform that generates royalty-free background music and sound effects via text-to-music and multimodal prompts. It offers APIs for music search and intelligence, serving creators across video, podcast, and gaming workflows.
Voice AIFirst-party researchedAI music generationText-to-music
PricingFree sign-up available · Sign-up is free · Free signup available
AccessWeb platform hosted at https://beatoven.ai and other sites owned/operated by Beatoven
Rime AI delivers enterprise conversational voice models with 600+ voices across 50+ languages and sub-100ms latency. Offers Mist v3 and Arcana v3 models, voice shaping controls, and cloud, on-prem, or VPC deployment with HIPAA and SOC 2 compliance.
PricingFree trial available ('Try for free') alongside sales-led demo booking · Starter plan priced at $0.05 per 1,000 characters of generated speech (usage-based) · Every new account receives 3,000 free minutes on the Starter plan · Enterprise plans use volume pricing tailored to organizational scale
Vapi is a developer-focused voice AI platform by Vapi, Inc. that enables building, testing, and deploying conversational voice agents in minutes. It scales to millions of calls with sub-500ms latency and offers enterprise-grade compliance and security.
PricingSMS/Chat $0.005/msg volume based · Model Provider Cost (STT, LLM, TTS) at cost, $0 if you bring your own API key · 10 call concurrency included; $10 per line/month additional · HIPAA add-on $2000/month
AccessPlatform for developers creating conversational voice AI
Wondercraft is an AI-native video and audio studio that lets users create videos, podcasts, and voiceovers by chatting with AI. It offers text-to-speech, voice cloning, video generation, and an API, serving 250,000 creatives and teams.
Voice AIFirst-party researchedAI-native video studioAI video
PricingFree plan at $0/month including 150 credits for individual creators · Creator plan priced at $25/month (or $21/month billed yearly) with 1,000 credits · Offers Pricing and Enterprise tiers · Offers Enterprise pricing tier
AccessWonda AI studio for creating video by chatting with AI
Cartesia is a voice AI platform offering real-time text-to-speech (Sonic-3.5), streaming speech-to-text (Ink), and a code-first voice agent development platform (Line). Built on State Space Models, it delivers sub-100ms latency across 40+ languages.
PricingFree plan is $0/month with 20K credits/month and $1 prepaid agents/month. · Pro plan is $5/month with 100K credits/month, commercial use license, and instant voice cl
AccessLine — platform for building and shipping enterprise voice agents
WellSaid is an AI voice generator that creates professional-quality voice overs from text using realistic AI voices. It offers a studio for script-to-voice conversion, collaboration features, API access, and integrations for creative workflows.
Deepgram provides enterprise voice AI APIs including speech-to-text, text-to-speech, and voice agents. It offers real-time and batch processing with cloud and self-hosted deployment options, supporting multilingual conversational capabilities.
Murf AI is a platform for generating human-like speech and deploying conversational AI voice agents, offering text-to-speech APIs and tools for use cases such as receptionists, recruiters, call centers, and customer service.