$299/month
- 8M credits/month for STT/TTS models plus $299/month prepaid for Line Voice Agents
Loading the latest directory information.
Start here for the decision-making essentials: what Cartesia does, who it is for, how it is accessed, and the first-party sources behind this profile.
$299/month
Cartesia AI is a voice AI platform built on State Space Models, offering real-time TTS (Sonic) and STT (Ink-2/Ink-Whisper) with 40ms latency across 44 languages. It supports cloud, on-premise, and on-device deployment with SOC 2, HIPAA, GDPR, and PCI compliance.
Read-only public captures of Cartesia’s homepage. Screenshots are dated, never live embeds, and open full-screen.
Short answers to the questions buyers and builders commonly ask about Cartesia. Each answer cites the shared ledger below, where every source is listed once.
AI Dubbing for seamless multilingual video dubbing with fine-grained control over pitch, s · Instant voice cloning from a 3-second clip that preserves speaking style, accent, and emot · Time-to-first-audio of 40ms for AI Dubbing · Provides real-time text-to-speech and streaming speech-to-text models for enterprise voice
Discover AI Dubbing for seamless multilingual video dubbing with Cartesia
Financial-services voice agents support real-time outbound verification of suspicious tran · Customer service automation including 24/7 self-service, IVR modernization, and outbound c · Voice agents · Dubbing / localized voiceovers
Real-time outbound verification calls on suspicious transactions, step-up authentication
Cartesia's services are designed only for U.S. users and are not intended for users outsid
The Services are designed for users in the United States only and are not intended for users located outside the United States.
Cartesia charges for credits and agent minutes used, with unlimited workspace seats on eve · Scale | $299/month | 8M credits/month plus $299/month in prepaid Managed Agent dollars
Pay for the credits and agent minutes you use, with unlimited workspace seats on every plan.
The Python TTS API provides a 40ms time-to-first-audio. · Single API call to add Cartesia voices to voice agents (no re-engineering required)
Experience a blazing fast 40ms time-to-first-audio with our TTS API for seamless interactions.
Quora's Poe platform integration - users can convert text to audio using Cartesia Sonic wi · ServiceNow AI Voice Agents (part of ServiceNow's AI Experience) · Voice reader integrates seamlessly into applications
Quora brings multilingual voice to millions of AI chats with Cartesia
This profile connects the jobs Cartesia is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.
Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.
HTTP 200 verified twice
Cartesia | AI that learns and interacts like humans
Cartesia AI, Inc. is the company name · Models built on State Space Models (SSMs) from Stanford research · Sonic is the real-time streaming TTS model
Cartesia AI, Inc.
Cartesia's models are built on State Space Models for low-latency, long-context reasoning and efficiency at scale.
Sonic-3.6 is Cartesia's real-time streaming text-to-speech model.
Ink-2 is Cartesia's speech-to-text model.
Cartesia's TTS supports natural, expressive voices with laughter in 44 languages for AI agents and interactive applications.
The Python TTS API provides a 40ms time-to-first-audio.
Financial-services voice agents support real-time outbound verification of suspicious transactions and step-up authentication.
Cartesia charges for credits and agent minutes used, with unlimited workspace seats on every plan.
Scale | $299/month | 8M credits/month plus $299/month in prepaid Managed Agent dollars
A concise view of the jobs, capabilities and integrations described in the recorded product sources.
Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.
A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.
This profile is being checked for dated releases and material product changes. Nothing appears here until the exact date and event can be verified from a recorded source.
Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.
Cartesia supports deploying the same models and agents across cloud, on-premise, and on-device environments.
Cartesia is SOC 2 Type II certified.
Cartesia's services are designed only for U.S. users and are not intended for users outside the United States.
AI Dubbing for seamless multilingual video dubbing with fine-grained control over pitch, speed, and emotion
Instant voice cloning from a 3-second clip that preserves speaking style, accent, and emotion
Time-to-first-audio of 40ms for AI Dubbing
Sonic supports 44 languages with a wide range of accents and native-speaker quality voices
Sonic - low-latency voice model for lifelike speech generation
Ink-2 - speech-to-text model for real-time voice agents, ranked #1 on Artificial Analysis's streaming leaderboard with the most accurate built-in turn detection of any provider
Line - modern voice agent development platform built to be code-first
SOC2 compliant with unlimited concurrency for enterprise customers
GDPR compliant text-to-speech platform
Quora's Poe platform integration - users can convert text to audio using Cartesia Sonic with 100+ default voices and 14 languages via the '@Cartesia' bot command
$64 million Series A led by Kleiner Perkins
Mamba-3B-SlimPJ released in partnership with Cartesia and Together under an Apache 2.0 license
Provides real-time text-to-speech and streaming speech-to-text models for enterprise voice AI applications
Sonic (text-to-speech, 90ms latency)
Ink-Whisper (fast streaming speech-to-text optimized for real-time transcription)
Sonic 3 (newest model, 27 new languages, custom pronunciations, speed and volume controls)
State-space models (SSMs)
Built on years of research at Stanford
ServiceNow AI Voice Agents (part of ServiceNow's AI Experience)
Single API call to add Cartesia voices to voice agents (no re-engineering required)
Fully air-gapped on-prem deployment of voice models, inference engine, and orchestration
India data center for India-resident processing and data residency
SOC-2 Type 2, HIPAA, GDPR, and PCI compliance
40+ languages with native support (not English models adapted for other languages)
Ultra-low latency voice generation with 40ms time-to-first-audio