$199/month
- per connection, SAML/OIDC, available on Contract and PAYG plans
Loading the latest directory information.
AssemblyAI provides speech-to-text, speech understanding, and voice agent APIs powered by the Universal-3.5 Pro model. It offers transcription in 99 languages, real-time streaming, LLM Gateway access, and enterprise-grade security with SOC 2, ISO 27001, and HIPAA compliance.
Start here for the decision-making essentials: what AssemblyAI does, who it is for, how it is accessed, and the first-party sources behind this profile.
$199/month
Read-only public captures of AssemblyAI’s homepage and verified pricing page. Screenshots are dated, never live embeds, and open full-screen.
Short answers to the questions buyers and builders commonly ask about AssemblyAI. Each answer cites the shared ledger below, where every source is listed once.
Flagship speech-to-text model Universal-3.5 Pro with highest accuracy to date across 18 la · Voice Agent sessions for real-time voice agents, with features like voice_focus background · Speech Understanding and Guardrails features (speaker diarization, speaker identification, · Voice Agent API offers low-latency speech-to-speech with turn detection, tool calling, and
Universal-3.5 Pro is now our flagship speech model, with the highest accuracy we've released to date across 18 languages.
Startups through Fortune 500 companies processing millions of audio minutes
From startups to Fortune 500s, see how companies process millions of audio minutes with AssemblyAI.
Inference platform powering enterprise Voice AI products, running 800M+ API calls per mont · Build Voice AI apps
Every feature runs on the same industry-leading Voice AI models and inference platform behind 800M+ API calls a month, with the accuracy, availability, and compliance controls enterprise deployments require.
Pre-recorded Speech-to-Text Universal-3.5 Pro | $0.21/hr | async, 18 languages, native cod · Realtime Speech-to-Text Universal-3.5 Pro Realtime | $0.45/hr | 18 languages, voice agents · Voice Agent API | $4.50/hr ($0.075/min) | managed orchestration, STT, LLM, TTS, turn detec · Voice Agent API billed at $4.50/hr ($0.075/min), all-inclusive of STT, LLM, TTS, turn dete
Universal-3.5 Pro The most accurate async speech-to-text model that transcribes every conversation exactly as it's heard – works across 18 languages, with native code switching and our most accurate speaker diarization yet. | $0.21 /hr
Realtime streaming endpoint at api_host="streaming.assemblyai.com" with streaming client S · Streaming WebSocket endpoint wss://streaming.assemblyai.com/v3/ws and Voice Agent WebSocke
client = StreamingClient( StreamingClientOptions( api_key=API_KEY, api_host="streaming.assemblyai.com", ) )
Documentation includes an integration prompt usable with Cursor, Claude Code, Copilot, and · Voice Agent API uses cascaded STT + LLM + TTS over self-hosted LiveKit; Twilio SIP integra · Documentation exposes an integration prompt usable from Cursor, Claude Code, Copilot, Devi · Voice Agent API works with Twilio, LiveKit, and any telephony provider out of the box
Pin the integration prompt URL in your agent's AGENTS.md / CLAUDE.md / .cursorrules, or copy the full prompt directly. Works with Cursor, Claude Code, Copilot, Devin, and any other AI coding assistant.
This profile connects the jobs AssemblyAI is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.
Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.
HTTP 200 verified twice
AssemblyAI | AI models to transcribe and understand speech
Voice AI platform for speech-to-text and speech understanding · Transcription supports up to 99 languages, with translation into 86 languages · Universal-3.5 Pro async speech-to-text priced at $0.21/hr across 18 languages
AssemblyAI
Voice AI products needing transcription, understanding, and agentic voice workflows from MVP to production scale.
Transcription supports up to 99 languages.
Pre-recorded Speech-to-Text Universal-3.5 Pro | $0.21/hr | async, 18 languages, native code switching, speaker diarization
Realtime Speech-to-Text Universal-3.5 Pro Realtime | $0.45/hr | 18 languages, voice agents
Voice Agent API | $4.50/hr ($0.075/min) | managed orchestration, STT, LLM, TTS, turn detection, telephony included
Realtime streaming endpoint at api_host="streaming.assemblyai.com" with streaming client SDK (assemblyai.streaming.v3).
Documentation includes an integration prompt usable with Cursor, Claude Code, Copilot, and Devin AI coding assistants.
Global infrastructure with redundancy, processing over 2 million hours of audio per day.
A concise view of the jobs, capabilities and integrations described in the recorded product sources.
Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.
A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.
Showing the newest updates and meaningful milestones. Open an entry for its summary and source.
AssemblyAI launched the Sync API for short-audio transcription: one HTTP request returns a finished Universal-3.5 Pro transcript without polling, WebSockets, chunking, or job management.
View source [17]AssemblyAI expanded Universal to 99 languages at a flat $0.27 per hour, with automatic language detection across all languages and speaker diarization in 95 languages.
View source [16]Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.
Support via support ticket form or in-app live chat; 400k+ developers building, 1B+ audio files processed, 99.9% uptime.
Voice AI platform for speech-to-text and speech understanding
Flagship speech-to-text model Universal-3.5 Pro with highest accuracy to date across 18 languages
Voice Agent sessions for real-time voice agents, with features like voice_focus background suppression and transcription_prompt domain biasing
Startups through Fortune 500 companies processing millions of audio minutes
AES-256 encryption at rest and TLS 1.3 in transit by default
SOC 2 Type 2 certified
PCI-DSS 4.0 Level 1 compliant as of March 31, 2025
GDPR compliant with completed third-party assessment
Data processing centers in the United States and Dublin, Ireland, with self-serve EU Data Residency via the customer dashboard
99.9% uptime provided to all contracted customers
Speech Understanding and Guardrails features (speaker diarization, speaker identification, summarization, sentiment analysis, entity detection, topic detection, language detection, PII redaction, content moderation) on a single API request
Universal-3.5 Pro async model billed at $0.21/hr
Universal-3.5 Pro Realtime streaming model billed at $0.45/hr base
Voice Agent API billed at $4.50/hr ($0.075/min), all-inclusive of STT, LLM, TTS, turn detection, interruption handling, and tool calling
$50 in free credits on signup, no credit card required
HIPAA BAA available with no premium, PCI-DSS on Voice Agent API, ISO 27001, SOC 2 Type 2, and GDPR included
Automatic language detection and transcription across 99 languages; translation into 86 languages in the same request
Streaming WebSocket endpoint wss://streaming.assemblyai.com/v3/ws and Voice Agent WebSocket endpoint wss://agents.assemblyai.com/v1/ws
Voice Agent API uses cascaded STT + LLM + TTS over self-hosted LiveKit; Twilio SIP integration planned for Q2 2026
Documentation exposes an integration prompt usable from Cursor, Claude Code, Copilot, Devin, and other AI coding assistants via AGENTS.md / CLAUDE.md / .cursorrules
Inference platform powering enterprise Voice AI products, running 800M+ API calls per month
Voice Agent API offers low-latency speech-to-speech with turn detection, tool calling, and unified billing at $4.50/hr
Pre-recorded STT delivers clean, customizable transcripts in 99 languages with industry-leading accuracy and natural language prompting
Real-time STT for notetakers, agents, and captions
Sync STT returns a transcript in milliseconds from a single HTTP request with no polling or session management
LLM Gateway is an OpenAI-compatible API for every frontier model with automatic fallbacks, zero markup, and zero data retention
PII Redaction automatically masks sensitive information before it reaches the LLM