$249/month
- platform fee best suited for small teams of up to 5 people with increased usage limits
Loading the latest directory information.
Start here for the decision-making essentials: what Braintrust does, who it is for, how it is accessed, and the first-party sources behind this profile.
$249/month
Braintrust is an enterprise-grade AI observability and evaluation platform for tracing production LLM applications, running experiments on datasets, comparing prompts and models, and catching regressions via online scoring and quality gates.
Read-only public captures of Braintrust’s homepage and verified pricing page. Screenshots are dated, never live embeds, and open full-screen.
Short answers to the questions buyers and builders commonly ask about Braintrust. Each answer cites the shared ledger below, where every source is listed once.
AI observability platform for tracing production, running evals, and catching regressions · Run experiments against real datasets, compare prompts and models side-by-side, and score · Discovery surfaces patterns in real time across task, issues, and sentiment, with online s · Enterprise-grade AI evaluation platform for building reliable LLM applications
Braintrust is the AI observability platform for tracing production, running evals, and catching regressions before they reach users.
Datasets are versioned collections of test cases that power repeatable evaluations and cap · Topics custom facets support domain-specific categories such as churn risk, feature reques
Versioned collections of test cases that power repeatable evaluations and capture real production behavior as your application evolves.
Pro | $249/month | best suited for small teams of up to 5 people with increased usage limi · Free starter plan available with no credit card required, including free usage for traces,
When you sign up for the Pro plan, you'll immediately be charged a prorated amount of the monthly $249 platform fee.
Gateway provides a unified API across OpenAI, Anthropic, Google, and AWS providers · Listed on the Vercel Marketplace · Tracing for Claude Code, Codex, OpenCode, and pi runs through the bt CLI · Braintrust supports Azure AI Gateway as an AI provider, handling models that use the OpenA
You get a unified API across OpenAI, Anthropic, Google, AWS, and other providers, with automatic caching and observability on every request.
Hybrid deployment available; Brainstore data plane can be deployed on your own infrastruct · Supports hybrid deployment, letting customers decide where their AI data lives · Braintrust Lambda Extension gives Python and TypeScript/JavaScript Lambda functions a loca
Hybrid deployment Deploy Brainstore data plane on your own infrastructure
SOC 2 Type II certified with independently audited security controls verified annually · HIPAA compliant with full compliance with HIPAA requirements to secure PII · GDPR compliant with full compliance with EU data protection regulations · SSO / SAML integration with identity providers for seamless authentication
SOC 2 Type II Independently audited security controls verified annually
This profile connects the jobs Braintrust is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.
Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.
HTTP 200 verified twice
Braintrust - The active observability platform for agents
AI observability platform for tracing production, evals, and catching regressions · Enterprise-grade AI evaluation platform for reliable LLM applications · Run experiments against real datasets, compare prompts and models side-by-side
AI observability platform for tracing production, running evals, and catching regressions before they reach users
Run experiments against real datasets, compare prompts and models side-by-side, and score outputs with LLMs, code, or humans
Discovery surfaces patterns in real time across task, issues, and sentiment, with online scoring catching regressions and quality gates blocking bad releases
SOC 2 Type II certified with independently audited security controls verified annually
HIPAA compliant with full compliance with HIPAA requirements to secure PII
GDPR compliant with full compliance with EU data protection regulations
SSO / SAML integration with identity providers for seamless authentication
Hybrid deployment available; Brainstore data plane can be deployed on your own infrastructure
Pro | $249/month | best suited for small teams of up to 5 people with increased usage limits and longer data retention
A concise view of the jobs, capabilities and integrations described in the recorded product sources.
Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.
A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.
Showing the newest updates and meaningful milestones. Open an entry for its summary and source.
Shipped official integrations that let coding agents instrument code with Braintrust, run evals, and read production traces during development.
View source [1]Released Harbor-powered sandboxed execution for agent evals, isolating code-based scorers and agent runs for safer, reproducible evaluation.
View source [1]Published behavior specs as an open standard for defining and supervising expected behavior in long-horizon agentic systems.
View source [1]Engineering post documenting the architecture behind Braintrust's continuous, real-time trace intelligence pipeline operating at scale.
View source [5]Promoted Topics, the automatic pattern-discovery feature for production AI traces, from preview to general availability across all Braintrust plans.
View source [16]Introduced a free Starter pricing plan, broadening free-tier access to Braintrust's observability and evaluation platform for individual builders.
View source [1]Braintrust published a product update allowing customers to decide where their AI data lives for security and residency purposes.
View source [1]Introduced a security model letting customers choose where their AI data is stored, addressing residency and compliance requirements across regions.
View source [1]Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.
Free starter plan available with no credit card required, including free usage for traces, evals, and your whole team
DeveloperApplication offered as a Web platform for AI model evaluation, LLM observability, performance monitoring, debugging, experiment tracking, dataset management, prompt engineering, and A/B testing
Braintrust Data, Inc., headquartered at 548 Market St PMB 96611, San Francisco, California 94104-5401, founded in 2022
Enterprise-grade AI evaluation platform for building reliable LLM applications
2023
Agent observability that traces every request, scores responses, and catches regressions in production
Gateway provides a unified API across OpenAI, Anthropic, Google, and AWS providers
SDK auto-instrumentation logs every supported provider call with inputs, outputs, latency, tokens, and cost
SOC 2 Type II compliance
$5M seed round (December 2023)
$36M Series A (October 2024)
Series B announced February 2026
Cookbook recipes are open-source self-contained examples hosted on GitHub
Listed on the Vercel Marketplace
Supports hybrid deployment, letting customers decide where their AI data lives
Loop is a natural-language agent for analyzing logs, optimizing prompts, building datasets, generating scorers, finding similar traces, and filing support tickets
Tracing for Claude Code, Codex, OpenCode, and pi runs through the bt CLI
Braintrust supports Azure AI Gateway as an AI provider, handling models that use the OpenAI Chat Completions, OpenAI Responses, or Anthropic Messages API
Braintrust Lambda Extension gives Python and TypeScript/JavaScript Lambda functions a local handoff path for traces
The Braintrust MCP server exposes write tools so coding agents can author prompts, scorers, classifiers, dashboards, alerts, scheduled jobs, evals and dataset rows
Braintrust Gateway lets teams add any AI provider or framework not natively supported
Datasets are versioned collections of test cases that power repeatable evaluations and capture real production behavior, stored in a modern data warehouse with no storage limits
Online scoring rules support Group scope, letting related multi-turn traces be evaluated as a single unit keyed by a session
Kimi K3, DeepSeek V4 Flash 0731, and GLM-5.2 are built-in open-source models served by Braintrust under the Braintrust provider, usable via playgrounds, prompts, scorers, and the Gateway
Loop agent supports Claude, GPT, and Gemini models, with recommended models labeled 'Recommended' and newly evaluated models labeled 'Beta'
Monitoring views became dashboards with a dedicated page per dashboard, project-scoped list, auto-save for chart, filter, and grouping changes, and the ability to clone the built-in Cost and quality dashboard
A dataset row can reference a group of up to 64 traces instead of a single trace, with each trace rendered inline and flagged when unavailable
Generates customizable React components from natural-language descriptions of the desired trace or dataset view interface via Loop
Braintrust introduced Topics as a way to automatically discover what matters in production traces, preceding its general availability in June 2026.
View source [1]Braintrust announced Topics, a feature that automatically clusters and labels production traces to surface patterns for evaluation.
View source [1]Announced Topics, an active observability feature that automatically surfaces patterns in production traces and turns them into evaluable signals.
View source [1]Released a security update that lets Braintrust customers decide where their AI data lives, addressing data residency requirements.
View source [1]Braintrust highlighted debugging capabilities in its Logs product, demonstrating trace-driven debugging patterns.
View source [1]Braintrust documented how Claude Code integrates with the platform for engineering workflows.
View source [5]Engineering integration post documenting how Claude Code hooks into Braintrust logging and evaluation workflows.
View source [5]Braintrust extended its observability SDK coverage beyond Python and TypeScript, broadening language support for AI engineering teams.
View source [5]Announced broadened multi-language SDK coverage for the Braintrust platform, reducing the need for a Python proxy when running AI on JVM and other stacks.
View source [5]Braintrust framed Brainstore as the foundation that makes AI observability at scale possible, supporting high-volume trace ingestion and querying.
View source [1]Product post positioning Brainstore as the storage layer enabling terabyte-scale AI observability inside the Braintrust platform.
View source [1]Brainstore, Braintrust's purpose-built database for high-scale AI observability, became generally available, enabling AI-grade ingestion and querying at scale.
View source [1]Engineering post explaining how Loop was designed as a team-sport eval tool that generates prompts and scorers from production data.
View source [1]Braintrust launched Loop, a team-oriented agentic feature that generates better prompts, scorers, and datasets to improve AI agents.
View source [1]Braintrust launched Loop, a product that uses production data to generate improved prompts, scorers, and datasets for AI applications.
View source [1]Braintrust introduced Loop, an agent that turns production signals into better prompts, scorers, and datasets.
View source [1]Launched Loop, an agent that monitors production traces, surfaces issues, and feeds improvements back into prompts, scorers, and datasets.
View source [1]Launched the Braintrust Java SDK, bringing AI observability and evaluation capabilities to JVM-based applications.
View source [1]Shipped the Braintrust Java SDK, extending native observability and evaluation support to JVM-based AI applications beyond Python and TypeScript.
View source [1]Braintrust became available on the Vercel Marketplace, simplifying adoption for Vercel-deployed AI applications.
View source [1]Released Braintrust on the Vercel Marketplace, making its observability and evals available as an installable integration for Vercel-deployed AI apps.
View source [1]Braintrust announced product capabilities under the 'AI that knows your data' theme, enabling AI experiences grounded in customer data.
View source [1]Braintrust shared its canonical agent architecture—a while-loop-with-tools pattern—as an engineering reference for AI agent design.
View source [5]Engineering overhaul of the experiments UI delivering roughly 10x faster performance for evaluation workflows.
View source [5]Engineering update delivering a 10x faster Experiments UI for running and comparing AI evaluations.
View source [5]Engineering post describing how Braintrust redesigned its observability pipeline to be resilient to failures and network issues.
View source [5]Launched Brainstore, a database designed for high-scale AI workloads, reporting median query times under one second across terabytes and 80x faster benchmarks than traditional observability stores.
View source [20]