HTTP 200 verified twice
Finding your next stop…
Loading the latest directory information.
Loading the latest directory information.
Start here for the decision-making essentials: what Patronus AI does, who it is for, how it is accessed, and the first-party sources behind this profile.
$10 / 1k small evaluator API calls
Patronus AI develops simulation research and infrastructure to accelerate progress toward human-aligned AGI. It offers digital world models and AI evaluation/benchmarking tools across software development, customer service, finance, research science, and product applications.
Read-only public captures of Patronus AI’s homepage. Screenshots are dated, never live embeds, and open full-screen.
Short answers to the questions buyers and builders commonly ask about Patronus AI. Each answer cites the shared ledger below, where every source is listed once.
Offers digital world models leading on coding, dialogue, research, and general tool use · Designing scenarios that target core model skills to uncover learning signals (Deep Resear · Custom eval model fine-tuning and eval dataset generation (AI Services) · AI evaluation, benchmarking, and development support testing for safety, accuracy, relevan
Patronus AI offers digital world models leading on coding, dialogue, research, and general tool use across public benchmarks
Software development — delivering software across many tools, services, and frameworks · Customer service — handling complex customer support cases end-to-end · Finance — M&A, private equity, quantitative trading, and strategic finance · Research science — exploring research comprehension, synthesis, and extension
Software Development Delivering software across many tools, services, and frameworks
Base plan limits: 2-week data access, 2 Projects, 5 Experiments per Project, with Logs and
Patronus Experiments Access to last 2 weeks 2 Projects 5 Experiments per Project Patronus Logs Access to last 2 weeks Patronus Traces Access to last 2 weeks
$10 / 1k small evaluator API calls · $20 / 1k large evaluator API calls · $10 / 1k eval explanations · $10 in free credits to get started
$10 / 1k small evaluator API calls
Patronus API (optional add-on) for evaluator calls and eval explanations
Patronus API (optional)
Webhooks (Premium Platform Features)
Premium Platform Features Patronus Evaluation Runs, webhooks.
This profile connects the jobs Patronus AI is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.
Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.
HTTP 200 verified twice
Patronus AI | Simulating the World's Intelligence
Mission: develop simulation research and infrastructure to accelerate progress toward huma · Offers digital world models leading on coding, dialogue, research, and general tool use ac · Benchmarks covered: InterCode, τ-bench, BFCL-v4, Toolathlon, SpeedrunBench, FinanceBench
Develop simulation research and infrastructure to accelerate progress toward human-aligned AGI
Offers digital world models leading on coding, dialogue, research, and general tool use
InterCode, τ-bench, BFCL-v4, and Toolathlon
Software development: delivering software across many tools, services, and frameworks
Customer service: handling complex customer support cases end-to-end
Finance: M&A, private equity, quantitative trading, and strategic finance
Research science — exploring research comprehension, synthesis, and extension
Product applications: navigating UI/UX across web and mobile applications
Designing scenarios that target core model skills to uncover learning signals (Deep Research, Multi-Turn Dialogue, Long Horizon, Memory)
A concise view of the jobs, capabilities and integrations described in the recorded product sources.
Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.
A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.
This profile is being checked for dated releases and material product changes. Nothing appears here until the exact date and event can be verified from a recorded source.
Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.
Patronus API (optional add-on) for evaluator calls and eval explanations
$10 / 1k small evaluator API calls
$20 / 1k large evaluator API calls
$10 / 1k eval explanations
$10 in free credits to get started
On-prem / dedicated VPC available
Custom data retention and SSO
Webhooks (Premium Platform Features)
Custom eval model fine-tuning and eval dataset generation (AI Services)
Base plan limits: 2-week data access, 2 Projects, 5 Experiments per Project, with Logs and Traces limited to last 2 weeks
Patronus AI, Inc.
Catching LLM mistakes at scale; testing existing and custom language models
Percival (Core Platform)
RL Envs (RL Environments)
Does not train models on customer data
Does not share customer data with foundation model companies such as OpenAI
Infrastructure managed on AWS
$50M Series B announced June 25, 2026
Digital World Model for AI Agent Training and Simulation
Financial Services use case supported
General contact available at contact@patronus.ai
Security contact available at security@patronus.ai
Patronus AI, Inc.
Develops simulation research and infrastructure to accelerate progress toward human-aligned AGI
Customer service chatbot evaluation for quality, context retrieval, hallucination, summarization, and safety
AI evaluation, benchmarking, and development support testing for safety, accuracy, relevance, behavior alignment, and multimodal use cases
LLM evaluation against FinanceBench on the Patronus AI platform to detect hallucinations on financial questions
Multi-step evaluation: assessing appropriate planning, delegation, and execution behaviors required for task completion