AI infrastructure · Tool

Groq

Researched

Groq delivers fast, low-cost AI inference through its custom LPU silicon, pioneered in 2016 as the first chip purpose-built for inference. Its GroqCloud platform serves 3M developers and enterprises via tokens-as-a-service pricing, OpenAI-compatible APIs, and global data center deployment.

Online Checked
Official site snapshots

See the official site at a glance

Read-only public captures of Groq’s homepage. Screenshots are dated, never live embeds, and open full-screen.

Visit live site
Homepage · captured Jul 21, 2026
At a glance

In one minute

Start here for the decision-making essentials: what Groq does, who it is for, how it is accessed, and the first-party sources behind this profile.

Pricing4 options

Tokens-as-a-Service on-demand pricing with per-million-token input and output rates

GPT OSS 120B (128k context) priced at $0.15 input / $0.60 output per million tokens at 500

Llama 3.1 8B Instant (128k context) priced at $0.05 input / $0.08 output per million token

Llama 3.3 70B Versatile (128k context) priced at $0.59 input / $0.79 output per million to

Platforms
GroqCloud powers 1M+ developers
API accessNot public
FoundedGroq pioneered the LPU in 2016 as the first chip purpose-built for inference.
AvailabilityWeb / remote
LicenseProprietary

Best suited to

Source-backed fit
Developers building latency-sensitive AI applications Enterprises needing large-scale inference capacity and dedicated support Teams migrating from OpenAI with minimal code changes Startups running inference-based AI businesses rather than training LLMs High-volume workloads suited to async batch processing Serves both developers and enterprises
Decision support

Common questions and adoption checks

6 sourced answers

Short answers to the questions buyers and builders commonly ask about Groq. Each answer cites the shared ledger below, where every source is listed once.

01What does Groq say it can do?

Groq delivers fast, low-cost inference via the LPU. · Fast, low-cost inference delivered by the Groq LPU · Fast, low cost inference via the Groq LPU for developers · Fast, low-cost AI inference

Groq is fast, low cost inference. The Groq LPU delivers inference with the speed and cost developers need.
02Who is Groq intended for?

Serves both developers and enterprises · Fortune 500 companies and more than 1.4 million developers building real-time AI applicati

fast, low latency API access to state-of-the-art models for everyone from the developer to the enterprise
03What use cases does Groq describe?

Running inference-based AI businesses (as opposed to training LLMs) · AI-generated text detection · Industry solutions for Healthcare, Defense, Finance, Entertainment, Gaming, and Telecommun · Retrieval Augmented Generation (RAG)

For training LLMs, there are options. But to run an inference-based AI business, you need something different. Groq was the right choice.
04What should teams verify before adopting Groq?

Cached tokens do not count towards GroqCloud rate limits

All cached tokens don't count towards GroqCloud rate limits.
05What pricing information is available for Groq?

Tokens-as-a-Service on-demand pricing with per-million-token input and output rates · GPT OSS 120B (128k context) priced at $0.15 input / $0.60 output per million tokens at 500 · Llama 3.1 8B Instant (128k context) priced at $0.05 input / $0.08 output per million token · Llama 3.3 70B Versatile (128k context) priced at $0.59 input / $0.79 output per million to

Groq On-Demand Pricing for Tokens-as-a-Service | Groq is fast, low cost inference.
06Does Groq document API access?

A free API key is available to developers · Developer Tier offers rate limits up to 10x higher than the free tier. · Free API key available for developers · OpenAI-compatible API endpoint at https://api.groq.com/openai/v1/responses, allowing migra

Free API key
Decision guide

Capabilities and operating fit

AI infrastructure

This profile connects the jobs Groq is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.

Common use cases

  • Running inference-based AI products at low latency worldwide
  • Pay-as-you-go inference via Developer Tier with a credit card
  • Migrating from OpenAI by changing three lines of code
  • High-volume parallel API workloads via the Batch API
  • Tool-enabled AI applications through Remote MCP integrations
  • Web automation, search, scraping, and invoicing via MCP servers like BrowserBase, Exa, Fir

Access signals

Pricing model
Tokens-as-a-Service on-demand pricing with per-million-token input and output rates · GPT OSS 120B (128k context) priced at $0.15 input / $0.60 output per million tokens at 500 · Llama 3.1 8B Instant (128k context) priced at $0.05 input / $0.08 output per million token · Llama 3.3 70B Versatile (128k context) priced at $0.59 input / $0.79 output per million to
API
Not publicly listed
Source links
17 recorded
Source-backed

Verified facts

Updated July 22, 2026

Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.

Official website

HTTP 200 verified twice

First-party description

Groq is fast, low cost inference.

Source-supported facts

Fast, low-cost inference delivered by the Groq LPU · Groq pioneered the LPU in 2016, the first chip purpose-built for inference · Custom silicon rather than GPUs alone

Capability

Groq delivers fast, low-cost inference via the LPU.

Founded

Groq pioneered the LPU in 2016 as the first chip purpose-built for inference.

Team size

Groq is used by 3M developers and teams.

Platform

GroqCloud is the cloud platform/console for running inference on Groq's LPU-based stack.

Architecture

Groq uses its own custom LPU silicon rather than relying on GPUs alone.

View 32 more verified facts
Deployment

Groq's LPU-based stack runs in data centers across the world for low-latency inference.

Pricing

Tokens-as-a-Service on-demand pricing with per-million-token input and output rates

[11]groq.com/pricing
Api

A free API key is available to developers

[11]groq.com/pricing
Pricing

GPT OSS 120B (128k context) priced at $0.15 input / $0.60 output per million tokens at 500 TPS

[11]groq.com/pricing
Pricing

Llama 3.1 8B Instant (128k context) priced at $0.05 input / $0.08 output per million tokens at 840 TPS

[11]groq.com/pricing
Pricing

Llama 3.3 70B Versatile (128k context) priced at $0.59 input / $0.79 output per million tokens at 394 TPS

[11]groq.com/pricing
Pricing

Developer Tier is a self-serve, pay-as-you-go plan requiring only a credit card to access GroqCloud.

[3]groq.com/blog/developer-tier-now-available-on-groqcloud
Api

Developer Tier offers rate limits up to 10x higher than the free tier.

[3]groq.com/blog/developer-tier-now-available-on-groqcloud
Pricing

The Batch API provides a 25% cost discount with guaranteed 24-hour processing time for high-volume parallel workloads.

[3]groq.com/blog/developer-tier-now-available-on-groqcloud
Model

Flex Tier (beta) supports llama-3.3-70b-versatile and llama-3.1-8b-instant models with up to 10x rate limits.

[3]groq.com/blog/developer-tier-now-available-on-groqcloud
Integration

GroqCloud offers OpenAI endpoint compatibility, allowing migration by changing three lines of code.

[3]groq.com/blog/developer-tier-now-available-on-groqcloud
Support

Groq provides an Enterprise tier for large-scale needs with enterprise-grade capacity and dedicated support.

[3]groq.com/blog/developer-tier-now-available-on-groqcloud
Capability

Fast, low-cost inference delivered by the Groq LPU

[8]groq.com/contact
Founded

2016

[8]groq.com/contact
Company

Groq is focused on inference

[8]groq.com/contact
Api

Free API key available for developers

[8]groq.com/contact
Support

Developer support via community resources and GroqCloud Dev Console chat; enterprise inquiries handled by sales team

[8]groq.com/contact
Capability

Fast, low cost inference via the Groq LPU for developers

[12]groq.com/security
Security

Groq is committed to protecting customers, systems, and data

[12]groq.com/security
Security

Proactively identifies and mitigates security risks using industry-standard practices

[12]groq.com/security
Security

Operates a private vulnerability disclosure program through HackerOne

[12]groq.com/security
Security

Eligible vulnerability submissions may be recognized or rewarded at Groq's discretion

[12]groq.com/security
Support

Groq Trust Center provides information on the security program, compliance posture, and supporting documentation

[12]groq.com/security
Company

Groq was established in 2016 focused on inference

[15]groq.com/blog/groq-vercel-partner-to-make-building-fast-and-simple
Capability

Fast, low-cost AI inference

[15]groq.com/blog/groq-vercel-partner-to-make-building-fast-and-simple
Audience

Serves both developers and enterprises

[15]groq.com/blog/groq-vercel-partner-to-make-building-fast-and-simple
Release date

Sep 23, 2025

[7]groq.com/changelog
Integration

Remote Model Context Protocol (MCP) server integration available in Beta on GroqCloud, connecting AI models to thousands of external tools through Anthropic's open MCP standard

[7]groq.com/changelog
Api

OpenAI-compatible API endpoint at https://api.groq.com/openai/v1/responses, allowing migration from OpenAI with zero code changes

[7]groq.com/changelog
Tool use

Remote MCP enables developers to connect any remote MCP server to models hosted on GroqCloud, enabling tool capabilities

[7]groq.com/changelog
Model

Remote MCP is available on all models that support tool use, including openai/gpt-oss-20b, openai/gpt-oss-120b, moonshotai/kimi-k2-instruct-0905, qwen/qwen3-32b, meta-llama/llama-4-maverick-17b-128e-instruct, meta-llama/llama-4-scout-17b-16

[7]groq.com/changelog
Integration

Supported MCP server integrations include BrowserBase, Browser Use, Exa, Firecrawl, HuggingFace, Parallel, Stripe, and Tavily

[7]groq.com/changelog
Practical capabilities

What it helps with

8 documented areas

A concise view of the jobs, capabilities and integrations described in the recorded product sources.

Use case

Running inference-based AI products at low latency worldwide

Use case

Pay-as-you-go inference via Developer Tier with a credit card

Use case

Migrating from OpenAI by changing three lines of code

Use case

High-volume parallel API workloads via the Batch API

Use case

Tool-enabled AI applications through Remote MCP integrations

Use case

Web automation, search, scraping, and invoicing via MCP servers like BrowserBase, Exa, Fir

Use case

Enterprise-grade custom inference solutions

Use case

Running inference-based AI businesses (as opposed to training LLMs)

Availability

Where it runs and where to get it

Source checked

Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.

Cost / license

Tokens-as-a-Service on-demand pricing with per-million-token input and output rates · GPT OSS 120B (128k context) priced at $0.15 input / $0.60 output per million tokens at 500 · Llama 3.1 8B Instant (128k context) priced at $0.05 input / $0.08 output per million token · Llama 3.3 70B Versatile (128k context) priced at $0.59 input / $0.79 output per million toGroq API Cookbook is an open-source community repository · compound-beta integrates multiple open-source models

Platforms

GroqCloud is the cloud platform/console for running inference on Groq's LPU-based stack.GroqCloud powers 1M+ developersGroq LPU, a custom-built inference chip, delivered via GroqCloudGroqCloud™ Dev Console is the platform where models from the Llama 3.2 suite are made avai
Implementation details

Adoption notes

DeploymentGroq's LPU-based stack runs in data centers across the world for low-latency inference.
LicenseGroq API Cookbook is an open-source community repository · compound-beta integrates multiple open-source models
Model supportFlex Tier (beta) supports llama-3.3-70b-versatile and llama-3.1-8b-instant models with up · Remote MCP is available on all models that support tool use, including openai/gpt-oss-20b, · Llama 4 inference is available live on GroqCloud · Llama 4 (Meta) available via the official Llama API on Groq · compound-beta and compound-beta-mini
Data controlGroq is committed to protecting customers, systems, and data · Proactively identifies and mitigates security risks using industry-standard practices · Operates a private vulnerability disclosure program through HackerOne · Eligible vulnerability submissions may be recognized or rewarded at Groq's discretion · Compliance-ready inference for handling classified defense documents, protected health inf
Learning curveIntermediate
Primary use casesRunning inference-based AI products at low latency worldwide, Pay-as-you-go inference via Developer Tier with a credit card, Migrating from OpenAI by changing three lines of code, High-volume parallel API workloads via the Batch API, Tool-enabled AI applications through Remote MCP integrations, Web automation, search, scraping, and invoicing via MCP servers like BrowserBase, Exa, Fir, Enterprise-grade custom inference solutions, Running inference-based AI businesses (as opposed to training LLMs), AI-generated text detection, Industry solutions for Healthcare, Defense, Finance, Entertainment, Gaming, and Telecommun, Retrieval Augmented Generation (RAG), Building real-time AI agents and agentic workflows with minimal configuration, Production AI workloads and real-time AI applications requiring consistent low latency and, The Llama 3.2 11B and 90B vision models on Groq support image reasoning use cases.

What to verify before adopting

  • Cached tokens do not count towards GroqCloud rate limits
Evolution and major updates

Groq timeline

A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.

Research in progress
Scheduled for research

Building a reliable release history.

This profile is being checked for dated releases and material product changes. Nothing appears here until the exact date and event can be verified from a recorded source.

Dated event Short explanation Original source
Citation ledger

Recorded sources

17 unique pages

Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.

  1. 1groq.com 10 facts · 1 answer · Official site
  2. 2groq.com/blog 8 facts · 1 answer · Official site
  3. 3groq.com/blog/developer-tier-now-available-on-groqcloud 6 facts · 2 answers · Official site
  4. 4groq.com/blog/gpt-oss-improvements-prompt-caching-and-lower-pricing 6 facts · 2 answers · Pricing
  5. 5groq.com/blog/how-to-build-your-own-ai-research-agent-with-one-groq-api-call 6 facts · 1 answer · Documentation
  6. 6groq.com/blog/retrieval-augmented-generation-with-groq-api 6 facts · 1 answer · Documentation
  7. 7groq.com/changelog 6 facts · 1 answer · Official site
  8. 8groq.com/contact 5 facts · 2 answers · Official site
  9. 9groq.com/customer-stories 6 facts · 1 answer · Official site
  10. 10groq.com/industry-solutions 6 facts · 1 answer · Official site
  11. 11groq.com/pricing 5 facts · 2 answers · Pricing
  12. 12groq.com/security 6 facts · 1 answer · Security
  13. 13groq.com/blog/meta-and-groq-continue-to-build-open-source-developer-ecosystem-as-llama-3-2-launches 6 facts · Official site
  14. 14groq.com/newsroom/meta-and-groq-collaborate-to-deliver-fast-inference-for-the-official-llama-api 5 facts · 1 answer · Documentation
  15. 15groq.com/blog/groq-vercel-partner-to-make-building-fast-and-simple 3 facts · 2 answers · Official site
  16. 16groq.com/blog/thank-you-1m-developers-building-with-groqcloud 5 facts · Documentation
  17. 17groq.com/blog/the-official-llama-api-accelerated-by-groq 5 facts · Documentation
Research status99 substantive facts · 17 source pages · quality score 100/100