AI infrastructure · Tool

Modal

Researched

Modal is a serverless cloud platform for AI and data teams offering sub-second cold starts, instant GPU autoscaling, and pay-per-use compute for inference, training, batch processing, and sandboxes.

Online Checked XLinkedInYouTubeXFollow updates
Official site snapshots

See the official site at a glance

Read-only public captures of Modal’s homepage and verified pricing page. Screenshots are dated, never live embeds, and open full-screen.

Visit live site
Pricing page · captured Jul 24, 2026
At a glance

In one minute

Start here for the decision-making essentials: what Modal does, who it is for, how it is accessed, and the first-party sources behind this profile.

Pricing5 options

Pay-per-use compute billed by the CPU cycle

customers never pay for idle resources.

Nvidia H100 GPU tasks are priced at $0.001097 per second.

Starter plan is $0/month plus compute, with $30/month in free credits, up to 3 workspace s

Team plan is $250/month plus compute, with $100/month in free credits, unlimited seats, 10

Platforms
Cloud computing serviceCore Platform
API accessNot public
FoundedNot disclosed by source
AvailabilityWeb / remote
LicenseProprietary

Best suited to

Source-backed fit
AI and data teams building compute-intensive applications Engineers and researchers deploying and serving large language models Startups and organizations scaling GPU workloads Developers who prefer infrastructure-free cloud compute Graduate students, labs, and early-stage startups eligible for free… AI and data teams
Decision support

Common questions and adoption checks

6 sourced answers

Short answers to the questions buyers and builders commonly ask about Modal. Each answer cites the shared ledger below, where every source is listed once.

01What does Modal say it can do?

Run inference, training, batch processing, and sandboxes with sub-second cold starts and i · Autoscale from 0 to 1000+ GPUs instantly · Run any model or inference engine on H100s, A100s, A10Gs and more, with scale-to-zero · SFT, LoRA, and full fine-tunes on B200s, H100s, A100s and more

Run inference, training, batch processing, and sandboxes with sub-second cold starts, instant autoscaling, and a developer experience that feels local.
02Who is Modal intended for?

AI and data teams · Graduate students, labs, and researchers are eligible to apply for up to $10,000 in free M · Early-stage startups can apply for free Modal compute credits. · Engineers and researchers building compute-intensive applications

The serverless platform for AI and data teams.
03What use cases does Modal describe?

Deploy and scale inference for LLMs, audio, and image/video generation · Fine-tune open-source models on single or multi-node clusters · Programmatically scale secure, ephemeral environments for running untrusted code · Language Models

Inference Deploy and scale inference for LLMs, audio, image/video generation.
04What should teams verify before adopting Modal?

Refusing cookies may prevent use of some portions of the Service

If you choose to refuse our cookies, you may not be able to use some portions of our Service
05What pricing information is available for Modal?

Pay-per-use compute billed by the CPU cycle; customers never pay for idle resources. · Nvidia H100 GPU tasks are priced at $0.001097 per second. · Starter plan is $0/month plus compute, with $30/month in free credits, up to 3 workspace s · Team plan is $250/month plus compute, with $100/month in free credits, unlimited seats, 10

you always pay for what you use and nothing more. You never pay for idle resources — just actual compute time, by the CPU cycle.
06Does Modal document API access?

Modal offers a gRPC API as the primary well-described interface for interactions.

Most interactions with Modal are well-described in a gRPC API
Decision guide

Capabilities and operating fit

AI infrastructure

This profile connects the jobs Modal is described as handling with its delivery model, access options and the subjects used to match it to related products in this directory.

Common use cases

  • Deploy and scale inference for LLMs, audio, and image/video generation
  • Fine-tune open-source models on single or multi-node GPU clusters
  • Programmatically scale secure, ephemeral sandboxes for running untrusted code
  • Run large-scale batch workflows and job queues
  • Serve models with hundreds of billions of parameters
  • Deploy OpenAI-compatible LLM services and OpenCode coding agents at scale

Access signals

Pricing model
Pay-per-use compute billed by the CPU cycle; customers never pay for idle resources. · Nvidia H100 GPU tasks are priced at $0.001097 per second. · Starter plan is $0/month plus compute, with $30/month in free credits, up to 3 workspace s · Team plan is $250/month plus compute, with $100/month in free credits, unlimited seats, 10
API
Not publicly listed
Source links
17 recorded
Source-backed

Verified facts

Updated July 22, 2026

Each fact points to a recorded source, making it easy to distinguish verified product information from claims that need checking.

Official website

HTTP 200 verified twice

First-party description

Modal: High-performance AI infrastructure

Source-supported facts

Serverless platform for AI and data teams (modal.com/) · Sub-second cold starts and instant autoscaling (modal.com/) · Autoscale from 0 to 1000+ GPUs, instantly (modal.com/)

Capability

Run inference, training, batch processing, and sandboxes with sub-second cold starts and instant autoscaling

Platform

Serverless platform for AI and data teams

Audience

AI and data teams

Use case

Deploy and scale inference for LLMs, audio, and image/video generation

Use case

Fine-tune open-source models on single or multi-node clusters

View 32 more verified facts
Use case

Programmatically scale secure, ephemeral environments for running untrusted code

Capability

Autoscale from 0 to 1000+ GPUs instantly

Capability

Run any model or inference engine on H100s, A100s, A10Gs and more, with scale-to-zero

Capability

SFT, LoRA, and full fine-tunes on B200s, H100s, A100s and more

Capability

Multi-node training up to 128 B200s with 3200 Gbps Infiniband networking

Capability

Sub-10ms overhead latency from globally distributed compute

Integration

Out-of-the-box support for token streaming, WebRTC, and WebSocket

Pricing

Pay-per-use compute billed by the CPU cycle; customers never pay for idle resources.

[2]modal.com/pricing
Pricing

Nvidia H100 GPU tasks are priced at $0.001097 per second.

[2]modal.com/pricing
Pricing

Starter plan is $0/month plus compute, with $30/month in free credits, up to 3 workspace seats, 100 containers, 10 GPU concurrency, limited Scheduled and Web Functions, real-time metrics and logs, and region selection.

[2]modal.com/pricing
Pricing

Team plan is $250/month plus compute, with $100/month in free credits, unlimited seats, 1000 containers, 50 GPU concurrency, unlimited Scheduled Functions, custom domains, static IP proxy, and deployment rollbacks.

[2]modal.com/pricing
Pricing

Enterprise plan pricing is custom with volume-based discounts, unlimited seats, higher GPU concurrency, and embedded ML engineering services.

[2]modal.com/pricing
Capability

Serverless compute that instantly autoscales up and down based on request volume.

[2]modal.com/pricing
Capability

Modal Sandbox + Notebooks lets users burst CPU and memory as needed without over-allocating in advance, paying only for what they use.

[2]modal.com/pricing
Support

Enterprise plan customers receive support via a private Slack channel.

[2]modal.com/pricing
Security

Enterprise plan includes audit logs, Okta SSO, and HIPAA support.

[2]modal.com/pricing
Distribution

Modal can be transacted through the AWS and GCP marketplaces so customers can use committed spend.

[2]modal.com/pricing
Audience

Graduate students, labs, and researchers are eligible to apply for up to $10,000 in free Modal compute credits.

[2]modal.com/pricing
Audience

Early-stage startups can apply for free Modal compute credits.

[2]modal.com/pricing
Deployment

Serverless cloud platform

[3]modal.com/docs
Audience

Engineers and researchers building compute-intensive applications

[3]modal.com/docs
Capability

Runs generative AI models, large-scale batch workflows, and job queues

[3]modal.com/docs
Capability

Run serverless cloud functions from the browser

[3]modal.com/docs
Capability

Deploy an OpenAI-compatible LLM service with a drop-in replacement for the OpenAI API

[3]modal.com/docs
Capability

Deploy OpenCode agents at scale in secure Sandboxes

[3]modal.com/docs
Capability

Design protein binders with ESMFold2, proposing and evaluating thousands of binders in parallel

[3]modal.com/docs
Capability

Transcribe speech in batches with Whisper, turning audio bytes into text at scale

[3]modal.com/docs
Capability

Serve language models with hundreds of billions of parameters

[3]modal.com/docs
Capability

Fold proteins with Boltz-2 to predict molecular structures and binding affinities using state-of-the-art open-source models

[3]modal.com/docs
Capability

Low-latency serving of interactive LLM applications via SGLang

[3]modal.com/docs
Capability

Serverless WebRTC streaming of YOLO detections on webcam footage in real time

[3]modal.com/docs
Company

Modal Labs

[9]modal.com/legal/privacy-policy
Practical capabilities

What it helps with

8 documented areas

A concise view of the jobs, capabilities and integrations described in the recorded product sources.

Use case

Deploy and scale inference for LLMs, audio, and image/video generation

Use case

Fine-tune open-source models on single or multi-node GPU clusters

Use case

Programmatically scale secure, ephemeral sandboxes for running untrusted code

Use case

Run large-scale batch workflows and job queues

Use case

Serve models with hundreds of billions of parameters

Use case

Deploy OpenAI-compatible LLM services and OpenCode coding agents at scale

Use case

Transcribe speech at scale with Whisper and design protein binders with ESMFold2

Use case

Fine-tune open-source models on single or multi-node clusters

Availability

Where it runs and where to get it

Source checked

Documented product formats, platforms and official distribution destinations. Availability can vary by region and plan.

Cost / license

Pay-per-use compute billed by the CPU cycle; customers never pay for idle resources. · Nvidia H100 GPU tasks are priced at $0.001097 per second. · Starter plan is $0/month plus compute, with $30/month in free credits, up to 3 workspace s · Team plan is $250/month plus compute, with $100/month in free credits, unlimited seats, 10

Application types

Products include Inference, Training, Sandboxes, Batch, and NotebooksModal Sandboxes for executing code in secure environments

Platforms

Serverless platform for AI and data teamsCloud computing serviceCloud infrastructure designed for AI workloadsCore Platform
Implementation details

Adoption notes

DeploymentServerless cloud platform · Burst to thousands of GPUs on demand and scale back down to zero · Fully serverless: Modal hosts everything and charges per second of usage · Serverless WebRTC support
LicenseNot disclosed by source
Model supportwhisper-tiny.en
Data controlEnterprise plan includes audit logs, Okta SSO, and HIPAA support. · No transmission or electronic storage method is 100% secure or reliable · Customer-supplied encryption keys supported · Modal builds software using memory-safe programming languages, including Rust (for worker · Modal makes decisions to minimize its attack surface, with most interactions well-describe
Learning curveIntermediate
Primary use casesDeploy and scale inference for LLMs, audio, and image/video generation, Fine-tune open-source models on single or multi-node GPU clusters, Programmatically scale secure, ephemeral sandboxes for running untrusted code, Run large-scale batch workflows and job queues, Serve models with hundreds of billions of parameters, Deploy OpenAI-compatible LLM services and OpenCode coding agents at scale, Transcribe speech at scale with Whisper and design protein binders with ESMFold2, Fine-tune open-source models on single or multi-node clusters, Programmatically scale secure, ephemeral environments for running untrusted code, Language Models, Image, Video, 3D Processing, Audio Processing, Batch Processing, Sandboxed Code, Computational Bio, Coding Agents

What to verify before adopting

  • Refusing cookies may prevent use of some portions of the Service
Evolution and major updates

Modal timeline

A concise history of software releases and material product changes. Events appear only when a dated source supports what changed.

2 dated updates
Latest first · exact dates

Showing the newest updates and meaningful milestones. Open an entry for its summary and source.

Citation ledger

Recorded sources

18 unique pages

Facts, answers, structured details, milestones and primary resource links cite this shared ledger. Each external page appears once; release tags from the same GitHub project are grouped under one release history.

  1. 1modal.com 16 facts · 3 answers · Official site
  2. 2modal.com/pricing 12 facts · 3 answers · Pricing
  3. 3modal.com/docs 12 facts · 2 answers · Documentation
  4. 4modal.com/docs/examples 14 facts · Documentation
  5. 5modal.com/customers 11 facts · 2 answers · Official site
  6. 6modal.com/products/platform 12 facts · 1 answer · Official site
  7. 7modal.com/docs/guide 10 facts · 1 answer · Documentation
  8. 8modal.com/blog 7 facts · 2 answers · Official site
  9. 9modal.com/legal/privacy-policy 6 facts · 3 answers · Security
  10. 10modal.com/docs/guide/security 4 facts · 1 answer · Documentation
  11. 11modal.com/blog/decagon-case-study 2 milestones
  12. 12modal.com/docs/examples/batched_whisper Documentation
  13. 13modal.com/docs/examples/chatterbox_tts Documentation
  14. 14modal.com/docs/examples/fine_tune_asr Documentation
  15. 15modal.com/docs/examples/generate_music Documentation
  16. 16modal.com/docs/examples/llm-finetuning Documentation
  17. 17modal.com/docs/examples/llm-voice-chat Documentation
  18. 18modal.com/docs/examples/vllm_inference Documentation
Research status120 substantive facts · 17 source pages · quality score 100/100