Modal: High-performance AI infrastructure

modal.com

Customers | Pricing | Docs | Log In

Tin mới

Engineered for inference.

From the proxy layer to the GPU scheduler, every part of Modal's stack is optimized for how inference workloads actually behave.

Built for the full training loop.

From single-GPU fine-tuning to parallel hyperparameter sweeps to multi-node runs, Modal handles all of your coding infrastructure in a single code file.

Designed to scale agents.

From interactive coding agents to long-running RL rollouts, Modal Sandboxes are the execution layer AI systems need: isolated, flexible, and built to scale.

Security and governance

65% Latency reduction

65% Latency reduction

Real-time, multi-node inference for Runway Characters

Real-time, multi-node inference for Runway Characters

Real-time, multi-node inference for Runway Characters

Real-time robot control running on Modal with 10–15 ms latency.

Real-time robot control running on Modal with 10–15 ms latency.

Real-time robot control running on Modal with 10–15 ms latency.

Real-time, multi-node inference for Runway Characters

4 months faster to launch

4 months faster to launch

Real-time, multi-node inference for Runway Characters

ML‑driven molecular design

ML‑driven molecular design

Real-time, multi-node inference for Runway Characters

Powering AI app generation at scale

Powering AI app generation at scale

Real-time, multi-node inference for Runway Characters

“We’re actively saving 2 engineers’ worth of ongoing time”

“We’re actively saving 2 engineers’ worth of ongoing time”

Real-time, multi-node inference for Runway Characters

3x latency decrease for document processing

Real-time, multi-node inference for Runway Characters

“Modal makes it easy to write code that runs on 100s of GPUs in parallel, transcribing podcasts in a fraction of the time.”

“Modal makes it easy to write code that runs on 100s of GPUs in parallel, transcribing podcasts in a fraction of the time.”

Real-time, multi-node inference for Runway Characters

Transcribe speech in batches with Whisper Turn audio bytes into text at scale

Voice chat with LLMs Build an interactive voice chat app

Transcribe speech with Kyutai STT Stream transcripts at the speed of speech

Make music Turn prompts into music with ACE-Step

Fine-tune Whisper on domain vocab Improve Whisper transcription accuracy on specialized vocabularies with fine-tuning

Deploy a TTS API with Chatterbox Serve text-to-speech with Chatterbox to generate natural audio from text

Serve your own LLM API

Create custom art of your pet

Deploy OpenCode agents in a cloud Sandbox

Video
▶ Video
Video
▶ Video
Video
▶ Video