provider · Cerebras Systems
Cerebras
Cerebras is a hardware + inference company that runs open-weight models (Llama 3.x and Llama 4 Scout/Maverick, Qwen3, DeepSeek-R1 distills, gpt-oss) on its wafer-scale CS-3 system — a single chip the size of a dinner plate (the WSE-3) with thousands of times the memory bandwidth of a GPU. The result is the fastest inference measured anywhere: roughly 1,800–2,600 tokens/sec on Llama models and ~3,000 tok/s on gpt-oss-120B, several times faster than GPU-based providers and even faster than Groq on the larger models. For a biology student the payoff is the same as Groq but more extreme: a long answer is finished before you finish reading the question, and a 50-paper batch job in n8n runs in well under a minute. The free tier is unusually generous — 1,000,000 tokens per day with no credit card — and the API is a drop-in for OpenAI's: change the base_url to https://api.cerebras.ai/v1 and your existing code works unchanged.
What it can help with
- fast inference
- open-weight models
- wafer-scale chip
- api compatibility
- generous free tier
- real-time loops