provider · Cerebras Systems

Cerebras

Cerebras is a hardware + inference company that runs open-weight models (Llama 3.x and Llama 4 Scout/Maverick, Qwen3, DeepSeek-R1 distills, gpt-oss) on its wafer-scale CS-3 system — a single chip the size of a dinner plate (the WSE-3) with thousands of times the memory bandwidth of a GPU. The result is the fastest inference measured anywhere: roughly 1,800–2,600 tokens/sec on Llama models and ~3,000 tok/s on gpt-oss-120B, several times faster than GPU-based providers and even faster than Groq on the larger models. For a biology student the payoff is the same as Groq but more extreme: a long answer is finished before you finish reading the question, and a 50-paper batch job in n8n runs in well under a minute. The free tier is unusually generous — 1,000,000 tokens per day with no credit card — and the API is a drop-in for OpenAI's: change the base_url to https://api.cerebras.ai/v1 and your existing code works unchanged.

  • free-tier
  • cloud
  • Commercial

What it can help with

  • fast inference
  • open-weight models
  • wafer-scale chip
  • api compatibility
  • generous free tier
  • real-time loops