model · OpenAI

gpt-oss-120b

gpt-oss-120b is OpenAI's flagship open-weight model — a mixture-of-experts with 117B total parameters but only ~5.1B active per token. Thanks to its native 4-bit MXFP4 format the weights are about 61 GB, so it runs on a single 80 GB datacentre GPU or, remarkably, a 128 GB AMD Ryzen AI mini-PC at roughly 31 tokens/sec. It is the most capable model a single person can realistically self-host, with a 128K context and an Apache-2.0 licence.

  • free
  • self-host
  • Commercial

What it can help with

  • mixture-of-experts
  • 4-bit quantization
  • self-hosted
  • single-gpu
  • high-context
  • open-source licence
  • amd mini-pc
  • token throughput