model · OpenAI
gpt-oss-120b
gpt-oss-120b is OpenAI's flagship open-weight model — a mixture-of-experts with 117B total parameters but only ~5.1B active per token. Thanks to its native 4-bit MXFP4 format the weights are about 61 GB, so it runs on a single 80 GB datacentre GPU or, remarkably, a 128 GB AMD Ryzen AI mini-PC at roughly 31 tokens/sec. It is the most capable model a single person can realistically self-host, with a 128K context and an Apache-2.0 licence.
What it can help with
- mixture-of-experts
- 4-bit quantization
- self-hosted
- single-gpu
- high-context
- open-source licence
- amd mini-pc
- token throughput