model · Meta

Llama 4 Scout

Llama 4 Scout is Meta's open mixture-of-experts model — 109B total parameters but only 17B active per token — built for a very large context window and strong general use. Quantized, it fits on a single high-end (80 GB) GPU. It is multimodal and broadly capable, backed by Meta's huge ecosystem of fine-tunes and tooling. The catch is licensing: Llama 4 uses Meta's Community Licence, which (unlike Apache/MIT) adds conditions for the largest platforms.

  • free
  • self-host
  • Commercial

What it can help with

  • mixture of experts
  • large context window
  • quantized model
  • single gpu inference
  • multimodal support
  • open licence