app · vLLM project contributors
GuideLLM
GuideLLM is a Python‑based framework that evaluates large language model deployments under realistic, production‑like traffic patterns. It runs on Linux or macOS (Python 3.10–3.13) and can be installed via pip or used from a multi‑arch Docker/Podman image. The tool is free and open source, generating detailed JSON/CSV/HTML reports for latency, throughput and token‑level metrics.
What it can help with
- benchmarking
- load testing
- inference latency
- throughput
- vllm