model · Google DeepMind
Gemini 3.1 Flash-Lite
Gemini 3.1 Flash-Lite is Google's ultra-cheap, low-latency tier — about $0.25 / $1.50 per 1M tokens — and unlike most budget models it is fully multimodal (text, image, video, audio and PDF in) with a 1M-token context window. That makes it a strong cheap leg for high-volume multimodal work: bulk document, image or transcript processing. It runs on Google's cloud, with EU data residency available through Vertex AI regions (e.g. Frankfurt). Closed and cloud-only; step up to Gemini 3.1 Pro when a task needs the frontier tier.
What it can help with
- budget multimodal
- high-volume processing
- large-context window
- document extraction
- image analysis
- audio transcription
- cloud-only deployment
- eu data residency