model · Google DeepMind

Gemini 3.1 Flash-Lite

Gemini 3.1 Flash-Lite is Google's ultra-cheap, low-latency tier — about $0.25 / $1.50 per 1M tokens — and unlike most budget models it is fully multimodal (text, image, video, audio and PDF in) with a 1M-token context window. That makes it a strong cheap leg for high-volume multimodal work: bulk document, image or transcript processing. It runs on Google's cloud, with EU data residency available through Vertex AI regions (e.g. Frankfurt). Closed and cloud-only; step up to Gemini 3.1 Pro when a task needs the frontier tier.

  • paid
  • cloud
  • Commercial

What it can help with

  • budget multimodal
  • high-volume processing
  • large-context window
  • document extraction
  • image analysis
  • audio transcription
  • cloud-only deployment
  • eu data residency