Definition
Gemini Flash-Lite is the Flash-Lite tier within Google DeepMind's Gemini model family. Google positions it for low-latency, low-cost serving and high-throughput tasks. Its versions have their own capabilities and prices.
Tasks suited to this tier can include extraction or document parsing when the chosen release meets the application's quality threshold. The threshold belongs to the application and its evaluation; the Lite label does not establish it.
Origin and attribution
Google introduced Gemini 2.0 Flash-Lite on February 5, 2025 in public preview. Koray Kavukcuoglu announced it on behalf of Google DeepMind's Gemini team, alongside separate Flash and Pro Experimental releases. Later Flash-Lite releases kept the tier name while changing features and operating characteristics.
Scope and limitations
The current Gemini 3.5 Flash-Lite documentation describes a multimodal model for high-volume workflows. It accepts text, image, video, audio, and PDF inputs and produces text. It also supports thinking and several tool interfaces. These are documented capabilities of 3.5 Flash-Lite; earlier and later releases need their own checks.
The same model page distinguishes supported input modalities from media generation and Live API support. Those endpoint-specific limits must be retained when building a workflow. Support for structured output controls representation without establishing that the extracted facts or selected values are correct.
Operational significance
Evaluate total cost per accepted result. A cheaper call can require retries, human correction, or escalation that changes its operational cost. Track missed errors alongside output format compliance, and test the request distribution the model will receive.
Flash-Lite can serve an early stage in a model cascade when an independent acceptance rule determines whether to stop or escalate. Define an explicit outcome when the evidence is inadequate. A stronger follow-up model also needs validation.
Distinguish it from nearby terms
- Gemini is the complete family; Flash-Lite is one tier within it.
- Flash is a separate tier with its own releases and operating characteristics.
- A reasoning model describes inference behavior. Some Flash-Lite releases support thinking, so Lite and non-reasoning cannot be treated as synonyms.
- Throughput measures completed work per unit time; low per-call cost or latency alone does not establish accepted-task throughput.
Check your understanding
Flash-Lite reduces API spending on document extraction, but more results need manual correction. Which measurements would show whether the switch improved the workflow, and what evidence should trigger escalation?