Gemini API pricing in 2026: 3.8 Flash, 3.1 Pro, Flash-Lite, Live, caching, batch, and free tier
Updated: Sep 17

Gemini API pricing in September 2026 spans several distinct model classes. The strongest current price-performance reference is Gemini 3.8 Flash at an introductory $0.75 input / $3.75 output per million tokens through December 31, 2026.
Gemini 3.1 Pro Preview costs substantially more and uses a two-band long-context price schedule, while Flash-Lite models target low-cost high-volume work. Free-tier availability also differs by model.
··········
THE CURRENT GEMINI API PRICE MAP
[[ADPLUS_300x250_1]]
Google prices Gemini models by input, output and sometimes modality, with separate rates for context caching, batch, priority inference, grounding and storage. Comparing only one headline token rate can understate production cost.
........
Model | Standard input / 1M | Standard output / 1M | Free tier |
|---|---|---|---|
Gemini 3.8 Flash | $0.75 through Dec. 31, 2026 | $3.75 through Dec. 31, 2026 | Yes |
Gemini 3.1 Pro Preview <=200K prompt | $2.00 | $12.00 | No paid-model free tier |
Gemini 3.1 Pro Preview >200K prompt | $4.00 | $18.00 | No paid-model free tier |
Gemini 3.1 Flash-Lite | $0.25 text/image/video | $1.50 | Yes |
Gemini 3.8 Live | $0.75 text; modality-specific audio/video rates | $4.50 text; modality-specific audio | Yes |
........
··········
GEMINI 3.8 FLASH HAS INTRODUCTORY 2026 PRICING
Through December 31, 2026, Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens at Standard rates. Google says those rates become $1.50 and $7.50 on January 1, 2027.
··········
GEMINI 3.1 PRO USES LONG-CONTEXT PRICE BANDS
Gemini 3.1 Pro Preview costs $2 input and $12 output per million tokens when the prompt is 200K tokens or less. Above 200K prompt tokens, the rates rise to $4 input and $18 output, so long-context workloads can change unit economics materially.
··········
FLASH-LITE IS THE LOW-COST END OF THE API
Gemini 3.1 Flash-Lite is priced at $0.25 per million text/image/video input tokens and $1.50 per million output tokens, with audio input priced separately. It targets high-volume agentic tasks, translation and simpler data processing.
··········
DS CALCULATION: NORMALIZED 10M INPUT + 2M OUTPUT
Using Standard rates and assuming Pro requests stay at or below 200K prompt tokens, the same 10M-input + 2M-output workload costs about $5.50 on Gemini 3.1 Flash-Lite, $15 on Gemini 3.8 Flash and $44 on Gemini 3.1 Pro Preview. This excludes caching, grounding and tool fees.
··········
CACHING, BATCH, SEARCH, AND PRIORITY INFERENCE CHANGE TOTAL COST
Context caching can reduce repeated-input cost, batch rates can cut eligible offline workloads, and Google Search grounding is charged separately after included monthly search requests. Priority inference also carries its own higher rate schedule.
··········
CONSUMER GEMINI PLANS DO NOT OFFSET API TOKEN CHARGES
Google AI Plus, Pro and Ultra are consumer subscriptions. API usage is billed through the developer platform under model-specific pricing and should be modeled as a separate cost center.
··········
FOLLOW US FOR MORE
··········
DATA STUDIOS
··········


