top of page

Gemini API pricing in 2026: 3.8 Flash, 3.1 Pro, Flash-Lite, Live, caching, batch, and free tier

Sep 30, 2025
2 min read

Updated: Sep 17

Gemini API pricing in September 2026 spans several distinct model classes. The strongest current price-performance reference is Gemini 3.8 Flash at an introductory $0.75 input / $3.75 output per million tokens through December 31, 2026.


Gemini 3.1 Pro Preview costs substantially more and uses a two-band long-context price schedule, while Flash-Lite models target low-cost high-volume work. Free-tier availability also differs by model.


··········

THE CURRENT GEMINI API PRICE MAP


[[ADPLUS_300x250_1]]



Google prices Gemini models by input, output and sometimes modality, with separate rates for context caching, batch, priority inference, grounding and storage. Comparing only one headline token rate can understate production cost.


........

Model

Standard input / 1M

Standard output / 1M

Free tier

Gemini 3.8 Flash

$0.75 through Dec. 31, 2026

$3.75 through Dec. 31, 2026

Yes

Gemini 3.1 Pro Preview <=200K prompt

$2.00

$12.00

No paid-model free tier

Gemini 3.1 Pro Preview >200K prompt

$4.00

$18.00

No paid-model free tier

Gemini 3.1 Flash-Lite

$0.25 text/image/video

$1.50

Yes

Gemini 3.8 Live

$0.75 text; modality-specific audio/video rates

$4.50 text; modality-specific audio

Yes

........


··········

GEMINI 3.8 FLASH HAS INTRODUCTORY 2026 PRICING


Through December 31, 2026, Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens at Standard rates. Google says those rates become $1.50 and $7.50 on January 1, 2027.


··········

GEMINI 3.1 PRO USES LONG-CONTEXT PRICE BANDS


Gemini 3.1 Pro Preview costs $2 input and $12 output per million tokens when the prompt is 200K tokens or less. Above 200K prompt tokens, the rates rise to $4 input and $18 output, so long-context workloads can change unit economics materially.


··········

FLASH-LITE IS THE LOW-COST END OF THE API


Gemini 3.1 Flash-Lite is priced at $0.25 per million text/image/video input tokens and $1.50 per million output tokens, with audio input priced separately. It targets high-volume agentic tasks, translation and simpler data processing.


··········

DS CALCULATION: NORMALIZED 10M INPUT + 2M OUTPUT


Using Standard rates and assuming Pro requests stay at or below 200K prompt tokens, the same 10M-input + 2M-output workload costs about $5.50 on Gemini 3.1 Flash-Lite, $15 on Gemini 3.8 Flash and $44 on Gemini 3.1 Pro Preview. This excludes caching, grounding and tool fees.


··········

CACHING, BATCH, SEARCH, AND PRIORITY INFERENCE CHANGE TOTAL COST


Context caching can reduce repeated-input cost, batch rates can cut eligible offline workloads, and Google Search grounding is charged separately after included monthly search requests. Priority inference also carries its own higher rate schedule.


··········

CONSUMER GEMINI PLANS DO NOT OFFSET API TOKEN CHARGES


Google AI Plus, Pro and Ultra are consumer subscriptions. API usage is billed through the developer platform under model-specific pricing and should be modeled as a separate cost center.


··········


FOLLOW US FOR MORE


··········


DATA STUDIOS


··········


Recent Posts

See All
bottom of page