50 models7 providersprices scraped nightly — no estimates
google · gemma family

Gemma 2 9B IT

9B params · 8K context · gemma · released 2024-07

Gemma 2 9B IT is Google DeepMind's mid-tier Gemma 2 model, a 9-billion-parameter instruction-tuned transformer released July 2024. It performs comparably to Llama 3.1 8B and Mistral 7B v0.3 on standard benchmarks, and Groq hosts it with some of the lowest latency numbers available for sub-10B models. Pricing across providers typically lands below $0.20 per million tokens, making it cost-competitive for high-throughput classification, extraction, and short-form generation. The 8K context window is the same bottleneck as the rest of the Gemma 2 line — document-QA or retrieval-augmented workloads that need 32K or more should look elsewhere. License is Gemma, permissive but not OSI-approved.

Cheapest now
no hosted pricing yet
90-day price Δ
cheapest-host trend
Fastest TTFT
no data
Providers
0
indexed providers

No pricing data available for this model yet.

No hosted pricing scraped for this model yet. Methodology · Raw data