Gemma 2 9B IT
Gemma 2 9B IT is Google DeepMind's mid-tier Gemma 2 model, a 9-billion-parameter instruction-tuned transformer released July 2024. It performs comparably to Llama 3.1 8B and Mistral 7B v0.3 on standard benchmarks, and Groq hosts it with some of the lowest latency numbers available for sub-10B models. Pricing across providers typically lands below $0.20 per million tokens, making it cost-competitive for high-throughput classification, extraction, and short-form generation. The 8K context window is the same bottleneck as the rest of the Gemma 2 line — document-QA or retrieval-augmented workloads that need 32K or more should look elsewhere. License is Gemma, permissive but not OSI-approved.
No pricing data available for this model yet.
Not enough history to display chart. View full history →
No benchmark data available.
Questions developers ask
How much does it cost to run Gemma 2 9B IT for 100M tokens?▾
Pricing varies by provider. Use the workload calculator to get an accurate estimate for your specific usage pattern.
What is the cheapest provider for Gemma 2 9B IT?▾
No pricing data is available yet. Check back after the next scrape run.
What context window does Gemma 2 9B IT support?▾
Gemma 2 9B IT supports a context window of 8,192 tokens. Individual providers may cap this lower.
What's the cheapest way to run Gemma 2 9B IT?▾
Compare provider prices in the table above and consider using prompt caching if supported.
Is there a free tier for Gemma 2 9B IT?▾
Free tiers vary by provider and change frequently. Check each provider's current pricing page for trial credits or free-tier limits. The prices shown on this page reflect paid API access.
How much does Gemma 2 9B IT cost for 1M tokens?▾
Pricing data for Gemma 2 9B IT is not yet available. Check back after the next data refresh.
Keep exploring
No hosted pricing scraped for this model yet. Methodology · Raw data