Gemma 4 31B API Pricing
5 available providers and 5 live price sources tracked, including 1 official provider.
Data snapshot:
30-second decision
- Lowest recorded input price: $0.1300/1M tokens
- Compare official API and relay quotes with their service terms
- Choose this model for everyday production work that needs a balance of capability and cost.
Task fit and model switching guide
Derived from official positioning and current ComputeUnion provider and pricing data.
Discount vs official API
Input-price difference for the same model; negative values are above official price.
Price composition
Input and output remain paired within each provider quote.
Capability radar
Task-fit levels derived from official positioning, not an independent benchmark.
Channel cost and evidence map
X-axis is observed example monthly cost; Y-axis is price-evidence strength.
External capability comparison for related models
Only evidence-backed related models with the same external metric are compared.
- Prices: current official and provider quotes in ComputeUnion, calculated within each provider.
- Capability: official task positioning plus safely matched Artificial Analysis metrics; neither is presented as a ComputeUnion benchmark.
- Evidence: official, live API, platform-submitted, and unreviewed quotes remain visibly separated.
Best-fit tasks
- Best suited to coding.
- Suitable for complex reasoning.
- Best suited to general production workloads.
- Best suited to workloads balancing quality and cost.
When to switch models
- Compare Gemma 4 26B A4B for batch, repeatable, or cost-sensitive workloads.
- Switch to Gemini 2.5 Pro when quality and complex reasoning take priority.
- Compare Gemini 3.1 Flash-Lite for batch, repeatable, or cost-sensitive workloads.
ComputeUnion market view
- ComputeUnion currently tracks 5 providers, including 1 official and 5 live price sources.
- The lowest recorded input price is 86.9% below the official input price; a lower price does not imply better reliability, limits, or refund terms.
Task guidance is generated by fixed rules from official positioning and ComputeUnion market data; it is not an independent benchmark or guarantee.
Provider summary and buying decision
Evidence tier first, then normalized price within each tier.
| Provider / operator | Price evidence | Public operating history | Input / 1M tokens | Output / 1M tokens | API docs | Refund policy | Payment / invoice |
|---|---|---|---|---|---|---|---|
| Cerebrasunknown | Official price sourcePlatform page ↗Observed: | —Unknown | $0.99 | $1.49 | View docs ↗ | Not provided | visa · mastercardunknown |
| Deep Infraunknown | Live API pricePlatform page ↗Observed: | —Unknown | $0.13 | $0.38 | View docs ↗ | Not provided | visa · mastercardunknown |
| Novita AIunknown | Live API pricePlatform page ↗Observed: | —Unknown | $0.14 | $0.40 | View docs ↗ | Not provided | visa · mastercardunknown |
| OpenRouterunknown | Live API pricePlatform page ↗Observed: | —Unknown | $0.14 | $0.40 | View docs ↗ | Not provided | visa · mastercard · cryptounknown |
| Together AIunknown | Live API pricePlatform page ↗Observed: | —Unknown | $0.39 | $0.97 | View docs ↗ | Not provided | visa · mastercardunknown |
Platform-submitted information or domain verification does not mean the price was independently checked and is not a ComputeUnion guarantee of pricing, availability, refunds, compliance, invoices, or service quality.
Gemma 4 31B price history
Observed input-price changes across providers over the last 30/90 days.
Unit: $/1M tokens
Estimate monthly cost
The estimate uses the displayed price without assuming discounts or cache hits.
FAQ
How is Gemma 4 31B API priced?
Input and output tokens are billed separately. The current recorded official baseline is $0.9900 input / $1.4900 output per 1M tokens.
Which provider is cheapest for Gemma 4 31B?
The lowest recorded input price is $0.1300/1M tokens from Deep Infra; verify provider terms and the observation date before production use.
How do official and relay prices differ?
Official and relay channels can differ in price, latency, limits, and service terms; evaluate them separately for production traffic.