GoogleLLMactive

Gemma 4 31B API Pricing

5 available providers and 5 live price sources tracked, including 1 official provider.

Data snapshot:

30-second decision

  • Lowest recorded input price: $0.1300/1M tokens
  • Compare official API and relay quotes with their service terms
  • Choose this model for everyday production work that needs a balance of capability and cost.
Official input / 1M$0.9900
Lowest input / 1M$0.1300
Maximum input gap86.9%Not a service-quality measure
Comparable providers55 live

Task fit and model switching guide

Derived from official positioning and current ComputeUnion provider and pricing data.

Official positioningOfficially positioned as a balance of capability and cost for everyday production work.

Discount vs official API

Input-price difference for the same model; negative values are above official price.

Price composition

Input and output remain paired within each provider quote.

Capability radar

Task-fit levels derived from official positioning, not an independent benchmark.

Channel cost and evidence map

X-axis is observed example monthly cost; Y-axis is price-evidence strength.

External capability comparison for related models

Only evidence-backed related models with the same external metric are compared.

Artificial Analysis
Sources and verification method
  • Prices: current official and provider quotes in ComputeUnion, calculated within each provider.
  • Capability: official task positioning plus safely matched Artificial Analysis metrics; neither is presented as a ComputeUnion benchmark.
  • Evidence: official, live API, platform-submitted, and unreviewed quotes remain visibly separated.

Best-fit tasks

  • Best suited to coding.
  • Suitable for complex reasoning.
  • Best suited to general production workloads.
  • Best suited to workloads balancing quality and cost.

When to switch models

  • Compare Gemma 4 26B A4B for batch, repeatable, or cost-sensitive workloads.
  • Switch to Gemini 2.5 Pro when quality and complex reasoning take priority.
  • Compare Gemini 3.1 Flash-Lite for batch, repeatable, or cost-sensitive workloads.

ComputeUnion market view

  • ComputeUnion currently tracks 5 providers, including 1 official and 5 live price sources.
  • The lowest recorded input price is 86.9% below the official input price; a lower price does not imply better reliability, limits, or refund terms.
Example usage1M input + 0.2M output
Official monthly cost$1.288
Lowest recorded monthly cost$0.206

Task guidance is generated by fixed rules from official positioning and ComputeUnion market data; it is not an independent benchmark or guarantee.

Provider summary and buying decision

Evidence tier first, then normalized price within each tier.

Market low: $0.13Lowest observed: $0.13
Gemma 4 31B provider price and evidence comparison
Provider / operatorPrice evidencePublic operating historyInput / 1M tokensOutput / 1M tokensAPI docsRefund policyPayment / invoice
CerebrasunknownOfficial price sourcePlatform pageObserved: Unknown$0.99$1.49View docsNot providedvisa · mastercardunknown
Deep InfraunknownLive API pricePlatform pageObserved: Unknown$0.13$0.38View docsNot providedvisa · mastercardunknown
Novita AIunknownLive API pricePlatform pageObserved: Unknown$0.14$0.40View docsNot providedvisa · mastercardunknown
OpenRouterunknownLive API pricePlatform pageObserved: Unknown$0.14$0.40View docsNot providedvisa · mastercard · cryptounknown
Together AIunknownLive API pricePlatform pageObserved: Unknown$0.39$0.97View docsNot providedvisa · mastercardunknown

Platform-submitted information or domain verification does not mean the price was independently checked and is not a ComputeUnion guarantee of pricing, availability, refunds, compliance, invoices, or service quality.

Gemma 4 31B price history

Observed input-price changes across providers over the last 30/90 days.

Unit: $/1M tokens

Estimate monthly cost

The estimate uses the displayed price without assuming discounts or cache hits.

Estimated monthly cost$0.51

FAQ

How is Gemma 4 31B API priced?

Input and output tokens are billed separately. The current recorded official baseline is $0.9900 input / $1.4900 output per 1M tokens.

Which provider is cheapest for Gemma 4 31B?

The lowest recorded input price is $0.1300/1M tokens from Deep Infra; verify provider terms and the observation date before production use.

How do official and relay prices differ?

Official and relay channels can differ in price, latency, limits, and service terms; evaluate them separately for production traffic.