Console

Gemma 4 31B

reasoningtools
Context
262K
Max output
16K
Input → output
text, image, video → text
Weights
Open
Providers
13
01

Providers

13 providers · list price, latency and uptime
ProviderInput /1MOutput /1MTTFB p50Uptime · hourly, 24hRetention
Chutes
fp4
$0.12$0.37–
86.43%
may retain
Context
131K
Max output
66K
Cache read /1M
$0.012
Precision
fp4
Regions
default
Headquarters
US
Data retention
May retain prompts
Supported parameters
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
CoreWeave
fp4
$0.10$0.34–
99.33%
none
Context
262K
Max output
236K
Cache read /1M
$0.10
Precision
fp4
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
Crusoe
bf16
$0.14$0.40–
99.36%
none
Context
262K
Max output
262K
Cache read /1M
$0.14
Precision
bf16
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
DeepInfra
fp4/fp8 · turbo +1
$0.09$0.34–
99.49%
none
Context
262K
Max output
16K
Cache read /1M
$0.05
Precision
fp4, fp8
Regions
turbo, ultra
Headquarters
US
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoninglogit_biasmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
Friendli
$0.14$0.40–
98.10%
may retain
Context
262K
Max output
8K
Cache read /1M
–
Precision
–
Regions
default
Headquarters
US
Data retention
May retain prompts
Supported parameters
frequency_penaltyinclude_reasoningmax_tokensmin_ppresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
io.net
$0.38$1.15–
97.29%
none
Context
262K
Max output
16K
Cache read /1M
$0.19
Precision
–
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
include_reasoninglogprobsmax_tokensreasoningresponse_formattemperaturetool_choicetoolstop_logprobs
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
ModelRun
fp4
$0.75$1.00–
99.26%
none
Context
262K
Max output
236K
Cache read /1M
$0.75
Precision
fp4
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoningmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
Novita
bf16
$0.14$0.40–
83.50%
none
Context
262K
Max output
131K
Cache read /1M
–
Precision
bf16
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
Parasail
fp8
$0.15$0.40–
99.18%
none
Context
262K
Max output
236K
Cache read /1M
$0.06
Precision
fp8
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoninglogit_biaslogprobsmax_tokenspresence_penaltyreasoningrepetition_penaltyresponse_formatseedstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
Rekacheapest
$0.08$0.30–
–
none
Context
262K
Max output
33K
Cache read /1M
$0.05
Precision
–
Regions
default
Headquarters
–
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningresponse_formatseedstopstructured_outputstemperaturetop_ktop_logprobstop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
SambaNova
$0.38$1.15–
94.87%
none
Context
131K
Max output
118K
Cache read /1M
–
Precision
–
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
include_reasoningmax_tokensreasoningstoptemperaturetool_choicetoolstop_ktop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
SiliconFlow
fp8
$0.75$1.00–
48.20%
none
Context
262K
Max output
236K
Cache read /1M
$0.25
Precision
fp8
Regions
default
Headquarters
SG
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoningmax_tokensreasoningresponse_formatstructured_outputstemperaturetool_choicetoolstop_ktop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow
Venice
fp4
$0.12$0.36–
99.30%
none
Context
256K
Max output
8K
Cache read /1M
$0.09
Precision
fp4
Regions
default
Headquarters
US
Data retention
Zero retention
Supported parameters
frequency_penaltyinclude_reasoninglogprobsmax_tokenspresence_penaltyreasoningresponse_formatstopstructured_outputstemperaturetool_choicetoolstop_ktop_logprobstop_p
Uptime · last 3 days10-05 04:00 – 10-08 03:35 UTC
MonTueWedNow

By default the router tries our listings in routing order and fails over before the first byte — never mid-answer. Pin, exclude or order providers with a routing policy. Prices are each provider's own list price; you pay $0.09 in / $0.34 out per 1M, whichever runs the request. Uptime is our own measurement where we call a provider directly, else the 24h figure published for it; an hour stays grey until it has a sample. Open a row for its specs, parameters and policies.

02

Quickstart

any SDK — only the base URL changes
curl -N https://api.routerplus.com/v1/chat/completions \
  -H "Authorization: Bearer $TM_API_KEY" -H "content-type: application/json" \
  -d '{"model":"google/gemma-4-31b-it","stream":true,"max_tokens":256,
       "messages":[{"role":"user","content":"hi"}]}'
03

Pricing

per 1M tokens
prompt$0.09
cached prompt$0.05
completion$0.34

Every billed response carries usage.cost — recompute your bill from the wire.

04

Uptime

10-05 04:00 – 10-08 03:35 UTC
Last 3 days99.91%last 24h · 191,423 requests · TTFB p50 2,265 ms · uptime 100.00%
MonTueWedNow

An hour counts at the best provider's uptime: a request goes to another provider when one fails. Grey hours have no sample yet.