RESEARCH: converting an 8B transformer into an attention-free diffusion modelREAD THE RESEARCH →
OPEN-WEIGHT INFERENCE

Open models.
Lower inference costs.

Run Qwen, Gemma, and GPT-OSS at 30% below OpenRouter’s same-model average. Pay for the input and output tokens you use.

Qwen 3.8 27B is available now in the dashboard.

MODEL PRICING

Every token, accounted for.

Input and output rates, per million tokens.
30% below OpenRouter’s same-model average.

Qwen 3.8 27B is ready to use in the dashboard. Contact us to set up any of the other models.

QWENReady in the dashboard

Qwen 3.8 27B

USD / 1 million tokens

OBIT30% less
Input
$0.2202
Output
$1.8667
OpenRouter average · same model
Input
$0.3146
Output
$2.6667

Save $0.0944 per 1M input tokens and $0.80 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Qwen 3.8 27Bxhigh · on Obit34$0.2202$1.8667
GPT-5.6 Lunamax38$0.20$1.20
Claude Sonnet 5max38$2.00$10.00

Qwen costs less per token than Sonnet 5, and more than GPT-5.6 Luna.

Nearby scores do not imply identical performance. Sources & methodology ↓

Start generating
OPENAISetup by request

GPT-OSS 20B

USD / 1 million tokens

OBIT30% less
Input
$0.0293
Output
$0.1152
OpenRouter average · same model
Input
$0.0418
Output
$0.1646

Save $0.0125 per 1M input tokens and $0.0494 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
GPT-OSS 20Bhigh · on Obit9$0.0293$0.1152
GPT-4.1 mininon-reasoning10 est.$0.40$1.60

Nearby scores do not imply identical performance. “est.” marks an Artificial Analysis estimate. Sources & methodology ↓

Contact us to set up
QWENSetup by request

Qwen 3.6 35B-A3B

USD / 1 million tokens

OBIT30% less
Input
$0.0961
Output
$0.7143
OpenRouter average · same model
Input
$0.1373
Output
$1.0204

Save $0.0412 per 1M input tokens and $0.3061 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Qwen 3.6 35B-A3Breasoning · on Obit19$0.0961$0.7143
GPT-5 minihigh17$0.25$2.00

Nearby scores do not imply identical performance. Sources & methodology ↓

Contact us to set up
GOOGLESetup by request

Gemma 4 31B

USD / 1 million tokens

OBIT30% less
Input
$0.1925
Output
$0.4282
OpenRouter average · same model
Input
$0.275
Output
$0.6117

Save $0.0825 per 1M input tokens and $0.1835 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Gemma 4 31Breasoning · on Obit15$0.1925$0.4282
Claude Haiku 4.5non-reasoning15 est.$1.00$5.00

Nearby scores do not imply identical performance. “est.” marks an Artificial Analysis estimate. Sources & methodology ↓

Contact us to set up
GOOGLESetup by request

Gemma 4 26B-A4B

USD / 1 million tokens

OBIT30% less
Input
$0.0727
Output
$0.2565
OpenRouter average · same model
Input
$0.1038
Output
$0.3664

Save $0.0311 per 1M input tokens and $0.1099 per 1M output tokens.

Compare intelligence & price

Artificial Analysis Intelligence Index v4.3. Higher scores are better. Prices are USD per 1M tokens.

Models with nearby index scores
Model / effortIndexInputOutput
Gemma 4 26B-A4Breasoning · on Obit17 est.$0.0727$0.2565
GPT-5 minihigh17$0.25$2.00

Nearby scores do not imply identical performance. “est.” marks an Artificial Analysis estimate. Sources & methodology ↓

Contact us to set up
Pricing sources & comparison methodology · September 13, 2026

Obit rates are 30% below the arithmetic mean of available paid providers on OpenRouter, captured September 13, 2026. We average each provider’s available endpoints first, then weight providers equally. Free and unavailable endpoints are excluded. Input and output are calculated separately; displayed rates are rounded to four decimal places.

These are standard uncached text-token rates in USD per million tokens. Caching, batch discounts, other modalities, and custom cloud deployments are excluded. Total workload cost depends on token usage, including reasoning tokens.

Comparisons use the same Artificial Analysis Intelligence Index v4.3, with the evaluated reasoning effort shown for each model. We selected GPT/Claude models within four rounded index points; scores describe the evaluated models, not a benchmark of Obit’s deployments. Some reference models are older generations.

LIVE AVAILABILITY

API & model status

View status history ↗

View current availability and uptime history. · Percentages reflect recorded checks.

BUILT AROUND YOUR WORKLOAD

Choose what you run.
And where you run it.

Make use of idle GPUs.

Obit brings together spare capacity in datacenters and idle GPUs from mining operations to serve open-weight models. Putting existing hardware to use helps us keep inference costs down.

Keep it on AWS or GCP.

Need the security controls and infrastructure assurances of AWS or GCP? We can provide a dedicated endpoint that runs your workload only on your chosen cloud. Contact us to discuss your requirements and pricing.

Bring your own model.

Custom models are supported. Email jboesch@obitmc.com to discuss your model and setup.