DeepSeek R1

deepseek/deepseek-r1

text in · text out · 163,840 context · 32,768 max output

toolsschemareasoningZDRno-train
Effective cost per 1k requests

cost = p_in·(input − cached) + p_cache_read·cached + p_out·output + p_req; cached tokens are the prefix a sticky endpoint already holds (part2b §9.3), which is why the cheapest list price is not always the cheapest request.

ProviderRegion · quantContext / max outInput / MOutput / MCache read / writeMarkupPer 1k reqUptime 30 dTTFT p50 / p95ThroughputLast 24 hFidelityDataCapabilitiesStatus
Fireworks AIus · fp8163,840 / 32,768$3.300$8.800— / —+10.0%$11.000———no traffic—ZDR no-train model defaultactive
OpenRouterglobal · unknown64,000 / 16,000$0.735$2.625— / —+5.0%$2.782———no traffic— no-train 30 d11 differactive

Quick start

curl https://api.ai.ml/v1/chat/completions \
  -H "Authorization: Bearer $AIML_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-r1","messages":[{"role":"user","content":"Say hello"}]}'
Get a key in the console

Changelog

  • new_model endpoint ep_openrouter_deepseek-r1_global (deepseek/deepseek-r1) added as active; price input 0.735 output 2.625 USD/M 10/5/2026