DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

text, image in · text out · 1,000,000 context · 384,000 max output · uptime 100.00% · TTFT 1105 ms

toolsschemavisionreasoningcacheno-train
Effective cost per 1k requests

cost = p_in·(input − cached) + p_cache_read·cached + p_out·output + p_req; cached tokens are the prefix a sticky endpoint already holds (part2b §9.3), which is why the cheapest list price is not always the cheapest request.

ProviderRegion · quantContext / max outInput / MOutput / MCache read / writeMarkupPer 1k reqUptime 30 dTTFT p50 / p95ThroughputLast 24 hFidelityDataCapabilitiesStatus
DeepSeekcn · none1,000,000 / 384,000$0.300$1.200$0.006 / —pass-through$0.906———no traffic— 365 dmodel defaultactive
OpenRouterglobal · unknown1,048,576 / 943,718$0.315$1.260$0.006 / —+5.0%$0.951100.00%1105 / 1105 ms14.4 tok/s100.00% · 1105 ms · 14.4 tok/s— no-train 30 d16 differactive

Quick start

curl https://api.ai.ml/v1/chat/completions \
  -H "Authorization: Bearer $AIML_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"deepseek/deepseek-v4.1-flash","messages":[{"role":"user","content":"Say hello"}]}'
Get a key in the console

Changelog

  • new_model endpoint ep_openrouter_deepseek-v4.1-flash_global (deepseek/deepseek-v4.1-flash) added as active; price input 0.315 output 1.260 USD/M 10/5/2026