Llama 3.3 70B Instruct

meta/llama-3.3-70b-instruct

text in · text out · 131,072 context · 16,384 max output · uptime 100.00% · TTFT 1419 ms

toolsschemaZDRno-train
Effective cost per 1k requests

cost = p_in·(input − cached) + p_cache_read·cached + p_out·output + p_req; cached tokens are the prefix a sticky endpoint already holds (part2b §9.3), which is why the cheapest list price is not always the cheapest request.

ProviderRegion · quantContext / max outInput / MOutput / MCache read / writeMarkupPer 1k reqUptime 30 dTTFT p50 / p95ThroughputLast 24 hFidelityDataCapabilitiesStatus
AWS Bedrockus · none131,072 / 16,384$0.720$0.720— / —+10.0%$1.800———no traffic— no-train 30 dmodel defaultactive
DeepInfraus · fp8131,072 / 16,384$0.253$0.440— / —+10.0%$0.726———no traffic—ZDR no-train 7 differactive
OpenRouterglobal · unknown131,072 / 16,384$0.105$0.336— / —+5.0%$0.378100.00%1419 / 1419 ms4.6 tok/s100.00% · 1419 ms · 4.6 tok/s— no-train 30 d12 differactive

Quick start

curl https://api.ai.ml/v1/chat/completions \
  -H "Authorization: Bearer $AIML_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"meta/llama-3.3-70b-instruct","messages":[{"role":"user","content":"Say hello"}]}'
Get a key in the console

Changelog

  • new_model endpoint ep_openrouter_llama-3.3-70b-instruct_global (meta-llama/llama-3.3-70b-instruct) added as active; price input 0.105 output 0.336 USD/M 10/5/2026