Policies as a file
Write routing rules once and bind them to the whole organization, a project or one key. They are type-checked before they are saved.
ai.ml is an LLM gateway. Keep the OpenAI, Anthropic or Gemini SDK you already use, change the base URL, and every request is translated, routed to the best host, streamed back and metered to the token.
from openai import OpenAI client = OpenAI(- base_url="https://api.openai.com/v1",+ base_url="https://api.ai.ml/v1", api_key=KEY,)client.chat.completions.create( model="openai/gpt-5-mini", ...)
An example of failover: the first host times out, so the request moves to the next one. You get one answer and pay for one.
Five wire formats go in, and each one comes back in the same format, whichever provider served it. Pick a format to see where it lives and what it carries.
client = OpenAI(base_url="https://api.ai.ml/v1", api_key=KEY)
Six stages sit between your SDK and the model. Click a stage, or play the whole path, to see what each one checks and what it writes down.
Hosts that cannot serve the request are filtered out by feature, country and data rules. The rest are ranked by your policy: price, speed or uptime.
# trace
5 hosts, 3 eligible ranked by priceEach model lists the hosts that serve it, with the price, the measured speed and the 30-day uptime. Sort them, keep traffic in one country, or click a row to pin one host. These are the hosts of openai/gpt-5-mini.
| Host | Country | Input / output per 1M | First token | 30-day uptime |
|---|---|---|---|---|
| Azure OpenAI (Global Standard) | Global | $0.25 / $2.00 | n/a | n/a |
| OpenAI | United States | $0.25 / $2.00 | n/a | n/a |
| OpenRouter | Global | $0.26 / $2.10 | 0.68 s | 100.00% |
If that host fails or slows down, the next one in the list takes the request. You are charged once.
# add to the request body "route": { "sort": "price" }
Live figures from the catalog, measured over the last 30 days.
Write routing rules once and bind them to the whole organization, a project or one key. They are type-checked before they are saved.
Send a small share of traffic to a new host. The share is cut automatically when its error rate or answer quality slips.
Replay live requests to another model in the background and compare cost, speed and stop reason. Your users never see it.
Run the real router over a past request or a described one and read which host it would pick, and why the others lost.
Hosts differ in what they support. The gateway knows each one, fills the gap where it safely can, and refuses with a clear error where it cannot. Hover or tap a cell.
| Model | Tools | Structured output | Vision | Reasoning | Prompt cache |
|---|---|---|---|---|---|
| anthropic/claude-haiku-4.5 | Supported | Supported | Supported | Supported | Supported |
| google/gemini-3.1-flash-lite | Supported | Supported | Supported | Supported | Supported |
| openai/gpt-5-mini | Supported | Supported | Supported | Supported | Supported |
| deepseek/deepseek-v4.1-flash | Supported | Supported | Supported | Supported | Supported |
| meta/llama-3.3-70b-instruct | Supported | Supported | Not available | Not available | Not available |
| aion-labs/aion-2.0 | Supported | Supported | Not available | Supported | Supported |
From the live catalog. Each model's page lists every host and what it supports.
You never have to guess which provider answered, how many attempts it took or what it cost. The answer is on the response, and the same record is in the request log for as long as you keep logs.
Every host that was considered and what happened at each.
HTTP/1.1 200 OK x-aiml-request-id: 01K7QW3M9T2ZB4 x-aiml-provider: bedrock x-aiml-region: us x-aiml-attempts: 2 x-aiml-cost-micro: 2676 x-aiml-route-trace: anthropic:timeout,bedrock:ok
Each key carries its own rate limits, model allow-list, allowed browser origins, expiry and spend budget. When a budget is spent, the next request is refused with a clear error and a webhook tells your team. Try it: start an agent stuck in a loop.
Waiting for the first request.Data rules are set per key and can only get stricter further down. Flip the switches to see what a key with those rules does.
A signed data processing agreement, the list of sub-processors and erasure on request are part of every account.
The setup command checks your key and model first, changes only the settings it owns, and keeps a backup so one more command puts everything back.
$ npx @aiml/cli setup claude-code ok key accepted ok model available ok backup saved ok wrote ~/.claude/settings.json npx @aiml/cli restore
The parts a team asks for after the first integration are already there.
Turn it on for a key and an identical request is answered from cache at no cost, streaming included.
One route for OpenAI, Cohere, Voyage and Jina style embedding models, with large batches split for you.
Bring the keys you already have. Requests go through them and you keep the logs, limits and routing.
Upload an image, PDF or audio file once and reference it from any request, on any provider.
TypeScript, Python and Go libraries, plus a tool-calling loop that stops at a spend cap you set.
Let your editor or agent read models, usage and requests, and create keys, through standard tools.
Score models and hosts on your own test set and let the result steer routing.
Signed events with retries, SAML or OIDC sign-in, and user provisioning from your directory.
Prepaid, pay per token. Top up in rupees by UPI, card or net banking, or in dollars by card. No subscription. Indian companies get a GST invoice on every top-up.
See pricingAnything else, ask in your access request and a person will reply.
The beta is by invitation. Request access, and when your request is approved you receive a link to set a password and open your account.
Yes. Events come back in the format your SDK sent, in the order it expects, whichever provider produced them. Tool calls and reasoning stream too.
You pay per token at the price shown for each host on the model's page. There is no subscription.
No. Keep the OpenAI, Anthropic or Gemini SDK you use today and change the base URL and the key. Streaming, tools and structured output work the same way.
The request moves to the next host that serves the same model. You are charged for the answer you received and nothing for the attempt that failed.
It adds a small fixed overhead per request and reports it on every response, separate from the provider's own time, so you can see it for yourself.
You choose. A key can be limited to hosts in one region, and requests that cannot be served there are refused.
We are letting teams in a few at a time so every account gets real help with its first integration.