One endpoint for every model, with routing and failover built in.

ai.ml is an LLM gateway. Keep the OpenAI, Anthropic or Gemini SDK you already use, change the base URL, and every request is translated, routed to the best host, streamed back and metered to the token.

  • Five wire formats, one key
  • Streaming, tools and structured output
  • Failover across hosts and clouds
from openai import OpenAI client = OpenAI(-   base_url="https://api.openai.com/v1",+   base_url="https://api.ai.ml/v1",    api_key=KEY,)client.chat.completions.create(    model="openai/gpt-5-mini", ...)
openai/gpt-5-mini
  • HostCountryInput / 1MFirst tokenStatus
  • Azure OpenAI (Global Standard)GLOBAL$0.25n/aTimed out
  • OpenAIUS$0.25n/aServed
  • OpenRouterGLOBAL$0.260.68 sStandby

What the response tells you

Served by
OpenAI, US
Attempts
2
First token
n/a

An example of failover: the first host times out, so the request moves to the next one. You get one answer and pay for one.

Models fromAnthropicGoogleOpenAIDeepSeekMeta LlamaAion-labsAmazonAnthracite-orgArcee-aiBytedanceCognitivecomputationsCohere

It speaks the API your code already speaks

Five wire formats go in, and each one comes back in the same format, whichever provider served it. Pick a format to see where it lives and what it carries.

Drop-in for the OpenAI SDK and everything built on it

client = OpenAI(base_url="https://api.ai.ml/v1", api_key=KEY)
  • Streaming, with usage on the last chunk
  • Function tools and parallel tool calls
  • JSON mode and JSON Schema output
  • Images, audio and files as input

What happens to a request, in order

Six stages sit between your SDK and the model. Click a stage, or play the whole path, to see what each one checks and what it writes down.

4. Route

Hosts that cannot serve the request are filtered out by feature, country and data rules. The rest are ranked by your policy: price, speed or uptime.

# trace
5 hosts, 3 eligible   ranked by price

See every host before you send a request

Each model lists the hosts that serve it, with the price, the measured speed and the 30-day uptime. Sort them, keep traffic in one country, or click a row to pin one host. These are the hosts of openai/gpt-5-mini.

HostCountryInput / output per 1MFirst token30-day uptime
Azure OpenAI (Global Standard)Global$0.25 / $2.00n/an/a
OpenAIUnited States$0.25 / $2.00n/an/a
OpenRouterGlobal$0.26 / $2.100.68 s100.00%

Your request goes to the cheapest host

If that host fails or slows down, the next one in the list takes the request. You are charged once.

# add to the request body
"route": {
  "sort": "price"
}

Live figures from the catalog, measured over the last 30 days.

Browse models

Policies as a file

Write routing rules once and bind them to the whole organization, a project or one key. They are type-checked before they are saved.

Canary a new host

Send a small share of traffic to a new host. The share is cut automatically when its error rate or answer quality slips.

Shadow a second model

Replay live requests to another model in the background and compare cost, speed and stop reason. Your users never see it.

Dry run before you ship

Run the real router over a past request or a described one and read which host it would pick, and why the others lost.

A feature works the same on every host, or you are told

Hosts differ in what they support. The gateway knows each one, fills the gap where it safely can, and refuses with a clear error where it cannot. Hover or tap a cell.

ModelToolsStructured outputVisionReasoningPrompt cache
anthropic/claude-haiku-4.5SupportedSupportedSupportedSupportedSupported
google/gemini-3.1-flash-liteSupportedSupportedSupportedSupportedSupported
openai/gpt-5-miniSupportedSupportedSupportedSupportedSupported
deepseek/deepseek-v4.1-flashSupportedSupportedSupportedSupportedSupported
meta/llama-3.3-70b-instructSupportedSupportedNot availableNot availableNot available
aion-labs/aion-2.0SupportedSupportedNot availableSupportedSupported

From the live catalog. Each model's page lists every host and what it supports.

Every response explains itself

You never have to guess which provider answered, how many attempts it took or what it cost. The answer is on the response, and the same record is in the request log for as long as you keep logs.

Every host that was considered and what happened at each.

HTTP/1.1 200 OK
x-aiml-request-id: 01K7QW3M9T2ZB4
x-aiml-provider: bedrock
x-aiml-region: us
x-aiml-attempts: 2
x-aiml-cost-micro: 2676
x-aiml-route-trace: anthropic:timeout,bedrock:ok

Limits that hold, per key and per project

Each key carries its own rate limits, model allow-list, allowed browser origins, expiry and spend budget. When a budget is spent, the next request is refused with a clear error and a webhook tells your team. Try it: start an agent stuck in a loop.

$0.00spent today
0requests served
0requests refused
Waiting for the first request.

You decide what is kept, and where it runs

Data rules are set per key and can only get stricter further down. Flip the switches to see what a key with those rules does.

Keep request logsStored encrypted for 7 to 365 days, your choice.
Hide personal details in logsEmails, phone numbers and card numbers are masked before storage.
Zero data retentionOnly hosts that keep nothing may serve the request.
No training on your dataSkip any host whose terms allow training on API traffic.
India onlyRequests are served by hosts in India or refused.

With these rules, this key will

  • Keep encrypted request logs with personal details masked.
  • Never send a request to a host that may train on it.
  • Run on the best host in any region.

A signed data processing agreement, the list of sub-processors and erasure on request are part of every account.

Your coding tools, pointed here by one command

The setup command checks your key and model first, changes only the settings it owns, and keeps a backup so one more command puts everything back.

$ npx @aiml/cli setup claude-code
ok  key accepted
ok  model available
ok  backup saved
ok  wrote ~/.claude/settings.json

npx @aiml/cli restore

Also in the box

The parts a team asks for after the first integration are already there.

Response cache

Turn it on for a key and an identical request is answered from cache at no cost, streaming included.

Embeddings and rerank

One route for OpenAI, Cohere, Voyage and Jina style embedding models, with large batches split for you.

Your own provider keys

Bring the keys you already have. Requests go through them and you keep the logs, limits and routing.

File uploads

Upload an image, PDF or audio file once and reference it from any request, on any provider.

SDKs and an agent loop

TypeScript, Python and Go libraries, plus a tool-calling loop that stops at a spend cap you set.

MCP server

Let your editor or agent read models, usage and requests, and create keys, through standard tools.

Evals on your prompts

Score models and hosts on your own test set and let the result steer routing.

Webhooks and single sign-on

Signed events with retries, SAML or OIDC sign-in, and user provisioning from your directory.

Prepaid, pay per token. Top up in rupees by UPI, card or net banking, or in dollars by card. No subscription. Indian companies get a GST invoice on every top-up.

See pricing

Questions people ask first

Anything else, ask in your access request and a person will reply.

How do I get in?

The beta is by invitation. Request access, and when your request is approved you receive a link to set a password and open your account.

Does streaming work the way my SDK expects?

Yes. Events come back in the format your SDK sent, in the order it expects, whichever provider produced them. Tool calls and reasoning stream too.

What does it cost?

You pay per token at the price shown for each host on the model's page. There is no subscription.

Do I have to rewrite my code?

No. Keep the OpenAI, Anthropic or Gemini SDK you use today and change the base URL and the key. Streaming, tools and structured output work the same way.

What happens when a provider goes down?

The request moves to the next host that serves the same model. You are charged for the answer you received and nothing for the attempt that failed.

How much latency does the gateway add?

It adds a small fixed overhead per request and reports it on every response, separate from the provider's own time, so you can see it for yourself.

Where is my data processed?

You choose. A key can be limited to hosts in one region, and requests that cannot be served there are refused.

Join the closed beta

We are letting teams in a few at a time so every account gets real help with its first integration.

How it worksTell us what you are building. When your request is approved, you get a link to set a password and open your account.