The decision layer for AI infrastructure

One endpoint in front of 4,500+ models. Alphe picks the cheapest model that still clears your quality bar, and cuts up to 70% of your inference bill.

Drop-in OpenAI-compatible endpoint · No rewrites · Live in under five minutes

Routes across
OpenAI Anthropic Google DeepMind Meta Mistral Cohere DeepSeek xAI Groq Together Fireworks Perplexity AWS Bedrock Azure OpenAI
Before the call

Every way to answer this. One of them gets sent.

Which model, which tool, whether one of your own agents already does the job: a space of answers, not a single choice. Alphe scores the whole space first and calls the cheapest route that still clears your bar. The other routes are what you would have paid for.

Request

Read the last 40 support tickets, group them by theme, and open a Linear issue for each theme.

Scored against 4,500+ models, 7,000+ tools, and every agent and workflow you have registered
  • Claude Opus 5 (max) + linear·create_issue $0.94 Clears the bar. Seventy-eight times the price of one that also clears it.
  • GPT-5.6 Sol (max) + linear·create_issue $0.51 Clears the bar, and still buying reasoning this job never asks for.
  • Gemini 3.6 Flash + linear·create_issue $0.012 Cheapest route that clears the bar on clustering and on tool calls. Sent
  • Step 3.7 Flash + linear·create_issue $0.004 Cheaper, and misses the bar. It splits one theme into four.
  • Your ticket-triage agent $0.03 Already registered, already tags tickets. It cannot open issues.

9 ms to score them. Nothing above has been called yet, so nothing above has been billed.

Watch one call find its model.

A real request, scored and priced in the open. Every number below is one the router had to produce before it was allowed to send anything.

alphe_classify

tag 12,000 tickets by intent 5 signals
  • TASK_CLASS scored
    Bulk classification, fixed label set
  • REASONING_DEPTH scored
    Shallow: single pass, no chain
  • CONTEXT_LENGTH scored
    1.2k tokens median, 4.1k p99

Plan Score, price, send cheapest pass

Holds Strict JSON · eu-west-1 only

Your application

Tag 12,000 support tickets by intent and escalate anything angry.

Alphe classifying

Scored as bulk classification. Shallow reasoning, 1.2k median context, strict JSON out.

Alphe pricing candidates

Four models clear the quality bar. Haiku 4.5 is the cheapest of them at this context length.

12,000 tagged, 341 escalated. $11.42 spent against $412 on the frontier model you were using.

Ask anything…
claude-haiku-4.5

alphe_candidates

QUALITY BAR: 0.92 · 39 PRICED

  • claude-haiku-4.5 $0.80/M picked
  • gpt-4o-mini $0.60/M 0.89, under bar

alphe_dispatch

ROUTE: none
POST /v1/chat/completions
tokens
1,284
cost
$0.0009

200 OK · 940 ms

route_policy

QUALITY BAR
0.92
FALLBACK
gpt-4o
REGION
eu-west-1

Watch it on your traffic

Route my first call
How it works

Five things happen between your call and the answer.

client.ts
const client = new OpenAI({
  baseURL: "https://api.alpheai.com/v1",
  apiKey: process.env.ALPHE_KEY,
});
01

Ingest

Change the base URL. Nothing else. Alphe speaks the OpenAI, Anthropic and raw HTTP shapes, so existing SDKs, agent frameworks and eval harnesses keep working.

  • One line of config, no SDK swap, no rewrite
  • OpenAI, Anthropic and raw HTTP on the same endpoint
  • Under 8 ms of gateway overhead on the path
alphe_classify 5 signals · 4 ms
  • TASK_CLASS bulk label
  • REASONING shallow
  • CONTEXT 1.2k median
  • TOOLS none
  • OUTPUT strict JSON

Scored before anything expensive sees the request

02

Classify

A small resident classifier reads the request before anything expensive sees it: task type, reasoning depth, context length, tool surface, output contract.

  • Five signals scored on-path, about 4 ms added
  • Resident, so classifying costs no extra provider call
  • That score is the input to every decision after it
route_table bar 0.92
  • classmodel$/M
  • bulk classifyclaude-haiku-4.50.80
  • long contextgemini-1.5-pro1.25
  • tool callsgpt-4o2.50
  • hard reasoningclaude-opus-415.00
  • embeddingstext-embedding-30.02

312 candidates repriced as providers move

03

Route

Every candidate model is priced against the live token cost, its measured quality on that task class, and its current latency. The cheapest one that clears your bar wins.

  • 4,500+ candidates, repriced as providers move
  • The quality bar is yours, set per endpoint
  • Decision made in under 2 ms
verify · ticket_tagger rubric v4
  • Valid JSON against the schema
  • Label drawn from the allowed set
  • Sentiment matches the ticket
claude-haiku-4.5 claude-opus-4

re-run passed · table updated

04

Verify

Outputs are scored against your rubric, not assumed correct. A response that misses is escalated to a stronger model automatically and the routing table learns from it.

  • Rubrics per endpoint, not one global score
  • Automatic escalation on a miss, same request id
  • The table learns, so that class does not miss twice
trace_4f9c21ab cache HIT
  • teamsupport-eng
  • featureticket-tagger
  • customeracme-corp
  • regioneu-west-1

1,284 tokens · $0.0009 · 940 ms

05

Return

The answer comes back with the semantic cache warmed, the trace written, and cost attributed to a team, a feature and a customer. You can answer "what did this feature cost last week" without instrumenting anything.

  • Semantic cache absorbs 30–50% of repeat traffic
  • Every call traced, with nothing to instrument
  • Cost attributed by team, feature and customer
Reach

One call, and all of it is already in range.

Models are the part everyone talks about. Alphe routes across agents, tools and workflows on the same call, weighs what each candidate costs and how fast it answers, then sends the request to whichever one clears your bar for the least money.

Alphe
  • Models
  • Agents
  • Tools
  • Workflows
  • Cost
  • Latency
  • Quality
  • Context
  • Fallbacks
  • Region
  • Chat
  • Reasoning
  • Search
  • Extraction
  • Embeddings
  • Speech
  • Code
  • Vision

None of it is a separate integration to wire up. One endpoint, one key, and the decision made per request.

Real work

Watch Alphe actually do the work.

One prompt. Several tools, more than one model, a receipt at the end. Pick a job and watch which route it takes.

Jobs Four replays
Your application ROUTE: none

Read the 74 vendor contracts in Drive and flag every auto-renewal clause.

drive·fetch_files 74 documents · 1.9M tokens 2.1 s

Scored as bulk extraction over long context. gemini-2.5-flash clears the bar at a fortieth of the frontier price, so 71 of the 74 go there.

alphe·escalate 3 documents · ambiguous indemnity claude-sonnet-4.5
notion·append_page Contract risk register 12 clauses

Receipt $1.14 · 3 m 12 s · $47.60 if every page had gone to one frontier model

Reply…
gemini-2.5-flash +1
Multi-workspace

Every workload, every key, every budget, at once.

One request rarely means one kind of work. Alphe splits it: the images and video one way, the documents another, the code a third. Each part is charged to the workspace and provider account that owns it, before it picks a model.

Swap models instantly

Stay on the best model.

Your tools, agents and workflows are connected to Alphe, not to a model. When something new wins at code review or long-context extraction, switch. The whole stack comes with it.

  • Keys, scopes and budgets stay on Alphe, so there is nothing to re-wire per model.
  • Every frontier and open model, your own agents, and whatever ships next month.
  • Pilot a new model on 5% of traffic, then roll it out. Same tools, no migration.
One hard question

Two points cost 65× more.

One question, five answers, one grader. Six weeks of a running flood, asked from the first breach to the day of the run, and every clause of it wants a different source. Alphe routes it for $0.0074 and grades 6.5 for accuracy and 7.5 for how it reads. The sharpest answer is two points better on accuracy at 65× the bill; the best-written one is a point and a half better at 95×. The fastest answer lands in 18 seconds and scores 3 out of 10 for being worth reading.

Accuracy points returned per dollar spent on this question. Alphe returns 878 of them; the best single model returns 130. Column height is logarithmic: the spread is 77×.

  1. 878 Alphe Alphe
  2. 130 Grok 4.20 SpaceXAI
  3. 100 Gemini 3.1 Pro Preview Google
  4. 18 GPT-5.6 Sol OpenAI
  5. 11 Claude Opus 5 Anthropic

Alphe's own measurement — one question, one grader, five ways of answering it. Accuracy and quality are 0–10 grades; cost is the whole bill for the run, provider tokens included; wall clock is time to the finished answer.

Comparison

Everyone else owns one column.

A gateway gives you every model behind one key but no opinion about which one to call. A framework gives you agents and workflows but leaves the model, the tool and the bill to you. Going direct gives you neither. Alphe is the decision layer over all four, and the price of the answer is part of the decision.

Capability Best Alphe One decision layer Model gateways OpenRouter Together Agent frameworks LangChain CrewAI Direct to provider OpenAI Anthropic
Route across models Pick the model per call, not per project.
Route across agents Hand the job to the agent that handles it.
Route across tools One catalogue, called on your behalf.
Route across workflows Multi-step runs picked the same way.
Picks the tool for the job Selection, not a list you maintain.
Price is part of the decision Cheapest route that still clears the bar.
Fails over mid-request A provider going down is not your outage.
One key for every provider No per-vendor accounts to keep alive.
Drops into code you already wrote Same request shape, one base URL changed.

Scroll the table sideways for the rest →

  • Yes
  • Partly, or by hand
  • No

Marks describe the category, not any one product in it. The named examples are there to say which category is meant.

Economics

Put your own number in.

The savings are not a discount. They are the difference between what you paid and what the request was worth.

Three compounding levers: routing each call to the cheapest model that clears your bar, serving semantic cache hits instead of re-billing near-identical prompts, and compressing context that never influenced the answer.

  • Routing: 75–85% reduction on the traffic that gets routed down
  • Semantic cache: 30–50% on repetitive workloads
  • Context compression: 20–40% fewer input tokens at equal output quality
Current monthly inference spend
$24K/mo
With Alphe
$7K/mo
Saved per year
$205K

70% reduction: routing, semantic cache hits and prompt compression, measured against your current provider mix.

0%
Token savings
0%
Cost reduction
0
Models routed
< 2s
Median response
Coverage

Four and a half thousand of them, behind one contract.

Models are benchmarked and added to the routing table as they ship, behind the same key, the same bill and the same rate limit. You deploy nothing.

OpenAI Anthropic Google Gemini Meta Llama Mistral AI Cohere DeepSeek xAI Grok Alibaba Qwen Microsoft Phi
Amazon Bedrock Google Gemma AI21 Labs Databricks Perplexity NVIDIA IBM Granite Stability AI Moonshot Kimi Zhipu GLM
MiniMax 01.AI Yi Baichuan Hugging Face Voyage AI Groq Together AI Fireworks AI Cerebras Ollama

+ 4,500 models on the same routing table

Tools

Each request gets the three it needs.

Handing a model the whole catalogue is how it picks the wrong thing. Alphe ranks 7,000+ integrations against the request, passes on the shortlist, and runs the calls in order, the same selection it does for models, one layer down.

Refund order 4471 and tell the customer why.

Gmail Slack GitHub Stripe Notion HubSpot Zoom Twilio Linear Google Drive Salesforce Sentry Figma Datadog Jira Google Sheets Discord Letta PostgreSQL Zendesk Asana Shopify ClickUp Telegram Snowflake Trello Intercom Google Calendar Glean Docker Okta Airtable MailChimp MongoDB WhatsApp GitLab Supabase PayPal Miro Box Confluence Databricks Calendly Elasticsearch Dropbox QuickBooks Postman Vercel Zapier + 7,000 more
Trust

A gateway sees everything. Here is what ours does with it.

Prompts and completions are held only as long as the request is in flight. Nothing is written to durable storage unless you switch on tracing for a specific endpoint, and retention is set per endpoint rather than per account.
Encrypted in transit and at rest. Access to production is broken-glass only, logged, and expires automatically. The full report is available under NDA.
Detection and redaction run before a request leaves our edge, so a provider never receives the raw value. Detected entities are replaced with stable tokens and rehydrated on the way back, which keeps the model output usable.
Pin a tenant to a region and Alphe will only consider models served from it, even when a cheaper candidate exists elsewhere. Residency is a routing constraint, not a policy document.
Retrieved documents and tool output are treated as untrusted input. Instructions found inside them are flagged and stripped of authority before the model sees them, and the decision is written to the trace.
Bring your own provider keys and Alphe routes through them, so your existing commitments, discounts and rate limits still apply. Or use ours and get a single invoice. Both can run side by side per tenant.

Stop paying frontier prices for classification work.

Early access is open. Point one endpoint at Alphe, keep your code, and read the difference off your own invoice.