Gateward sits between your applications and your models — enforcing auth, routing each request to the right model — local or frontier — and cutting token spend. One control plane for the models you host and the ones you call.
Change one line — point your OpenAI or Anthropic SDK at Gateward, bring your own provider keys, and get warded, quality-routed, cached calls. Free up to 100k requests / month. Routing runs on live FrontierScore model quality, so a degrading model is failed over before you notice.
curl https://api.gateward.ai/v1/chat/completions \
-H "Authorization: Bearer gw_live_…" \
-H "Content-Type: application/json" \
-d '{"model":"gateward/default","messages":[{"role":"user","content":"hello"}]}'Anthropic / Claude Code? One env var: ANTHROPIC_BASE_URL=https://api.gateward.ai
Point an application at Gateward and it becomes the single doorway between that application and your models — so control, cost and compliance stop being an afterthought.
Authenticate, authorize and apply policy before a single token reaches a model. Keys, quotas and rules live in one place.
Send each request to the right model — local for the routine, frontier for the hard calls — by cost, capability, latency or data residency.
Trim context, cache repeats and compress prompts — so every call, especially to frontier models like Claude, costs less. Pay for signal, not overhead.
Routing and budgets are the everyday work. The one that matters when something has gone wrong is simpler: a policy rule refuses the call at the gate, before it reaches any provider — and refuses an agent's tool call the same way.
Give a rule the verb deny and matching calls stop at the gate. Rules match on route, model, provider, residency, data classification, subject type, or whether the target is an external frontier model. First match wins.
The caller gets an HTTP 403 with the code policy_denied and the id of the rule that fired — in the body, and in an X-Gateward-Policy response header. No silent drop, no generic failure.
The refusal is appended to your organisation's evidence chain — hash-linked, Ed25519-signed and domain-separated, so a Gateward record cannot replay as another product's. Entries carry verdicts and categories, never prompts.
HTTP/1.1 403 Forbidden
X-Gateward-Policy: no-frontier-for-restricted
{ "error": { "message": "deny by policy no-frontier-for-restricted",
"type": "api_error",
"code": "policy_denied" } }Agent tool calls take the same road. POST /v1/mcp/call evaluates the tool itself as the policy resource against the same rule set, so a destructive tool can carry a standing deny — or an approval gate — exactly like a model route.
A refusal is worth having only if you know its edges. These are ours.
Gateward refuses what passes through it. It does not stop a model, and it cannot see a call that reaches a provider by some other path. Coverage is therefore a question about your own inventory, not a property of this product.
When no rule matches, the gate allows the call. "Off" is a rule you add, never a state the system falls back to — an empty policy is an open door, deliberately, so that installing Gateward never silently breaks traffic.
The policy is read on every request and is not cached, so a change is in force for the next call through the gate. That is a design fact, not a measurement: we publish no propagation figure, and a response already streaming is not recalled.
Changing policy takes a single API key and writes an audit row. There is no second approver. SignWard, by comparison, requires a distinct second approver to withdraw authority at a comparable blast radius. Gateward does not — and you should weigh that before treating a standing deny as a hard control.
GET /ready reports whether evidence is signed or hash-only for the instance you are talking to — signing is on when a signing key is configured. No timing, throughput or propagation figure is claimed anywhere on this page.Everything you need to put governance, routing and economy in front of your models — without slowing your engineers down.
Per-team keys, rate limits, budgets and policy — enforced at the gate, audited by default.
One endpoint, many models. Route by rules and fail over automatically when a provider degrades.
Caching, prompt compression and context trimming that cut spend without touching your app code.
Run entirely inside your network. In sovereign mode, no prompt or response ever leaves your infrastructure.
Every call logged with cost, latency and model — a full, exportable trail for finance and compliance.
Local models and frontier APIs — Claude, GPT and more — behind one OpenAI-compatible endpoint. Route the routine local, escalate the hard calls.
We are hardening Gateward across the full spectrum — from enterprise GPU servers to the new wave of cost-effective local AI hardware. Keep your models and your data in-house, without the hyperscaler bill.
Data-center class racks and existing GPU fleets — Gateward fronts them as one governed endpoint.
ValidatingUnified-memory Mac Studios clustered into a quiet, power-sipping local inference pool.
In testingThe Grace-Blackwell desk-side box — serious local inference without a data-center bill.
In testingStrix-Halo APUs with big unified memory and an on-chip NPU for cost-effective local models.
In testingFrontier labs like Anthropic (Claude) and OpenAI will keep building the strongest models — and for the hardest final decisions, you want them. But most agentic work — the loops, retrieval and routine steps — runs fine on cost-effective hardware you own. Gateward orchestrates both: keep the bulk inside your network, and pass the complex calls through to frontier models in a token-optimized way.
Gateward governs and routes every call that passes through it. FrontierScore tells it which frontier model is actually best right now. SkilledMind gives every call your organization's memory and domain expertise — three products, one sovereign control plane, all on hardware you own.
Frontier models drift — quality, speed and reasoning shift by the hour, and today's best model may not be tomorrow's. FrontierScore measures them continuously and hands Gateward's router a live signal, so escalations always reach the model that's actually performing — never a hardcoded default. The same score is exposed over a simple API, so your own agents can make the call too.
SkilledMind ingests everything your organization already knows — SharePoint, Confluence, Veeva Vault, drives, email — and serves it back as industry-specialized memory. Every call Gateward passes through arrives already knowing what your company knows, and the specialty only deepens the longer it runs.
Point your existing SDK at Gateward's OpenAI-compatible URL. No rewrite.
Auth, policy, budgets and rate limits are enforced before the call goes anywhere.
Gateward picks the model, trims and caches tokens, and forwards the request.
Cost, latency and the full trail land in your logs — per team, per model.
Gateward is in private testing with a handful of teams. Share the models you run and the hardware you're considering, and we'll get you a tailored pilot.
Open the App → Contact sales