LLM control plane · on-prem ready

The warded gateway for your models.

Gateward sits between your applications and your models — enforcing auth, routing each request to the right model — local or frontier — and cutting token spend. One control plane for the models you host and the ones you call.

Built by Digital One · Dublin, Ireland
Free · self-serve

One endpoint. Every model. Free.

Change one line — point your OpenAI or Anthropic SDK at Gateward, bring your own provider keys, and get warded, quality-routed, cached calls. Free up to 100k requests / month. Routing runs on live FrontierScore model quality, so a degrading model is failed over before you notice.

curl https://api.gateward.ai/v1/chat/completions \
  -H "Authorization: Bearer gw_live_…" \
  -H "Content-Type: application/json" \
  -d '{"model":"gateward/default","messages":[{"role":"user","content":"hello"}]}'

Anthropic / Claude Code? One env var: ANTHROPIC_BASE_URL=https://api.gateward.ai

The gate

Every prompt passes through Gateward first.

Point an application at Gateward and it becomes the single doorway between that application and your models — so control, cost and compliance stop being an afterthought.

Ward

Authenticate, authorize and apply policy before a single token reaches a model. Keys, quotas and rules live in one place.

Route

Send each request to the right model — local for the routine, frontier for the hard calls — by cost, capability, latency or data residency.

Save

Trim context, cache repeats and compress prompts — so every call, especially to frontier models like Claude, costs less. Pay for signal, not overhead.

The refusal

The gate can say no.

Routing and budgets are the everyday work. The one that matters when something has gone wrong is simpler: a policy rule refuses the call at the gate, before it reaches any provider — and refuses an agent's tool call the same way.

A standing deny

Give a rule the verb deny and matching calls stop at the gate. Rules match on route, model, provider, residency, data classification, subject type, or whether the target is an external frontier model. First match wins.

A 403 that names the rule

The caller gets an HTTP 403 with the code policy_denied and the id of the rule that fired — in the body, and in an X-Gateward-Policy response header. No silent drop, no generic failure.

Recorded, not just returned

The refusal is appended to your organisation's evidence chain — hash-linked, Ed25519-signed and domain-separated, so a Gateward record cannot replay as another product's. Entries carry verdicts and categories, never prompts.

HTTP/1.1 403 Forbidden
X-Gateward-Policy: no-frontier-for-restricted

{ "error": { "message": "deny by policy no-frontier-for-restricted",
             "type": "api_error",
             "code": "policy_denied" } }

Agent tool calls take the same road. POST /v1/mcp/call evaluates the tool itself as the policy resource against the same rule set, so a destructive tool can carry a standing deny — or an approval gate — exactly like a model route.

What that does not mean

Four limits, stated plainly.

A refusal is worth having only if you know its edges. These are ours.

It stops calls through the gate

Gateward refuses what passes through it. It does not stop a model, and it cannot see a call that reaches a provider by some other path. Coverage is therefore a question about your own inventory, not a property of this product.

The default is allow

When no rule matches, the gate allows the call. "Off" is a rule you add, never a state the system falls back to — an empty policy is an open door, deliberately, so that installing Gateward never silently breaks traffic.

It binds on the next request

The policy is read on every request and is not cached, so a change is in force for the next call through the gate. That is a design fact, not a measurement: we publish no propagation figure, and a response already streaming is not recalled.

A policy write is one API key

Changing policy takes a single API key and writes an audit row. There is no second approver. SignWard, by comparison, requires a distinct second approver to withdraw authority at a comparable blast radius. Gateward does not — and you should weigh that before treating a standing deny as a hard control.

Checkable rather than asserted: the deny path, the evidence append, the policy lint and the agent tool-call gate each have a named test in the governance suite, and GET /ready reports whether evidence is signed or hash-only for the instance you are talking to — signing is on when a signing key is configured. No timing, throughput or propagation figure is claimed anywhere on this page.
Capabilities

A control plane, not just a proxy.

Everything you need to put governance, routing and economy in front of your models — without slowing your engineers down.

Access control & keys

Per-team keys, rate limits, budgets and policy — enforced at the gate, audited by default.

Model routing & fallback

One endpoint, many models. Route by rules and fail over automatically when a provider degrades.

Token & cost optimization

Caching, prompt compression and context trimming that cut spend without touching your app code.

On-prem & sovereign

Run entirely inside your network. In sovereign mode, no prompt or response ever leaves your infrastructure.

Observability & audit

Every call logged with cost, latency and model — a full, exportable trail for finance and compliance.

Multi-provider gateway

Local models and frontier APIs — Claude, GPT and more — behind one OpenAI-compatible endpoint. Route the routine local, escalate the hard calls.

Runs where your data lives

From data-center racks to a desk-side cluster.

We are hardening Gateward across the full spectrum — from enterprise GPU servers to the new wave of cost-effective local AI hardware. Keep your models and your data in-house, without the hyperscaler bill.

Enterprise GPU servers

Data-center class racks and existing GPU fleets — Gateward fronts them as one governed endpoint.

Validating

Apple Mac Studio clusters

Unified-memory Mac Studios clustered into a quiet, power-sipping local inference pool.

In testing

NVIDIA DGX Spark

The Grace-Blackwell desk-side box — serious local inference without a data-center bill.

In testing

AMD Ryzen AI Max

Strix-Halo APUs with big unified memory and an on-chip NPU for cost-effective local models.

In testing
Relative cost to serve, by deployment
Indexed to managed cloud API = 100 · lower is cheaper · illustrative, validation in progress
Cloud API (managed)
0
Enterprise GPU server
0
Mac Studio cluster
0
NVIDIA DGX Spark
0
AMD Ryzen AI Max
0
Figures are directional estimates we are validating across the hardware above — not a published benchmark. The pattern we keep seeing: once traffic is steady, owning the silicon undercuts per-token cloud pricing, and on-prem keeps data in your walls.
Local-first, frontier-ready

Run your agents locally. Escalate to the best models when it counts.

Frontier labs like Anthropic (Claude) and OpenAI will keep building the strongest models — and for the hardest final decisions, you want them. But most agentic work — the loops, retrieval and routine steps — runs fine on cost-effective hardware you own. Gateward orchestrates both: keep the bulk inside your network, and pass the complex calls through to frontier models in a token-optimized way.

Local ecosystem
Agents, loops & retrieval
On your Mac Studio, DGX Spark or Ryzen AI hardware
optimize · route
Gateward
Auth · token economy · escalation
hard calls only
Frontier models
Claude · GPT · …
The complex, final decisions — token-optimized
The Digital One AI stack

Gateward is the gate. Two more pieces make it a complete stack.

Gateward governs and routes every call that passes through it. FrontierScore tells it which frontier model is actually best right now. SkilledMind gives every call your organization's memory and domain expertise — three products, one sovereign control plane, all on hardware you own.

SkilledMind
Memory & skill layer
What your models know
memory
Gateward
Govern · route · optimize
live scores
FrontierScore
Live model health
Which model is best now
↓  one endpoint for your apps & agents — nothing leaves your network
Live model health

Route to the model that's best right now.

Frontier models drift — quality, speed and reasoning shift by the hour, and today's best model may not be tomorrow's. FrontierScore measures them continuously and hands Gateward's router a live signal, so escalations always reach the model that's actually performing — never a hardcoded default. The same score is exposed over a simple API, so your own agents can make the call too.

Gateward routes · FrontierScore tells it where.
Visit frontierscore.ai →
Memory & skill layer

Give every model your institutional memory.

SkilledMind ingests everything your organization already knows — SharePoint, Confluence, Veeva Vault, drives, email — and serves it back as industry-specialized memory. Every call Gateward passes through arrives already knowing what your company knows, and the specialty only deepens the longer it runs.

Gateward governs the call · SkilledMind gives it memory.
Visit skilledmind.ai →
How it works

Four steps, one endpoint.

01

Connect

Point your existing SDK at Gateward's OpenAI-compatible URL. No rewrite.

02

Ward

Auth, policy, budgets and rate limits are enforced before the call goes anywhere.

03

Route & optimize

Gateward picks the model, trims and caches tokens, and forwards the request.

04

Observe

Cost, latency and the full trail land in your logs — per team, per model.

Why it matters

Bring your models in-house. Keep your speed.

0
Endpoint for every model & provider
0%
On-prem option — data stays in your network
~0%
Target token-cost reduction in pilots
0
Hardware targets in active testing
Private beta

Tell us your stack. We'll set up a pilot.

Gateward is in private testing with a handful of teams. Share the models you run and the hardware you're considering, and we'll get you a tailored pilot.

Open the App → Contact sales