Authenticate and apply policy
The gateway key identifies the application, team and user. The policy attached to it decides which models are allowed, what the budget and rate limit are, and which protections are on.
Your applications talk to one familiar API. The gateway does the rest, in the same order every time.
Any OpenAI or Anthropic SDK, with a gateway key instead of a provider key.
Key, team, model permissions, budget and rate limit are checked first.
PII, secrets and prompt-injection checks run. Content is redacted, tokenised, blocked or just monitored.
Safe prompt clean-up, cache lookup, then the best model for the job, with failover ready.
The gateway holds the provider keys. Your apps never do.
Response DLP runs, tokens are restored, and spend and audit records are written.
The gateway key identifies the application, team and user. The policy attached to it decides which models are allowed, what the budget and rate limit are, and which protections are on.
Detectors look for personal information (including Australian identifiers), secrets and credentials, and prompt-injection patterns. Each finding is handled by the mode you chose: monitor, redact, tokenise or block.
Safe normalisation trims wasted tokens. The cache is checked: exact match first, and similarity matching only if you have enabled it.
The router picks the model within what the key is allowed to use and records a plain-language reason. If the provider fails, retries and failover take over.
Output is checked with the same detectors. Tokens swapped in earlier are restored so your application receives a natural answer.
Usage, cost, savings and policy actions are attributed to the key, team, user, model and provider, written to the audit log and exported over OpenTelemetry if configured.
We set up your organisation and Microsoft Entra sign-in.
Start in monitor mode, review what would be caught, then enforce.
Give each app a gateway key and point your SDK at the gateway.
If a provider is unavailable, approved failover models take over. If a request breaks a blocking rule, your application receives a clear error rather than silence.
Detection is probabilistic. We recommend running in monitor mode against representative data before turning on blocking.
Developer docsRequest access and we will help you set up your first policy, key and budget.