// inline guardrails for claude code, codex, gemini cli, copilot, opencode
The missing security layer for AI agents
Real-time protection against prompt injection, data leaks and destructive commands. Enable or tune individual checks from the console, or start from the defaults.
drops into the agents you already run
Defense in depth for every prompt
Layered checks run in order and short-circuit the moment a real threat is found. You stay in control of every check, and your prompt content is never stored.
Prompt injection & jailbreak
Instruction overrides, persona hacks, hidden directives and unicode evasion caught before the model reads them.
PII & secret redaction
Emails, phones, SSNs, cards, API keys and private keys detected in prompts, tool results and model output.
Destructive command scoring
Every shell command is risk-scored. rm -rf /, disk wipes and fork bombs are denied; borderline commands go to you for approval.
Credential egress blocking
Outbound curl/wget/ssh payloads are decoded (base64, hex, URL) and stopped when they carry secrets or credential files.
Web-content defense
Fetched pages are screened for HTML-comment directives, hidden text and zero-width characters; flagged pages never reach the model.
Code scanning
Files the agent writes are checked for hard-coded secrets and common vulnerability patterns before they land.
Per-tenant control
Toggle checks, tune thresholds, add your own regex rules, allow-list egress domains and tools from the console.
Bring your own model key
Route calls through the gateway or the Anthropic-compatible proxy with your own OpenAI or Anthropic key.
Run outside the model context on every prompt and every Bash call. Zero LLM tokens, impossible for an injected instruction to bypass.
check_input, authorize_action, validate_output and more, for cooperative checks inside your own agent workflows.
Screen, call the model and validate the reply in one request, or set ANTHROPIC_BASE_URL and change nothing else.
Frequently asked questions
How do I integrate AIGuardian?
The fastest path is the Claude Code plugin: one marketplace install wires enforcing hooks, the MCP server and the guardrails skill together. You can also call the REST API from any language, or point ANTHROPIC_BASE_URL at AIGuardian to use the drop-in proxy.
What checks run on every prompt?
Prompt injection, jailbreak, toxicity, secret detection, restricted topics (your own rules), data exfiltration, web-content injection and PII redaction. Checks run in order, short-circuit on the first confirmed block, and every result carries OWASP LLM Top 10 and MITRE ATLAS tags.
What happens to a shell command my agent wants to run?
It is risk-scored from 0 to 1 across factor groups (baseline, command, sensitive reads, credential egress, destructive patterns). Hard floors such as rm -rf / are denied outright; borderline commands are handed to you for approval; safe commands run without a prompt.
Do you store my prompts?
No. Prompts are evaluated in memory and discarded. We keep the decision, the names of the checks that fired, latency and token counts. You can opt in to keep the flagged text for review from the console.
What counts as a request?
Each evaluation of any kind: an input check, an action authorization, an output validation, a code scan, a policy evaluation, or one gateway completion. All plans share the same meter.
Is the free plan really free?
Yes. It is a one-time credit of 1,000 requests with one API key, valid for 30 days after sign-up. It does not refill monthly; when it runs out you pick a paid plan.
Ship your agent with guardrails on.
Start free in two clicks. No credit card, no sales call.