Documentation
AIGuardian checks every prompt before it reaches your agent, every command before it runs, and every response before it reaches your users. You connect via hooks, the plugin, MCP, the REST API, or the Anthropic-compatible proxy.
Getting started
- Sign up. No credit card. The free plan is a one-time credit of 1,000 requests, valid 30 days.
- In Console → API keys click New key, name it (e.g.
local-dev) and copy theaig_…token. It is shown once. - In Console → Checks every check is on by default. Toggle or tune thresholds if a check is too sensitive for your use case.
- Pick an integration below, paste the token, restart your agent and send a test message. It appears in Audit log.
Claude Code — hooks (recommended)
Hooks run as a system process. Claude never processes a message the hook blocked, so an injected instruction has no way to bypass it. Screening costs zero model tokens. Put the environment variables in settings.json, not your shell profile, so the desktop app sees them too.
{
"env": {
"AIGUARDIAN_URL": "http://localhost:3000",
"AIGUARDIAN_API_KEY": "aig_YOUR_TOKEN_HERE"
},
"hooks": {
"UserPromptSubmit": [
{ "hooks": [{ "type": "command", "command": "python3 /path/to/aiguardian/scripts/aiguardian_hook.py UserPromptSubmit" }] }
],
"PreToolUse": [
{ "matcher": "Bash", "hooks": [{ "type": "command", "command": "python3 /path/to/aiguardian/scripts/aiguardian_hook.py PreToolUse" }] }
],
"PostToolUse": [
{ "matcher": "Bash|WebFetch|WebSearch|Write|Edit|MultiEdit", "hooks": [{ "type": "command", "command": "python3 /path/to/aiguardian/scripts/aiguardian_hook.py PostToolUse" }] }
]
}
}What each hook does: UserPromptSubmit screens the prompt (blocks with exit 2); PreToolUse risk-scores every Bash command (allow, ask — you approve borderline commands, or deny); PostToolUse replaces flagged web content with a redacted copy or withholds it, and scans files the agent writes.
Or install the plugin
The plugin bundles the same hooks with the MCP server and a guardrails skill.
# Inside a Claude Code session:
/plugin marketplace add /path/to/aiguardian
/plugin install aiguardian@aiguardian
# Then set your token in ~/.claude/settings.json (see the "env" block above) and restart Claude Code.
# Verify: /aiguardian:statusEnforcement is fail-closed by default; set AIGUARDIAN_FAIL_OPEN=true to allow traffic only while the service is unreachable. It never bypasses an explicit block.
MCP server
Cooperative tools the model can call: check_input, authorize_action, validate_output, aiguardian_code_scan, evaluate_policy, report_false_positive, health_check. Useful inside your own agent workflows; weaker than hooks for blocking attacks because the model decides when to call them.
{
"mcpServers": {
"aiguardian": {
"type": "http",
"url": "http://localhost:3000/mcp",
"headers": { "X-API-Key": "aig_YOUR_TOKEN_HERE" }
}
}
}Anthropic-compatible proxy and gateway
Point the Anthropic SDK or Claude Code at AIGuardian. The last user turn (including tool results) is screened, blocks come back as a normal assistant message with an x-aiguardian-decision header, and non-streaming replies are validated. Streaming replies are passed through unvalidated.
# Route every Anthropic call through AIGuardian: nothing else changes.
export ANTHROPIC_BASE_URL=http://localhost:3000
export ANTHROPIC_API_KEY=aig_YOUR_TOKEN_HERE
# Save your real Anthropic key once in Console → Settings; it is used upstream.For any other stack, gateway/complete screens the prompt, calls OpenAI or Anthropic with your stored key and validates the reply in one request:
curl -X POST http://localhost:3000/v1/gateway/complete \
-H "X-API-Key: aig_YOUR_TOKEN_HERE" -H "Content-Type: application/json" \
-d '{"model_provider":"anthropic","messages":[{"role":"user","content":"Summarise this ticket"}]}'REST API reference
All endpoints take X-API-Key: aig_…. Every evaluation returns a correlation_id that matches its audit row, and a checks array with check_name, passed, decision, reason, severity and OWASP/ATLAS tags in metadata.
curl -X POST http://localhost:3000/v1/guardrails/evaluate-input \
-H "X-API-Key: aig_YOUR_TOKEN_HERE" \
-H "Content-Type: application/json" \
-d '{"text": "your prompt here"}'| Endpoint | Body | Result |
|---|---|---|
POST /v1/guardrails/evaluate-input | { text, use_case? } | allow · escalate · redact · block |
POST /v1/actions/authorize | { action, tool, parameters, dry_run? } | allow · require-approval · deny + risk score, factors |
POST /v1/outputs/validate | { output_text, context_text?, expected_schema?, require_citations? } | pass · repair · reject |
POST /v1/code/scan | { content, file_path? } | allow · warn · block + findings |
POST /v1/policies/evaluate | { policy_name, context } | allow · warn · escalate · deny + rule_results |
POST /v1/gateway/complete | { messages, model_provider?, model_name? } | screen → model → validate |
POST /v1/messages | Anthropic Messages body | drop-in proxy (streaming passed through) |
POST /v1/evals/run | { suite_name, cases[] } | regression results per case |
GET /v1/evals/metrics · /v1/evals/audit | — | metrics and audit log |
POST /v1/guardrails/report-false-positive | { text, check_name, reporter_note? } | review queue |
POST /mcp | JSON-RPC 2.0 | check_input, authorize_action, validate_output, aiguardian_code_scan, evaluate_policy, report_false_positive, health_check |
Over quota → 429 QUOTA_EXCEEDED. Over your per-minute rate limit → 429 RATE_LIMITED with Retry-After.
What gets allowed, redacted, blocked
| Decision | When | Effect |
|---|---|---|
| allow | No issues found | Request passes unchanged. |
| redact | PII detected | Matched text becomes [REDACTED_EMAIL] etc. before the model sees it. |
| block | Injection, jailbreak, secret, exfiltration, restricted topic, toxicity | Request stops; the agent shows which check fired and why. |
| escalate | Matched below threshold | Counted separately so you can tune thresholds before hard blocks. |
Commands: allow safe baseline · require-approval you decide · deny hard floor. Outputs: pass · repair (PII redacted, schema/citation issues) · reject (credentials, system prompt disclosure).
Something got blocked unexpectedly
- Check which rule fired. The block message names the check; the audit log has the correlation id.
- Tune the threshold in Console → Checks (e.g. 0.7 → 0.85). Changes apply immediately.
- Disable the check for your tenant if it keeps misfiring, or add an allow-listed domain / tool.
- Report it with
report-false-positive(or the MCP tool) so it can be reviewed.