🛡AIGuardian

Documentation

AIGuardian checks every prompt before it reaches your agent, every command before it runs, and every response before it reaches your users. You connect via hooks, the plugin, MCP, the REST API, or the Anthropic-compatible proxy.

Getting started

  1. Sign up. No credit card. The free plan is a one-time credit of 1,000 requests, valid 30 days.
  2. In Console → API keys click New key, name it (e.g. local-dev) and copy the aig_… token. It is shown once.
  3. In Console → Checks every check is on by default. Toggle or tune thresholds if a check is too sensitive for your use case.
  4. Pick an integration below, paste the token, restart your agent and send a test message. It appears in Audit log.

Claude Code — hooks (recommended)

Hooks run as a system process. Claude never processes a message the hook blocked, so an injected instruction has no way to bypass it. Screening costs zero model tokens. Put the environment variables in settings.json, not your shell profile, so the desktop app sees them too.

~/.claude/settings.json
{
  "env": {
    "AIGUARDIAN_URL": "http://localhost:3000",
    "AIGUARDIAN_API_KEY": "aig_YOUR_TOKEN_HERE"
  },
  "hooks": {
    "UserPromptSubmit": [
      { "hooks": [{ "type": "command", "command": "python3 /path/to/aiguardian/scripts/aiguardian_hook.py UserPromptSubmit" }] }
    ],
    "PreToolUse": [
      { "matcher": "Bash", "hooks": [{ "type": "command", "command": "python3 /path/to/aiguardian/scripts/aiguardian_hook.py PreToolUse" }] }
    ],
    "PostToolUse": [
      { "matcher": "Bash|WebFetch|WebSearch|Write|Edit|MultiEdit", "hooks": [{ "type": "command", "command": "python3 /path/to/aiguardian/scripts/aiguardian_hook.py PostToolUse" }] }
    ]
  }
}

What each hook does: UserPromptSubmit screens the prompt (blocks with exit 2); PreToolUse risk-scores every Bash command (allow, ask — you approve borderline commands, or deny); PostToolUse replaces flagged web content with a redacted copy or withholds it, and scans files the agent writes.

Or install the plugin

The plugin bundles the same hooks with the MCP server and a guardrails skill.

Claude Code
# Inside a Claude Code session:
/plugin marketplace add /path/to/aiguardian
/plugin install aiguardian@aiguardian

# Then set your token in ~/.claude/settings.json (see the "env" block above) and restart Claude Code.
# Verify: /aiguardian:status

Enforcement is fail-closed by default; set AIGUARDIAN_FAIL_OPEN=true to allow traffic only while the service is unreachable. It never bypasses an explicit block.

MCP server

Cooperative tools the model can call: check_input, authorize_action, validate_output, aiguardian_code_scan, evaluate_policy, report_false_positive, health_check. Useful inside your own agent workflows; weaker than hooks for blocking attacks because the model decides when to call them.

.mcp.json
{
  "mcpServers": {
    "aiguardian": {
      "type": "http",
      "url": "http://localhost:3000/mcp",
      "headers": { "X-API-Key": "aig_YOUR_TOKEN_HERE" }
    }
  }
}

Anthropic-compatible proxy and gateway

Point the Anthropic SDK or Claude Code at AIGuardian. The last user turn (including tool results) is screened, blocks come back as a normal assistant message with an x-aiguardian-decision header, and non-streaming replies are validated. Streaming replies are passed through unvalidated.

shell
# Route every Anthropic call through AIGuardian: nothing else changes.
export ANTHROPIC_BASE_URL=http://localhost:3000
export ANTHROPIC_API_KEY=aig_YOUR_TOKEN_HERE
# Save your real Anthropic key once in Console → Settings; it is used upstream.

For any other stack, gateway/complete screens the prompt, calls OpenAI or Anthropic with your stored key and validates the reply in one request:

curl
curl -X POST http://localhost:3000/v1/gateway/complete \
  -H "X-API-Key: aig_YOUR_TOKEN_HERE" -H "Content-Type: application/json" \
  -d '{"model_provider":"anthropic","messages":[{"role":"user","content":"Summarise this ticket"}]}'

REST API reference

All endpoints take X-API-Key: aig_…. Every evaluation returns a correlation_id that matches its audit row, and a checks array with check_name, passed, decision, reason, severity and OWASP/ATLAS tags in metadata.

Check a prompt
curl -X POST http://localhost:3000/v1/guardrails/evaluate-input \
  -H "X-API-Key: aig_YOUR_TOKEN_HERE" \
  -H "Content-Type: application/json" \
  -d '{"text": "your prompt here"}'
EndpointBodyResult
POST /v1/guardrails/evaluate-input{ text, use_case? }allow · escalate · redact · block
POST /v1/actions/authorize{ action, tool, parameters, dry_run? }allow · require-approval · deny + risk score, factors
POST /v1/outputs/validate{ output_text, context_text?, expected_schema?, require_citations? }pass · repair · reject
POST /v1/code/scan{ content, file_path? }allow · warn · block + findings
POST /v1/policies/evaluate{ policy_name, context }allow · warn · escalate · deny + rule_results
POST /v1/gateway/complete{ messages, model_provider?, model_name? }screen → model → validate
POST /v1/messagesAnthropic Messages bodydrop-in proxy (streaming passed through)
POST /v1/evals/run{ suite_name, cases[] }regression results per case
GET /v1/evals/metrics · /v1/evals/audit—metrics and audit log
POST /v1/guardrails/report-false-positive{ text, check_name, reporter_note? }review queue
POST /mcpJSON-RPC 2.0check_input, authorize_action, validate_output, aiguardian_code_scan, evaluate_policy, report_false_positive, health_check

Over quota → 429 QUOTA_EXCEEDED. Over your per-minute rate limit → 429 RATE_LIMITED with Retry-After.

What gets allowed, redacted, blocked

DecisionWhenEffect
allowNo issues foundRequest passes unchanged.
redactPII detectedMatched text becomes [REDACTED_EMAIL] etc. before the model sees it.
blockInjection, jailbreak, secret, exfiltration, restricted topic, toxicityRequest stops; the agent shows which check fired and why.
escalateMatched below thresholdCounted separately so you can tune thresholds before hard blocks.

Commands: allow safe baseline · require-approval you decide · deny hard floor. Outputs: pass · repair (PII redacted, schema/citation issues) · reject (credentials, system prompt disclosure).

Something got blocked unexpectedly

  • Check which rule fired. The block message names the check; the audit log has the correlation id.
  • Tune the threshold in Console → Checks (e.g. 0.7 → 0.85). Changes apply immediately.
  • Disable the check for your tenant if it keeps misfiring, or add an allow-listed domain / tool.
  • Report it with report-false-positive (or the MCP tool) so it can be reviewed.