01 · Prompt injection through trusted content
A poisoned PR title rides into agent context through a legitimate API read. What the injection can and cannot cause — and which downstream actions CIRVIX refuses regardless.
Every piece follows the same shape: problem, reproduction, limitations, the CIRVIX control, and a test you can run.
A poisoned PR title rides into agent context through a legitimate API read. What the injection can and cannot cause — and which downstream actions CIRVIX refuses regardless.
An MCP server exposes more than its description admits. How gateway framing, drift detection, and default-deny interact before the upstream server sees anything.
Credential-shaped arguments, egress after secret contact, and brokered handles that let an agent use a secret without seeing it.
A sandbox constrains where code runs; a policy constrains what a call may do. Why enforcement needs one of each, and what happens with only one.
Destructive shell, workspace escape writes, production deploys without approval. Reproductions, limits, and the rules that hold the line.
Every piece ships with fixtures in the public repo. Run npm run verify:adversarial — 11,629 attack cases, zero bypasses — and add your own.
These pieces describe what the engine enforces today. Where CIRVIX does not detect something — prompt injection itself among them — we say so plainly and show the boundary that still holds: the malicious instruction may reach the agent, but the prohibited action does not run.
Install Free Attack Lab MCP Demo Flagship Demo Security Research Benchmarks Book a walkthrough
Discuss custom VPC deployments, policy requirements, or dedicated enterprise support.
Prefer to talk? Book a 30-minute call