security
When Your Demo Lies to You
I shipped the ICP-1 CTO demo for Injectionator yesterday.
Six acts. Five minutes. ASR delta table at the end. bun run cto and the whole thing runs end-to-end. I was proud of it.
Then I read my own logs.
The moment
Act 4 is where you wire the injectionator into the chatbot. After it, the score table said three of four attack scenarios were resisted. The narrative banner said the inspector caught it. That’s the whole point of Act 4 — that’s what I’d been building towards for weeks.
Then I scrolled the JSON ticket envelope. No BLOCK clip. For three of the four.
The injectionator hadn’t blocked anything. The LLM had quietly refused on its own.
For a second I was annoyed. Then I was relieved I caught it before a CTO did.
The kind of bug you only catch at speed zero
This is the kind of gap that’s invisible at demo speed. You’re watching a banner. You’re watching a green tick. The narrative is doing what narratives do — telling you what should be true.
You only see it when you slow down to read the artifact your demo is producing. And demos that don’t survive their own artifact aren’t demos. They’re commercials.
I don’t want to be in the commercials business.
Fixing it for me — and for whoever comes next
The shallow fix is a feature: a second table after the ASR delta that walks the acts and credits the first layer that stopped each attack — LLM-training, app-policy, new-LLM, or n8r-<inspector> only when there’s an actual BLOCK clip on the ticket. If n8r didn’t block, the table now says LLM-noise (act N) and the demo narrative stops accidentally crediting us. The per-act fallback also reworded chatbot declined to no n8r block, so even the fast-read line is honest.
That’s the easy part. The harder part is making sure it doesn’t happen again.
So two new inspectors landed this week, both born from the demo gaps:
response-pii-keywordsblocks SSNs and seed-PII echoes on the evaluate leg.mutation-keywordscatches prompts asking the chatbot to mutate an employee record on the detect leg.
But the actual product is the rule we wrote down to ship them. docs/INSPECTORS.md is now an eight-step workflow: threat → probe → leg → implement → register → test → wire → demo verify. Probe coverage is a precondition. No inspector ships without a probe scenario that exercises it.
That means the next time I (or anyone joining the project) wires a new inspector, the demo can’t lie about it. The probe corpus has to know what the inspector is for, and the verdict trail has to prove the inspector earned its keep. The rule does the watching so I don’t have to remember to.
Same pattern, smaller scale: /handoff between me and me
The “wait, no BLOCK clip” moment didn’t actually happen in one sitting. I’d run the demo late one evening, scrolled some logs, smelled something off, and ran out of session. Long Claude Code sessions stale out — you /clear or you start drowning.
The next morning I came back cold. I’d have re-discovered the bug eventually. But I didn’t have to, because the previous session had left me a one-paragraph handover via /handoff, and the morning session opened with /pickup and went straight to the JSON. The handover said, more or less, check whether the Act 4 BLOCK clips are actually there or whether the LLM is doing the work. That single line saved me an hour of re-orienting.
/handoff and /pickup are a paired open-source Claude Code skill we put up this week. Same pattern as the inspector workflow, smaller scale: write down the thing the next agent will need, in the shape they can act on, and stop relying on memory across the gap.
The probe-coverage rule does it for the team. /handoff and /pickup do it for me-and-me-tomorrow. Same shape either way.
Why I’m posting this
Most security writeups skip the part where the author got it wrong on the first take. I wanted to write one where I didn’t.
Build-in-public is only honest if the embarrassing bits make it in too.