Notes on security, mysteries, and being wrong.
I break things to understand them — safety controls, cold cases, and my own settled opinions. Everything here is provisional: a theory I'm still testing, held only as tightly as the evidence earns.
Recent writing
I built a harness that talks AI safety filters into ignoring their own instructions — no exploit code, just words. Then it taught me why you can't grade a language model with a keyword.
Currently making
Iago — a red-team harness that probes an LLM's guardrails with a library of persuasion attacks, judges what slips through, and writes it up like a pentest. Read the writeup →