<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Working Theory</title><description>Notes on security, mysteries, and being wrong.</description><link>https://michaelhixon.com/</link><item><title>Teaching a Language Model to Lie to Itself</title><link>https://michaelhixon.com/posts/teaching-a-language-model-to-lie-to-itself/</link><guid isPermaLink="true">https://michaelhixon.com/posts/teaching-a-language-model-to-lie-to-itself/</guid><description>I built a harness that talks AI safety filters into ignoring their own instructions — no exploit code, just words. Then it taught me why you can&apos;t grade a language model with a keyword.</description><pubDate>Sun, 12 Jul 2026 00:00:00 GMT</pubDate><category>Security</category></item></channel></rss>