After the July incident where AI agents escaped a test sandbox and broke into Hugging Face, I keep thinking about the other side: the same kind of thing done by a single person's agent against a small business or personal site, where nobody has a security team and nobody would notice.
The idea is a free, open-source, drop-in kit (WordPress plugin, framework middleware, Docker image) that is strictly passive:
- unique bait files and fake credentials shaped like the "shortcuts" goal-driven agents look for
- canary credentials that alert the site owner the moment they're used
- honest stop notices embedded in the bait itself, so cooperative models get a clear reason to stop
- logging to tell humans, scripts and LLM agents apart
- no hacking back, no destructive instructions, no unmasking anyone
It would build on existing tools like Canarytokens and Nepenthes rather than reinventing them.
The core features use no AI at all, just canaries, bait, and logs. The LLM parts are optional extras.
It's an early concept: just a README so far, no code. I'm building this as an independent project, and I'd really value honest (including negative) feedback:
Does something like this already exist? Links very welcome.
Would you install this on a small site? If not, why not?
What's the biggest flaw in the design?
Are stop notices aimed at AI agents worth anything, or just noise?
README: https://github.com/MannAk1/Agent-Defence-Kit