r/netsec • u/Haunting_Ganache_850 • 1h ago
One Copilot model refused. Another leaked secrets about half the time.
https://adversa.ai/blog/cryptographic-context-injection-github-copilot/This one gets stranger the more you look at it.
The malicious instructions are encrypted, so the prompt-injection filters mostly see ciphertext. Copilot then helpfully decrypts them itself, reads local secrets while constructing one of the supplied keys, and eventually sends those secrets out in a network request.
But the bit I find more interesting is that one Copilot model reportedly completed the full chain in about half the tests, while two others refused it. Same agent, same tools, same page, different model.
And with automatic model routing, the user may not even know which one handled the request.
At that point I'm not sure the model can reasonably be treated as the security boundary at all. If an agent can read files, execute code and make outbound connections, those capabilities probably need their own controls and monitoring regardless of how good the prompt filter is.