A hash chain proves the ordering. Four attacks prove it's not enough.

Sam Labbé explores the critical distinction between ordering and authenticity in AI agent systems. While hash chains and DAGs effectively prove the sequence of events, they do not guarantee the integrity of the agents performing the actions. The article details a catalog of four common 'gate attacks'—self-review, rewritten exams, farmed oracles, and forged verdicts—that exploit the gap between sequential logging and genuine verification. By replaying these attacks against the NoireBox reconciliation layer, the author demonstrates how cryptographic controls and strict adherence to Unix-like primitives can mitigate these vulnerabilities. The piece emphasizes that audit logs should serve as evidence of duty rather than shields from accountability. Through iterative testing and schema updates, the project successfully closed security seams, proving that robust, verifiable trust layers are essential for AI agents operating on critical infrastructure or payment rails.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
Security researchers have analyzed a new version of the DarkSword exploit chain, dubbed 'P7', which targets users on iOS 18.4–18.7. Previously availab…
A recent TechRadar report clarifies the limitations of using a VPN while interacting with AI services like ChatGPT and Claude. While a VPN effectively…
Bitwarden, the popular open-source password management solution, has recently addressed community discussions regarding its dual licensing model. The…


