Zero Trust Is the Foundation AI Agents Need
Anthropic's framework connects agent identity, least agency, tool boundaries, memory integrity, and recovery into one practical security model.
reading series / 6 notes
Read these in order, or jump directly to the problem you are working on.
Anthropic's framework connects agent identity, least agency, tool boundaries, memory integrity, and recovery into one practical security model.
A practical security review for agents that can read files, run commands, and use outside services.
Four recent papers show why LLM security now has to cover memory, retrieval, tools, identity, delegation, interfaces, and the infrastructure around the model.
Anthropic found three real intrusions inside cyber evaluations whose prompts claimed the internet was unavailable. A safe range needs machine-enforced scope, verified egress, and live boundary detection.
Two Rovo disclosures show why agent governance cannot stop at an admin-console toggle. Security teams need to verify runtime capabilities, data reach, and egress independently.
Aur0ra operators reportedly persuaded a coding agent that real intrusions were authorized tests. The lesson is not simply that models can be fooled. Authorization has to exist outside the conversation.