The Wire / 2026-08-10
The week one vendor made its permission classifier the default and another tried to build a whole agentic internet in three days, while a second frontier lab published what happened when its models were handed to third-party evaluators. The theme, if there is one, is infrastructure being declared before anyone has agreed what it is for.
-
Auto mode is now the default in Claude Code for Pro, Max, and Team plans
We adjudicated this one against the primary record for this week's issue, so it is here as checked rather than relayed. The change is real and the argument for it is measured: a published study found testers caught a clearly dangerous command about one time in seven and got worse the longer they worked, while the classifier caught it about nine times in ten and stayed flat. The part worth carrying into your own thinking is that a boundary you state in conversation is not stored as a rule. The documentation recommends a deny rule if you want a hard guarantee.
-
The Agent Access Model
One vendor's proposal for how sites should decide what an agent may do, published in the middle of a week where the same vendor shipped an agent-first browser, a programmable wallet and a way to give any website an MCP interface. Read it as a position paper rather than a standard. It is the most complete public attempt so far to answer a question the harness vendors have been routing around, which is what the other end of the connection is supposed to do about any of this.
-
Give any website a WebMCP interface
The mechanical half of the same argument. If it lands, the surface an agent talks to stops being a page it scrapes and becomes an interface a site declares. Worth watching for the reason every protocol on this beat is worth watching: the security model arrives after the adoption, and the adoption is the easy part.
-
Third-party cyber evaluations involving OpenAI models
Read it directly after the Anthropic disclosure in our 2026-08-03 wire, where a model reached real systems from inside an evaluation environment. Two labs, a week apart, both describing what happens when a capable model meets a harness built to be permissive on purpose. The evaluation environment is turning out to be a category of production system that nobody staffed as one.
-
Zawinski's Law of MultiAgents
Every agent product expands until it can orchestrate other agents. Funny, and close enough to true that our own watchlist grew a meta-harness category this quarter and then a second entrant a fortnight later. If the joke keeps holding, the interesting question stops being which harness you run and becomes which one is holding the others.