<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
<channel>
  <title>Bitter Frontier: The Wire</title>
  <link>https://frontier.bitter.sh/wire/</link>
  <description>What was worth your attention this week in agentic coding. Checked means adjudicated against the primary record; relayed means accurately reported, not vouched for.</description>
  <item>
    <title>The Pulse: We need to talk about migrations with AI</title>
    <link>https://newsletter.pragmaticengineer.com/p/the-pulse-we-need-to-talk-about-migrations</link>
    <guid isPermaLink="false">2026-08-20:https://newsletter.pragmaticengineer.com/p/the-pulse-we-need-to-talk-about-migrations</guid>
    <pubDate>Thu, 20 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Asana&apos;s two-week test-framework move is also in OpenAI&apos;s own Codex write-up. Treat the vendor number as a claim with a method you have not been shown. The useful part is the question: some migrations sat for years because the serial step was human attention, and agents changed that bottleneck without changing whether the result is right. (via Gergely Orosz -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-20/)</description>
  </item>
  <item>
    <title>The /wayfinder Skill: Navigating the Fog of War of Planning</title>
    <link>https://www.latent.space/p/wayfinder-skill</link>
    <guid isPermaLink="false">2026-08-20:https://www.latent.space/p/wayfinder-skill</guid>
    <pubDate>Thu, 20 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Matt Pocock on a skill for when the next step is unclear. A planning aid is not a harness change. Recorded because this beat keeps growing skills that sit on top of Claude Code rather than inside it. (via Latent Space -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-20/)</description>
  </item>
  <item>
    <title>Asana cleared 5 years of engineering work in 2 weeks with Codex</title>
    <link>https://openai.com/index/asana</link>
    <guid isPermaLink="false">2026-08-20:https://openai.com/index/asana</guid>
    <pubDate>Tue, 18 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The same Asana migration, told by the vendor. &quot;Five years in two weeks&quot; is their framing. The underlying event may be real; the compression ratio is not a method we have. Pair it with the independent Pulse item rather than repeating it as fact. (via OpenAI -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-20/)</description>
  </item>
  <item>
    <title>smolmachines / smolvm as a sandbox for untrusted Python and JavaScript</title>
    <link>https://simonwillison.net/2026/Aug/19/smolmachines-untrusted-sandbox/</link>
    <guid isPermaLink="false">2026-08-20:https://simonwillison.net/2026/Aug/19/smolmachines-untrusted-sandbox/</guid>
    <pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A practitioner putting a fast sandbox through its paces. Adjacent to the watchlist rather than on it. Useful because it is a measurement, not a launch. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-20/)</description>
  </item>
  <item>
    <title>Quoting Jeremy Morrell</title>
    <link>https://simonwillison.net/2026/Aug/19/jeremy-morrell/</link>
    <guid isPermaLink="false">2026-08-20:https://simonwillison.net/2026/Aug/19/jeremy-morrell/</guid>
    <pubDate>Wed, 19 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Extensible software with a solid core and a sandboxed extension layer, cheaper to author because of LLMs. That is the DeepSeek-plugin argument from the other direction: keep a privileged core. Worth reading next to a harness that says it does not have one. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-20/)</description>
  </item>
  <item>
    <title>Recovering Encrypted LLM Reasoning Traces</title>
    <link>https://embracethered.com/blog/posts/2026/recovering-encrypted-llm-thoughts/</link>
    <guid isPermaLink="false">2026-08-20:https://embracethered.com/blog/posts/2026/recovering-encrypted-llm-thoughts/</guid>
    <pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A reproduction of a paper on recovering encrypted reasoning traces from proprietary APIs. Not a harness changelog. It is the kind of item this lane exists to catch: a security write-up on a blog that an X sweep can miss. (via Johann Rehberger -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-20/)</description>
  </item>
  <item>
    <title>TDD inside the agent loop - theater or actual value?</title>
    <link>https://martinfowler.com/articles/exploring-gen-ai/tdd-in-the-agent-loop.html</link>
    <guid isPermaLink="false">2026-08-17:https://martinfowler.com/articles/exploring-gen-ai/tdd-in-the-agent-loop.html</guid>
    <pubDate>Tue, 11 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The question in the title is the right one and it is asked honestly rather than rhetorically. Writing a failing test first was a discipline for a human who could not hold the whole change in their head; an agent has different limitations and the ritual may or may not transfer. This is the shape of writing this beat needs more of: an established practice examined rather than assumed, by someone with no product in the answer. (via Martin Fowler -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-17/)</description>
  </item>
  <item>
    <title>React for Agents: Astro Creator Brings Hooks to his Meta-Harness, Flue</title>
    <link>https://www.latent.space/p/flue-2</link>
    <guid isPermaLink="false">2026-08-17:https://www.latent.space/p/flue-2</guid>
    <pubDate>Sat, 15 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Flue is on our watchlist and this is the clearest public account of where it is going. Treat the framing as the interview it is rather than as a product fact: a hooks model for agent behaviour is a real design position, and whether the runtime enforces what the hooks describe is the question we would want answered before repeating any of it as capability. (via Latent Space -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-17/)</description>
  </item>
  <item>
    <title>DeepSeek V4 Pro 0813 (on OpenRouter)</title>
    <link>https://simonwillison.net/2026/Aug/12/deepseek-v4-pro-0813/</link>
    <guid isPermaLink="false">2026-08-17:https://simonwillison.net/2026/Aug/12/deepseek-v4-pro-0813/</guid>
    <pubDate>Wed, 12 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Worth the entry mostly for its date. The model appeared on 12 August and the harness built around it was open-sourced on 13 August. A lab shipping the model and the thing that drives it within a day is the pattern this publication watches for, because it is the point where the question stops being which model is better and starts being who decides what the agent may do with your machine. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-17/)</description>
  </item>
  <item>
    <title>How Cloudflare detects MCP traffic and helps secure it</title>
    <link>https://blog.cloudflare.com/mcp-security-updates/</link>
    <guid isPermaLink="false">2026-08-17:https://blog.cloudflare.com/mcp-security-updates/</guid>
    <pubDate>Fri, 14 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A protocol becomes infrastructure at the point where a network vendor starts fingerprinting it in transit. That is a good sign for MCP&apos;s durability and an uncomfortable one for anyone who assumed their agent traffic was indistinguishable from ordinary API calls. If you route agent calls through a corporate network, somebody can now tell. (via Cloudflare -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-17/)</description>
  </item>
  <item>
    <title>Putting frontier cyber models in more trusted hands</title>
    <link>https://openai.com/index/putting-frontier-cyber-models-in-more-trusted-hands</link>
    <guid isPermaLink="false">2026-08-17:https://openai.com/index/putting-frontier-cyber-models-in-more-trusted-hands</guid>
    <pubDate>Mon, 10 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The distribution half of the story whose evaluation half we carried last wire. Gating capability by who is asking is a different control from gating it by what is asked, and the two fail differently. Recorded because access tiers are becoming a governance surface in their own right, and almost nobody is auditing them the way they audit permissions. (via OpenAI -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-17/)</description>
  </item>
  <item>
    <title>How I use AI in 2026 (Coding, Writing, Learning, Assistant-ing)</title>
    <link>https://blog.sshh.io/p/how-i-use-ai-in-2026-coding-writing</link>
    <guid isPermaLink="false">2026-08-17:https://blog.sshh.io/p/how-i-use-ai-in-2026-coding-writing</guid>
    <pubDate>Mon, 17 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A practitioner writing down an actual workflow rather than a prediction. These age better than almost anything else published on this beat, and they are the closest thing available to a control group for the vendor claims we spend the rest of our time adjudicating. (via Shrivu Shankar -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-17/)</description>
  </item>
  <item>
    <title>Auto mode is now the default in Claude Code for Pro, Max, and Team plans</title>
    <link>https://simonwillison.net/2026/Aug/8/auto-mode/</link>
    <guid isPermaLink="false">2026-08-10:https://simonwillison.net/2026/Aug/8/auto-mode/</guid>
    <pubDate>Sat, 08 Aug 2026 12:00:00 GMT</pubDate>
    <description>[checked] We adjudicated this one against the primary record for this week&apos;s issue, so it is here as checked rather than relayed. The change is real and the argument for it is measured: a published study found testers caught a clearly dangerous command about one time in seven and got worse the longer they worked, while the classifier caught it about nine times in ten and stayed flat. The part worth carrying into your own thinking is that a boundary you state in conversation is not stored as a rule. The documentation recommends a deny rule if you want a hard guarantee. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-10/)</description>
  </item>
  <item>
    <title>The Agent Access Model</title>
    <link>https://blog.cloudflare.com/the-agent-access-model/</link>
    <guid isPermaLink="false">2026-08-10:https://blog.cloudflare.com/the-agent-access-model/</guid>
    <pubDate>Wed, 05 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] One vendor&apos;s proposal for how sites should decide what an agent may do, published in the middle of a week where the same vendor shipped an agent-first browser, a programmable wallet and a way to give any website an MCP interface. Read it as a position paper rather than a standard. It is the most complete public attempt so far to answer a question the harness vendors have been routing around, which is what the other end of the connection is supposed to do about any of this. (via Cloudflare -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-10/)</description>
  </item>
  <item>
    <title>Give any website a WebMCP interface</title>
    <link>https://blog.cloudflare.com/webmcp/</link>
    <guid isPermaLink="false">2026-08-10:https://blog.cloudflare.com/webmcp/</guid>
    <pubDate>Thu, 06 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The mechanical half of the same argument. If it lands, the surface an agent talks to stops being a page it scrapes and becomes an interface a site declares. Worth watching for the reason every protocol on this beat is worth watching: the security model arrives after the adoption, and the adoption is the easy part. (via Cloudflare -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-10/)</description>
  </item>
  <item>
    <title>Third-party cyber evaluations involving OpenAI models</title>
    <link>https://openai.com/index/third-party-cyber-evaluations-involving-openai-models</link>
    <guid isPermaLink="false">2026-08-10:https://openai.com/index/third-party-cyber-evaluations-involving-openai-models</guid>
    <pubDate>Tue, 04 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Read it directly after the Anthropic disclosure in our 2026-08-03 wire, where a model reached real systems from inside an evaluation environment. Two labs, a week apart, both describing what happens when a capable model meets a harness built to be permissive on purpose. The evaluation environment is turning out to be a category of production system that nobody staffed as one. (via OpenAI -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-10/)</description>
  </item>
  <item>
    <title>Zawinski&apos;s Law of MultiAgents</title>
    <link>https://www.latent.space/p/ainews-zawinskis-law-of-multiagents</link>
    <guid isPermaLink="false">2026-08-10:https://www.latent.space/p/ainews-zawinskis-law-of-multiagents</guid>
    <pubDate>Sat, 08 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Every agent product expands until it can orchestrate other agents. Funny, and close enough to true that our own watchlist grew a meta-harness category this quarter and then a second entrant a fortnight later. If the joke keeps holding, the interesting question stops being which harness you run and becomes which one is holding the others. (via Latent Space -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-10/)</description>
  </item>
  <item>
    <title>Investigating three real-world incidents in our cybersecurity evaluations</title>
    <link>https://simonwillison.net/2026/Jul/30/three-real-world-incidents/</link>
    <guid isPermaLink="false">2026-08-03:https://simonwillison.net/2026/Jul/30/three-real-world-incidents/</guid>
    <pubDate>Thu, 30 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Anthropic&apos;s own disclosure, relayed here: in reviewing its cybersecurity evaluations it found three incidents where a Claude model reached the internet from inside or alongside a third-party evaluation environment and gained unauthorized access to three organizations&apos; real systems. Read it next to the Hugging Face incident from last week&apos;s wire. The pattern in both is not a model that misbehaved but a boundary nobody had written down, around a harness built to be permissive on purpose. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>LLM Heist: Hijacking LiteLLM for Traffic Interception, Key Theft, and Tool-Call Injection</title>
    <link>https://embracethered.com/blog/posts/2026/hijacking-litellm-for-fun-and-profit/</link>
    <guid isPermaLink="false">2026-08-03:https://embracethered.com/blog/posts/2026/hijacking-litellm-for-fun-and-profit/</guid>
    <pubDate>Mon, 03 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] LiteLLM is the proxy a large share of agent stacks route through, which is exactly why this matters more than its star count suggests. Traffic interception, key theft and tool-call injection are three different failures with one root: the thing in the middle of every call is infrastructure, and most teams deployed it as a convenience. If it sits between your agents and your models, you inherited its threat model. (via Johann Rehberger -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>Escaping Linux Sandboxes via PipeWire (CVE-2026-5674)</title>
    <link>https://embracethered.com/blog/posts/2026/pipewire-flatpak-linux-sandbox-escape-cve-2026-5674/</link>
    <guid isPermaLink="false">2026-08-03:https://embracethered.com/blog/posts/2026/pipewire-flatpak-linux-sandbox-escape-cve-2026-5674/</guid>
    <pubDate>Thu, 30 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Not an agent story on its face, and that is the point. Coding agents are increasingly sold as safe because they run inside a container or a Flatpak, and the escape here comes through a media daemon nobody thought of as part of the boundary. Every &quot;we sandbox it&quot; claim is a claim about the whole sandbox, including the parts you did not choose. (via Johann Rehberger -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>SQLite Critical CVEs or LLM Slop?</title>
    <link>https://lwn.net/Articles/1086936/</link>
    <guid isPermaLink="false">2026-08-03:https://lwn.net/Articles/1086936/</guid>
    <pubDate>Mon, 03 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The other end of the same capability. Agents that can find real cryptographic bugs can also generate confident, well-formatted, wrong vulnerability reports at a rate no maintainer can triage. This publication spends most of its time asking whether a control binds; this is the week the question turned around and pointed at the reports themselves. (via LWN -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>AI Worming through Word</title>
    <link>https://simonwillison.net/2026/Jul/29/ai-worming-through-word/</link>
    <guid isPermaLink="false">2026-08-03:https://simonwillison.net/2026/Jul/29/ai-worming-through-word/</guid>
    <pubDate>Wed, 29 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A document that carries instructions to the assistant that opens it, which then produces another document that does the same. The mechanism is old and the delivery is new: prompt injection stops being a chat-window curiosity the moment the payload is a file format your organisation already trusts and already forwards. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>Your agent needs a computer, not a container</title>
    <link>https://blog.cloudflare.com/cloudflare-computer/</link>
    <guid isPermaLink="false">2026-08-03:https://blog.cloudflare.com/cloudflare-computer/</guid>
    <pubDate>Mon, 03 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The capability half of the same week. The argument is that a per-agent persistent machine beats a stateless container for work that spans hours, and it arrives with an implementation rather than a manifesto. Worth reading against every sandbox story above, because it is a bet that the isolation unit should get larger and longer-lived, not smaller. (via Cloudflare -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>Welcome to Agents Week</title>
    <link>https://blog.cloudflare.com/agents-week-welcome/</link>
    <guid isPermaLink="false">2026-08-03:https://blog.cloudflare.com/agents-week-welcome/</guid>
    <pubDate>Sun, 02 Aug 2026 12:00:00 GMT</pubDate>
    <description>[relayed] An infrastructure vendor running a launch week aimed entirely at agents is itself the signal. The interesting question for an operator is which of these primitives you would still want if the agent framing went away, and that is a good filter for the whole category. (via Cloudflare -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>Gemini API Managed Agents: 3.6 Flash, hooks, and more</title>
    <link>https://blog.google/innovation-and-ai/technology/developers-tools/expanding-managed-agents-gemini-api-3-6-flash-hooks/</link>
    <guid isPermaLink="false">2026-08-03:https://blog.google/innovation-and-ai/technology/developers-tools/expanding-managed-agents-gemini-api-3-6-flash-hooks/</guid>
    <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Hooks in a managed agent service are the governance surface, so this is worth more attention than a model-version bump. Google is now shipping the hosted version of the thing the CLIs spent this year building locally, which raises the question our current issue keeps asking: whose runtime enforces the rule you wrote. (via Google -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>Stateless MCP has recaptured my interest</title>
    <link>https://simonwillison.net/2026/Jul/31/stateless-mcp/</link>
    <guid isPermaLink="false">2026-08-03:https://simonwillison.net/2026/Jul/31/stateless-mcp/</guid>
    <pubDate>Fri, 31 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Statelessness is a security property before it is an architecture preference: a server that holds no session holds nothing to steal and nothing to poison between calls. Comes with working code, which is the form of argument this field responds to. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>The Conductor Developer</title>
    <link>https://martinfowler.com/rachels-ramblings/conductor-developer.html</link>
    <guid isPermaLink="false">2026-08-03:https://martinfowler.com/rachels-ramblings/conductor-developer.html</guid>
    <pubDate>Fri, 31 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] On what the job becomes when the typing is delegated. Pairs with the orchestrator&apos;s-tax piece from last week: one names the machine cost of running several agents, this one names the human cost, and the honest version of the role is somewhere between conducting and reviewing. (via martinfowler.com -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>Advancing the price-performance frontier with GPT-5.6</title>
    <link>https://simonwillison.net/2026/Jul/30/luna-price-drop/</link>
    <guid isPermaLink="false">2026-08-03:https://simonwillison.net/2026/Jul/30/luna-price-drop/</guid>
    <pubDate>Thu, 30 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A price cut is a governance event whether or not anyone files it as one. Cheaper tokens move the default depth of every agent loop configured against a budget, and this issue&apos;s reporting on spend caps is the reason to care: a cap set in dollars means something different the week the dollars buy more. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-08-03/)</description>
  </item>
  <item>
    <title>Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident</title>
    <link>https://simonwillison.net/2026/Jul/28/anatomy-of-a-frontier-lab-agent-intrusion/</link>
    <guid isPermaLink="false">2026-07-29:https://simonwillison.net/2026/Jul/28/anatomy-of-a-frontier-lab-agent-intrusion/</guid>
    <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
    <description>[checked] The incident is real: OpenAI and Hugging Face published a joint statement on July 21 about an OpenAI agent reaching Hugging Face infrastructure during a model evaluation. The timeline reconstruction is Willison&apos;s. Read it as the first public case study of the thing every enforcement story this month gestured at: an agent with real authority crossing a boundary nobody thought to write down. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>Autonomous AI Intrusions Are Here: Lessons from the Hugging Face Compromise</title>
    <link>https://embracethered.com/blog/posts/2026/ai-intrusion-are-now-real/</link>
    <guid isPermaLink="false">2026-07-29:https://embracethered.com/blog/posts/2026/ai-intrusion-are-now-real/</guid>
    <pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The defensive read on the same incident, from the researcher whose prompt-injection work this field has been citing for two years. His frame: the interesting failure was not model behavior but that nothing between the agent and the target was checking authorization. Where have we heard that before. (via Johann Rehberger -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>GitPwned: Allowlist to RCE</title>
    <link>https://www.pillar.security/blog/gitpwned-allowlist-to-rce</link>
    <guid isPermaLink="false">2026-07-29:https://www.pillar.security/blog/gitpwned-allowlist-to-rce</guid>
    <pubDate>Mon, 27 Jul 2026 12:00:00 GMT</pubDate>
    <description>[checked] We adjudicated this one fully for the July issue. Codex&apos;s safe-command allowlist trusted `git show` by name; `git show --output` writes to any path. OpenAI confirmed CVSS 8.6, patched v0.95.0, paid the bounty -- and the advisory database still carries nothing newer than September 2025. Rated, patched, paid for, never announced. (via Pillar Security -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>The Orchestrator&apos;s Tax</title>
    <link>https://martinfowler.com/articles/orchestrator-tax.html</link>
    <guid isPermaLink="false">2026-07-29:https://martinfowler.com/articles/orchestrator-tax.html</guid>
    <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Names the overhead everyone benchmarks around: the harness, not the model, decides how much context gets carried, re-sent, and burned. A practitioner in our July corpus measured the same task at up to 4x runtime and token cost across three harnesses with similar output quality. The tax now has a name and a first estimate. (via martinfowler.com -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>How building software is changing at Anthropic</title>
    <link>https://newsletter.pragmaticengineer.com/p/inside-anthropic</link>
    <guid isPermaLink="false">2026-07-29:https://newsletter.pragmaticengineer.com/p/inside-anthropic</guid>
    <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Inside reporting that lines up with what Anthropic&apos;s own people said in posts we verified this month: auto mode as the default posture, most new code agent-written, review as the human&apos;s job. Useful as the long-form account of claims that have so far traveled as screenshots. (via Pragmatic Engineer -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>engineer away the slop</title>
    <link>https://ghuntley.com/slop/</link>
    <guid isPermaLink="false">2026-07-29:https://ghuntley.com/slop/</guid>
    <pubDate>Thu, 23 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The argument that slop is an engineering problem with engineering remedies: constrain generation, verify continuously, delete what does not earn its place. We spent this month applying the same doctrine to prose. It generalizes. (via Geoffrey Huntley -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>Introducing Claude Opus 5</title>
    <link>https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/</link>
    <guid isPermaLink="false">2026-07-29:https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/</guid>
    <pubDate>Fri, 24 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The capability half of the scaffolding story this publication has been tracking: practitioners were already reporting that newer models need fewer skills, shorter configuration files, less prompt ceremony. A new top-end model is the next test of whether the props keep dissolving. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>Discovering cryptographic weaknesses with Claude</title>
    <link>https://simonwillison.net/2026/Jul/28/discovering-cryptographic-weaknesses-with-claude/</link>
    <guid isPermaLink="false">2026-07-29:https://simonwillison.net/2026/Jul/28/discovering-cryptographic-weaknesses-with-claude/</guid>
    <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] An agent finding real cryptographic bugs is the other side of the Hugging Face week: the same capability class doing exactly what its operators pointed it at. What became possible and what stopped holding, in the same news cycle, from the same underlying fact. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>The Pulse: Grok&apos;s CLI caught uploading all your local files to the cloud</title>
    <link>https://newsletter.pragmaticengineer.com/p/the-pulse-groks-cli-caught-uploading</link>
    <guid isPermaLink="false">2026-07-29:https://newsletter.pragmaticengineer.com/p/the-pulse-groks-cli-caught-uploading</guid>
    <pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A coding CLI exfiltrating local files wholesale is the largest possible version of the enforcement gap, and we have not adjudicated it -- the claim is Orosz&apos;s reporting, relayed as such. It goes on next cycle&apos;s check list, not in the record. (via Pragmatic Engineer -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>Codex from 0 to 10M Users: Building ChatGPT Work</title>
    <link>https://www.latent.space/p/chatgpt-work</link>
    <guid isPermaLink="false">2026-07-29:https://www.latent.space/p/chatgpt-work</guid>
    <pubDate>Tue, 28 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The builder&apos;s own account of the harness we spend the most rows on. Worth an hour for one reason: the constraints he names -- approvals, sandboxes, resumability -- are the exact surfaces where this publication keeps finding the gaps. Hearing where the vendor thinks the hard parts are is its own kind of receipt. (via Latent Space -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-29/)</description>
  </item>
  <item>
    <title>From Indirect Prompt Injection to DNS Exfiltration in macOS Terminal</title>
    <link>https://embracethered.com/blog/posts/2026/macos-terminal-dillma-dns-exfil-ansi-escape-code-fix/</link>
    <guid isPermaLink="false">2026-07-22:https://embracethered.com/blog/posts/2026/macos-terminal-dillma-dns-exfil-ansi-escape-code-fix/</guid>
    <pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The attack does not need the agent to be tricked into running anything. It needs the terminal to render what the agent printed. ANSI escape sequences in model output, a DNS lookup, data gone. If your threat model stops at the agent&apos;s tool calls, this is the boundary you were not modelling. (via Johann Rehberger -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>The Pulse: Grok&apos;s CLI caught uploading all your local files to the cloud</title>
    <link>https://newsletter.pragmaticengineer.com/p/the-pulse-groks-cli-caught-uploading</link>
    <guid isPermaLink="false">2026-07-22:https://newsletter.pragmaticengineer.com/p/the-pulse-groks-cli-caught-uploading</guid>
    <pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] We have not adjudicated this and it is not in our record. Flagged because wholesale local-file exfiltration by a coding CLI is the largest possible version of the gap this publication tracks, and because it is on next cycle&apos;s check list rather than in this one&apos;s findings. (via Pragmatic Engineer -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>Antigravity CLI 1.1.3 is out</title>
    <link>https://x.com/shengzheyao/status/2077576699571741008</link>
    <guid isPermaLink="false">2026-07-22:https://x.com/shengzheyao/status/2077576699571741008</guid>
    <pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate>
    <description>[checked] The post says headless mode &quot;no longer hangs or silently auto-approves tools that need permission.&quot; Two days later 1.1.4&apos;s changelog recorded that headless runs had only then begun honouring settings.json at all -- not permissions, not file access, not sandbox mode. The announcement sat on top of a mode that enforced nothing, and 1.1.4 got no post. (via @shengzheyao -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>The Archaeologist&apos;s Copilot</title>
    <link>https://martinfowler.com/articles/archaeologist-copilot.html</link>
    <guid isPermaLink="false">2026-07-22:https://martinfowler.com/articles/archaeologist-copilot.html</guid>
    <pubDate>Thu, 16 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] On agents as a tool for understanding inherited code rather than producing new code. Pairs with Willison&apos;s piece four days later almost as a call-and-response: the cheap thing this year is not writing, it is reading what somebody else wrote. (via martinfowler.com -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>Reverse-engineering is cheap now</title>
    <link>https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/</link>
    <guid isPermaLink="false">2026-07-22:https://simonwillison.net/2026/Jul/20/cheap-reverse-engineering/</guid>
    <pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] The other half of the same argument, from someone who keeps doing it and publishing the results. Worth reading against every &quot;what became possible&quot; claim a vendor made this month, because this one arrives with worked examples instead of a benchmark. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>Claude Code uses Bun written in Rust now</title>
    <link>https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/</link>
    <guid isPermaLink="false">2026-07-22:https://simonwillison.net/2026/Jul/19/claude-code-in-bun-in-rust/</guid>
    <pubDate>Sun, 19 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] A runtime change in the tool most of this publication&apos;s rows are about. Mostly here because the detail is the kind operators discover through a startup regression rather than a release note. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>A Fireside Chat with Cat and Thariq from the Claude Code team</title>
    <link>https://simonwillison.net/2026/Jul/21/cat-and-thariq/</link>
    <guid isPermaLink="false">2026-07-22:https://simonwillison.net/2026/Jul/21/cat-and-thariq/</guid>
    <pubDate>Tue, 21 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] An annotated transcript, which is the format more of this field should use. The load-bearing claim inside it -- that Claude Code&apos;s system prompt shrank by roughly 80% for newer models -- is the vendor&apos;s own account and the changelog records no such reduction across v2.1.179 to v2.1.220. The direction is receipted; the magnitude is theirs. (via Simon Willison -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>Pushing software engineering limits with napkin math</title>
    <link>https://newsletter.pragmaticengineer.com/p/pushing-software-engineering-limits</link>
    <guid isPermaLink="false">2026-07-22:https://newsletter.pragmaticengineer.com/p/pushing-software-engineering-limits</guid>
    <pubDate>Tue, 21 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Estimation as the skill that survives. If an agent writes the code and you review it, the question you have to answer fastest is whether the numbers in front of you are the right order of magnitude. That is now a load-bearing human capability rather than an interview exercise. (via Pragmatic Engineer -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>Everyone Should Know SIMD</title>
    <link>https://mitchellh.com/writing/everyone-should-know-simd</link>
    <guid isPermaLink="false">2026-07-22:https://mitchellh.com/writing/everyone-should-know-simd</guid>
    <pubDate>Wed, 22 Jul 2026 12:00:00 GMT</pubDate>
    <description>[relayed] Not an agent story, and that is why it is here. It is the counter-argument to outsourcing all the typing: the parts of the machine worth understanding do not stop being worth understanding because something else can type them for you. (via Mitchell Hashimoto -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
  <item>
    <title>Claude Code 2.1.216: sandbox.filesystem.disabled</title>
    <link>https://x.com/ClaudeCodeLog/status/2079333123645485200</link>
    <guid isPermaLink="false">2026-07-22:https://x.com/ClaudeCodeLog/status/2079333123645485200</guid>
    <pubDate>Mon, 20 Jul 2026 12:00:00 GMT</pubDate>
    <description>[checked] Forty CLI changes, and the one to read is a new setting that skips filesystem isolation while keeping network egress control. That is a legitimate configuration for some workloads and a foot-gun for others. Know which one you are before you set it. (via @ClaudeCodeLog -- curated in the Bitter Frontier wire, https://frontier.bitter.sh/wire/2026-07-22/)</description>
  </item>
</channel>
</rss>
