The Harness Leaves The Chat Box
Operator Brief
The action in coding agents has left the model and the transcript. Two weeks of commits across eight projects are about goals, memory, visible computers, permissions, gateways, and supervision layers -- the environment around the agent getting thicker -- and the four sources new to this read (OpenClaw, Agent Zero, Paperclip, OpenHands) each show a different wall of the same building. The durable question is who owns the loop around all of it.
- Try
- Run at least one visible-computer harness. Agent Zero's browser, file browser, screenshots, and desktop surface expose failure modes terminal chat hides. Signal
- Prefer memory that asks first: Gemini's Auto Memory inbox proposes changes for review instead of writing them silently. Signal
- Read the permissions and sandbox story before an agent touches real credentials -- this fortnight shows who is actually doing that work. Signal
- Watch
Two weeks ago, Agent Zero fired the agent that used a browser and gave the agent a browser of its own. It replaced a browser-use module with a native browser, then added a Chromium runtime, tabs, screenshot previews, a searchable file browser, Linux desktop controls, a document canvas, a LibreOffice runtime, and OAuth and quota visibility. The "workcell" stopped being a metaphor. The agent has a computer now, and the operator can watch it work.
That is the loudest version of what every commit stream on this expanded watchlist said in the same fortnight: the interesting action in coding agents is no longer confined to the model or the chat transcript. Codex is adding persistent goals, session metadata, plugin controls, and cloud executor paths. Gemini CLI is treating memory as a reviewable patch. Hermes is sanding the rough edges off persistent personal agents. Pi keeps proving the opposite lesson -- a thin harness moves fast precisely because its integrations are disposable. And the four projects new to this read each expose a different wall of the same building: OpenClaw the front door (messaging surfaces, onboarding, visible progress), Agent Zero the machine room, Paperclip the management floor, OpenHands the whole leased office. The frontier is not one winning agent. It is the environment around agents getting thicker, and the durable question is who owns the loop around all of it.
State becomes product
The strongest single signal is still Codex
/goal,
and the telling part is not the feature but the follow-through: goal
validation, paste handling, queued-command behavior, user guidance. When a
persistent objective earns that much plumbing, it has stopped being a UX
affordance and become operating state. Gemini's
Auto Memory
inbox makes the same point from the other side, and makes it better than
anyone: memory should be proposed, reviewed, and accepted, not silently
smeared into hidden context. Hermes added memory scoping and
Curator
commands; OpenClaw put agent progress into the chat itself with
timeline spans.
Agent-side state is becoming durable, visible, and operational -- which means
a serious run now has to be able to answer what goal, memory, session, or
thread state shaped it.
The visible computer
Agent Zero's browser-and-desktop build-out leads this thread, but the platform side is converging on it too. OpenHands is grouping execution into sandbox groups with app-server routing, user secrets, and model profiles behind it. Paperclip is doing remote provisioning and sandbox-provider work. Codex is building cloud executor paths and hardening its sandbox. The chat box is not enough for serious agent work, and the projects that understand that are racing to show the operator the actual machine: the browser, the files, the runtime, the screenshots, the credentials, the artifacts.
The authority model comes to the foreground
This window is full of permissions work, and the spread is the story. Codex shipped permission profiles, sandbox profiles, plugin sharing controls, and Linux sandbox hardening. Gemini added workspace trust, private memory-patch allowlists, shell-safety evals, and approval-mode-aware subagents. OpenHands tightened redaction and deleted a log that had been recording secrets. OpenClaw fixed allowlists, subagent security docs, OAuth labels, and live exec output limits. Paperclip added security roles and sandbox-provider contracts; Agent Zero keeps its browser and office surfaces opt-in and exposes OAuth disconnect. The harness is starting to show its authority model, which is the right direction -- and the operator's question is finally answerable in some of these tools: what could this agent read, change, execute, install, send, or leak?
Accessibility is a frontier capability
OpenClaw is the corrective to an overly technical reading of this market. Its fortnight is setup recovery, stale plugin repair, Discord voice behavior, Telegram reactions, WhatsApp identity mapping, OAuth labels, progress previews, chat drafts, install recovery, and group allowlists -- work whose only purpose is letting a normal person start, understand, recover, and control an agent without learning the project's private ontology. Hermes is doing the adjacent work: setup fixes, voice push-to-talk parity, gateway restart readiness, provider pickers. Agent Zero's screenshot previews make the computer legible; Pi's quickstart and terminal work lower the floor; Gemini's reviewable memory and headless auth, and OpenHands' visible model names, do the same from their corners. None of this is softness. Accessibility is distribution, trust, and operator leverage, and the projects treating it as real engineering are buying something the benchmark chasers are not.
The control plane arrives
Paperclip makes the management problem explicit: runtime specs, sandbox providers, cost summaries, roles, liveness, stale-session recovery, ordered sub-issues, pause and resume. OpenHands is consolidating around its app server; Hermes runs kanban task workers, gateway lifecycle, Curator, and providers under a dashboard; Codex is reshaping skills, goals, sessions, and executors into app-server-shaped surfaces; OpenClaw manages gateway sessions, subagents, and plugin metadata. This is the factory problem in miniature: once agents coordinate across tasks and machines, something has to keep the system legible, and that something is becoming a product layer of its own.
Integrations are weather
Pi added providers, removed providers, and changed its Codex transport inside a single window. Hermes is moving model providers into plugins; OpenClaw is externalizing channel plugins; OpenHands is replacing config surfaces with app-server services; Codex and Gemini rework plugin, MCP, memory, and approval surfaces weekly. This is not a reason to avoid frontier tools. It is the reason to hold them through a loop that stays stable -- objective, permissions, execution environment, evidence, review, memory -- while the best agent, provider, runtime, and plugin change under it every week.
That loop is the fortnight's real subject. Every project above is building a piece of it inside its own walls. The operator who wants to switch walls without losing the work keeps the loop outside.
How this was read: this is a commit-harvest window -- commit metadata was broad-sampled across all eight projects, with diff-level review only on selected high-signal commits. Claude Code is absent because its v0 source contract defines no public commit stream. OpenClaw's high commit volume means its durable product movement is the hardest to separate from rapid stabilization; that caveat stands until a release-note review.
Revised 2026-07-02 (artifact_version 4): editorial pass to the current house standard -- lede, structure, and operator brief. Claims, receipts, and window judgments are unchanged from the 2026-05-07 publication.
Top signals from this issue
- Codex Persistent agent state is becoming a product surface
- Agent Zero The agent interface is becoming a visible computer
- Codex Permissions, secrets, and sandboxes are moving into the foreground
- OpenClaw Accessibility is a frontier capability, not marketing polish
- Paperclip Agent systems are growing control planes
- Pi Coding Agent Integrations are volatile; the operating loop has to be durable
Projects reviewed in this research run
Research artifacts and publication history are open in the repository.
Sources
Primary links, including exact changelog lines when available.
- commit diff reviewedValidate /goal objective length in TUIgithub.com/openai/codex/commit/f09e1936e0fd464dcea78fe55b84bd20f721cad6commitGoal lifecycle metricsgithub.com/openai/codex/commit/91b735018779daed7c40f86aab9bec9abc9922e8commitSpawn MCP for memoriesgithub.com/openai/codex/commit/ca257b6ce5db5c2710ec8da290b25b263154e402commitSession idgithub.com/openai/codex/commit/a98623511ba433154ec811fc63091617f5945438commitMCP turn metadata includes thread idgithub.com/openai/codex/commit/fe24a180ab6f6b3639b682cc6a1e71150fea6d48commitPlugin share access controlsgithub.com/openai/codex/commit/5119680f85ed01fe039ee8fba0245de24f3a5e37commitBundled Linux sandboxgithub.com/openai/codex/commit/26f355b67b75b040ff16990d1b2e4e8093479213
- commit diff reviewedAuto Memory inbox flow with canonical patch contractgithub.com/google-gemini/gemini-cli/commit/a7beb890d093e2cf66ed1ac8debff690b75e1f6dcommitTighten private Auto Memory patch allowlistgithub.com/google-gemini/gemini-cli/commit/7fb5146c6b084888b38dea05af6a4e95ea48810acommitWorkspace trust visible in MCP list UXgithub.com/google-gemini/gemini-cli/commit/a38f393af77c0ccf50da10d73c84cfb594dd8175commitShell command safety evalsgithub.com/google-gemini/gemini-cli/commit/82f6ea5b61a6321748d81a62d34c62bf7d2c9fa2commitSubagents aware of active approval modesgithub.com/google-gemini/gemini-cli/commit/40b384de2c1d251c9d13a6359216a9e6cff5a254commitJSON output for AgentExecutionStoppedgithub.com/google-gemini/gemini-cli/commit/469092a72cbe368b69df25c0caeefbc911b6d6fd
- commitSystemd restart readiness for gatewaygithub.com/NousResearch/hermes-agent/commit/d797755a1c17566b0aef4d77548a4b460142d26acommitSetup wizard does not dead-end on system-scope unitgithub.com/NousResearch/hermes-agent/commit/3cdbf334d5074aff0de857c0f94f278f06745e6bcommitVoice push-to-talk paritygithub.com/NousResearch/hermes-agent/commit/04cf4788ccc05003785992682e3cb25205e509cccommitDefault-large dashboard themegithub.com/NousResearch/hermes-agent/commit/6388aafbd6cbfd22c26036291d884d4055b5f6bccommitSearXNG native search backendgithub.com/NousResearch/hermes-agent/commit/5c906d70266c1bbce88fd227ea98a3f7646551fecommitPluggable model provider modulesgithub.com/NousResearch/hermes-agent/commit/9022804d78e88253d138d448e9107a3884b2b96ccommitLong-term memory scoping headergithub.com/NousResearch/hermes-agent/commit/fe8560fc1249b4a7e448b5c3b80a7d213df9d78fcommitCurator archive and prune subcommandsgithub.com/NousResearch/hermes-agent/commit/436672de0efd8bcc50c6043a16223c102d30d71b
- commitCached Codex websocket transportgithub.com/badlogic/pi-mono/commit/4745a9589883fb8200981ddfecb94a593d6e95a2commitFallback from Codex websocket to SSEgithub.com/badlogic/pi-mono/commit/370fdae6fa23881b044efbab571fb7bf6267ed6ecommitRemove Gemini CLI and Antigravity supportgithub.com/badlogic/pi-mono/commit/fe66edd943691f8eac295fef68ce36930c35fa05commitAdd Cloudflare AI Gateway providergithub.com/badlogic/pi-mono/commit/24fb6b833b7263df3d08889cc492b03d46d3779bcommitSearchable auth provider login flowgithub.com/badlogic/pi-mono/commit/010e9acfe959f437613bcba7139b264012ca43a4commitSession dir envgithub.com/badlogic/pi-mono/commit/8191d59c170c9bb336a82771e1826d25bb7ec1e0commitCompact read renderinggithub.com/badlogic/pi-mono/commit/588639fa97567a661dac876dc2b1970c8a3497ae
- commit diff reviewedRecover externalized channel plugin from stale configgithub.com/openclaw/openclaw/commit/329580c64d13657592c3fabb97ff567c2e292bb6commitLabel Claude CLI OAuth statusgithub.com/openclaw/openclaw/commit/2b4b60b5514b47d8e242b9b11d9b395037e6674bcommitPrevent Discord voice self-feedbackgithub.com/openclaw/openclaw/commit/1c2832526f65cf23b469e9a1dc5694915c5be548commitHonor Telegram access group allowlistsgithub.com/openclaw/openclaw/commit/b6ae0b83a61a1f779ee41b5d639b6049bfd422cecommitDocument sub-agent security boundariesgithub.com/openclaw/openclaw/commit/33b112ad314dc8d9dfe0f5a68caed4811a23245acommitBound live exec output eventsgithub.com/openclaw/openclaw/commit/3ee7c02bcacfdf6327747c1fe24dd6d11de8612acommitCoarse agent turn timeline spansgithub.com/openclaw/openclaw/commit/61223a74a43fd8768c426d5b22f1633dbad37477commitShow Codex tool progress in channel draftsgithub.com/openclaw/openclaw/commit/3f210b10ce3a19ef6a04205aa7420353945567a2
- commit diff reviewedAdapters declare runtime command spec for remote provisioninggithub.com/paperclipai/paperclip/commit/90631b09b36fa028ad24ca5375bfa50e3602799ccommitFix remote workspace environment shapinggithub.com/paperclipai/paperclip/commit/856c6cb192e53a992875821297b5fd8d29c95c2dcommitAdd sandbox callback bridge for remote environment API accessgithub.com/paperclipai/paperclip/commit/a4ac6ff133fbe8bdb82f4046fda85f7cb372b6a9commitAdd E2B sandbox provider plugingithub.com/paperclipai/paperclip/commit/4ef969f0840810527333aa6ee44fed89f4551f7ccommitIssue cost summariesgithub.com/paperclipai/paperclip/commit/c4269bab59fff7a73ff31797578cc97ece7f160fcommitFirst-class security agent rolegithub.com/paperclipai/paperclip/commit/c036bbfa98494dcfe2521aab65019a4cd021c769commitPause and resume sidebar agentsgithub.com/paperclipai/paperclip/commit/43b0f2ae582b18f2872ae60bf468f54b99b614ba
- commit diff reviewedReplace browser-use agent with native browsergithub.com/agent0ai/agent-zero/commit/983d431a5eb785eb9deba9fdfd471fa93f349603commitPersistent full Chromium runtime for Browsergithub.com/agent0ai/agent-zero/commit/fa7eef1919901093b117a98ad6e402d809687cf6commitBrowser multi-tab awareness and modifier-key clickgithub.com/agent0ai/agent-zero/commit/5012dd3128aa6218cc55f6cbce8be42b2db2fee4commitBrowser screenshot previews in tool messagesgithub.com/agent0ai/agent-zero/commit/c2fb2c3c94e1e1c85b783252332b3fc003f39f2bcommitLinux Desktop skill controlsgithub.com/agent0ai/agent-zero/commit/62ac20e7b248179825e05664c1df97ebc6214c54commitDesktop document canvasgithub.com/agent0ai/agent-zero/commit/24dd548ebf221e397323b5aa3a509f037fb1b9aecommitOAuth disconnect and remaining quota visibilitygithub.com/agent0ai/agent-zero/commit/0da8f3dc2b640efbce22499053507837101fdf6f
- commit diff reviewedStrengthen log redaction for API keysgithub.com/OpenHands/OpenHands/commit/61e3dc2cadbefd4e0649b7c141ac2335c021ad2bcommitRemove debug log exposing hook_config secretsgithub.com/OpenHands/OpenHands/commit/0c6c461555f8651347ed140f1c555ff8a88ddf56commitExpose sandbox grouping strategy UIgithub.com/OpenHands/OpenHands/commit/90cf5f8003c247597481bcbef9a5aa73eb899e10commitProxy Tavily MCP through app servergithub.com/OpenHands/OpenHands/commit/949a15a560ef90cd3dd7f18baf6955430401edb4commitMove server content to app_servergithub.com/OpenHands/OpenHands/commit/5232d96dab0ca98e691d6307bd0759e943220d1ccommitInject user secrets into ACP subprocess envgithub.com/OpenHands/OpenHands/commit/cf156b0073350ca8e93067bc2f4ae18b90537a0acommitSelf-hosted GitLab supportgithub.com/OpenHands/OpenHands/commit/4e63531fa6595ec55102f08ef129845931fcd8ffcommitRemoved V0 runtimegithub.com/OpenHands/OpenHands/commit/e86067c15b54242fd611877aa9038a2f7a219658