Watched Sources

The intake promise.

A source contract is the public commitment that defines where the loop looks for evidence, what it accepts as a finding, and what it refuses before a profile or digest can carry the claim. The cards expose the reader-facing contract: official surfaces, accepted and refused evidence, standing questions, and review state. The complete YAML remains linked.

codex / active / tier 1 / daily

Codex / OpenAI

Watch Codex as provider-native frontier capability, not just as an open-source CLI. Pay special attention to features that change how operators run it: long-horizon work, goals, subagents, workflows, sandboxing, permissions, AGENTS.md behavior, skills, plugins, MCP, browser/computer-use surfaces, non-interactive execution, SDKs, cost reporting, and enterprise governance.

Primary surfaces

Accepts as evidence

  • official changelog
  • official docs
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • official blog or developer post
  • package registry release
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
pricing or usage change
observe

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

High-signal patterns

goal / long-horizon / subagent / memory / workflow / sandbox / approval / permission / command / hook / MCP / plugin / skill / AGENTS.md / local environment / browser / computer use / automation / non-interactive / SDK / cost / usage

Operator questions

  • Does this change how an operator should run, trust, or govern Codex?
  • Does this affect long-horizon autonomous software work?
  • Does this affect verification, replay, resume, permissions, memory, or receipts?
  • Does this change whether teams should wrap, test, adapt, or ignore Codex capability?
  • Does this make serious agent work easier to start, inspect, or control without hiding authority?

Discovery state

last verified: 2026-05-06 / manual web / high confidence

  • Which GitHub releases, tags, and npm package versions should be treated as canonical when they disagree with the official Codex changelog?
  • Which provider-native long-horizon features should be detected through local probes rather than relying on release notes?

claude-code / active / tier 1 / daily

Claude Code / Anthropic

Watch Claude Code as a fast-moving provider-native coding environment with strong session, hook, plugin, skill, permission, and enterprise surfaces. Its changelog is granular; promote findings only when they change how developers should run it, trust it, review its output, or wrap it inside a longer-lived project workflow.

Primary surfaces

Accepts as evidence

  • official changelog
  • official docs
  • official whats new
  • package registry release
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
pricing or usage change
observe

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

High-signal patterns

recap / resume / rewind / plan / subagent / task / hook / permission / managed setting / plugin / skill / slash command / MCP / SDK / headless / telemetry / prompt caching / usage / model picker / enterprise

Operator questions

  • Does this change how an operator should govern Claude Code sessions?
  • Does this affect agent lifecycle, resume behavior, planning, permissions, or hooks?
  • Does this change whether teams should defer to Claude Code-native capability?
  • Does this create a new adapter, eval, receipt, or capability-profile need for teams running it in production?
  • Does this make serious agent work easier to start, inspect, or control without hiding authority?

Discovery state

last verified: 2026-05-06 / manual web / high confidence

  • Which GitHub source backing the published changelog should be captured directly in addition to the rendered official docs?
  • Which Claude Code behaviors should be probed locally because the changelog is too granular to imply operator impact by itself?

gemini-cli / active / tier 1 / daily

Gemini CLI / Google

Watch Gemini CLI as a large open-source terminal agent with rapid release channels, explicit context-file behavior, tool and extension surfaces, checkpointing, sandboxing, IDE/GitHub integrations, and Google account or Vertex/enterprise authentication paths. Separate stable operator guidance from preview/nightly churn.

Primary surfaces

Accepts as evidence

  • official docs
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • security advisory
  • package registry release
  • official google post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
pricing or usage change
observe

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

High-signal patterns

checkpoint / resume / context file / GEMINI.md / tool call / shell / web fetch / search grounding / MCP / extension / sandbox / trusted folder / permission / IDE / GitHub Action / output format / stream-json / authentication / enterprise / telemetry / preview channel / security

Operator questions

  • Does this change how an operator should trust Gemini CLI in a repo?
  • Does this affect checkpointing, context files, sandboxing, tools, auth, or output contracts?
  • Does this change adapter, receipt, eval, or run-contract assumptions for teams running it in production?
  • Does the release-channel cadence change how operators should test stable versus preview behavior?
  • Does this make serious agent work easier to start, inspect, or control without hiding authority?

Discovery state

last verified: 2026-05-06 / manual web / high confidence

  • Should nightly and preview releases be harvested into findings or only used for adapter-probe canaries?
  • Which security advisories should be treated as direct signals even when they do not change public docs?

antigravity / active / tier 1 / daily

Antigravity CLI / Google

Antigravity CLI (the `agy` binary) is Google's closed-source, Go successor to consumer Gemini CLI. Google announced the transition on 2026-05-19 and stopped serving Gemini CLI to consumer tiers (AI Pro/Ultra, free individual Code Assist, new GitHub-org installs) on 2026-06-18; enterprise Code Assist retained access and the open-source gemini-cli repo remains Apache-2.0 and enterprise-serving. Watch Antigravity as two things at once: the market-succession case (a tier-1 vendor retiring an open CLI and force-migrating consumers to a closed one) and the closed-source-governance case (its approval/sandbox model is real and active -- strict "Always Approve" rule matching, project-over-global permission precedence, proceed-in-sandbox, subagent auto-approval -- but you must trust the changelog rather than read the enforcement). The high-signal tension is exactly that: it hardens some gates and auto-opens others in the same release train.

Primary surfaces

Accepts as evidence

  • official docs
  • github release
  • official changelog
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • marketing claim without doc or code

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
governance change
test
lifecycle change
adapt

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • security
  • ecosystem
  • philosophy
  • lifecycle

Research lenses

  • governance
  • capability
  • coordination control plane
  • market succession
  • closed source verifiability

High-signal patterns

approval / Always Approve / permission rule / regex / sandbox / proceed-in-sandbox / --sandbox / --dangerously-skip-permissions / subagent / always proceeds / auto approve / project permissions / dangerous paths / migration / deprecation

Operator questions

  • Does a subagent's "always proceeds" auto-approval remove the human gate the parent agent was subject to?
  • What does the sandbox actually isolate, and which commands does proceed-in-sandbox auto-run without asking?
  • Antigravity is closed source -- how does an operator verify a governance claim they cannot read the code for?
  • What is the real migration cost and behavior change moving from consumer Gemini CLI to Antigravity CLI (agy)?
  • Which permission scope wins when project and global configs disagree, and can that be exploited?

Discovery state

last verified: 2026-07-01 / manual web / medium confidence

  • The code is closed; governance claims rest on docs/changelog, not readable enforcement. What is independently verifiable via local probe?
  • Is github.com/google-antigravity/antigravity-cli the canonical ship channel, or is antigravity.google the primary and the repo a mirror for releases/changelog?
  • What license, if any, governs the distributed binaries?

hermes-agent / active / tier 1 / daily

Hermes Agent / Nous Research

Hermes should be watched as a broad self-improving agent platform, not just as a coding CLI. Pay special attention to memory, skills, automations, messaging surfaces, subagents, sandboxing, runtime portability, and research trajectory generation. The operator question around tools like this is the project workflow that surrounds them: permissions, evidence, review, memory, and what the next run should know.

Primary surfaces

Accepts as evidence

  • official docs
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
pricing or usage change
observe

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

High-signal patterns

memory / skill / self-improvement / subagent / delegate / toolset / terminal backend / sandbox / container / SSH / Modal / Daytona / cron / messaging gateway / Telegram / Discord / Slack / MCP / context file / SOUL.md / llms.txt / trajectory / RL

Operator questions

  • Does this change how an operator should think about self-improving agents?
  • Does this affect long-horizon work, memory, skills, subagents, runtime portability, or messaging gateways?
  • Does this change whether teams should wrap Hermes as an agent engine, compare it, or adopt an adapter assumption?
  • Does this expose a governance, permission, receipt, or replay gap for teams running it in production?
  • Does this make serious agent work easier to start, inspect, or control without hiding authority?

Discovery state

last verified: 2026-05-06 / manual web / high confidence

  • Which docs domain should be considered canonical if GitHub README links and deployed docs diverge?
  • Which social or Discord announcements are maintainer-authored enough to include, and how should they be cited?

pi-coding-agent / active / tier 1 / daily

Pi Coding Agent / earendil-works / Mario Zechner

Watch Pi as a minimal, extensible terminal coding harness. It is important partly because of what it chooses not to include by default: subagents, plan mode, permission popups, MCP, and other governance features. That deliberate minimalism throws into relief the project workflow operators must build around a coding agent: durable goals, permissions, evidence, verification, and memory.

Primary surfaces

  • site official site https://pi.dev/
    watch: positioning / installation / modes / providers / design principles / package ecosystem
  • docs official docs https://pi.dev/docs/latest
    watch: quickstart / usage / sessions / context files / system prompt files / compaction / skills / extensions / prompt templates / themes / packages / rpc / sdk / providers / settings
  • watch: releases / tags / commits / pull requests / issues / packages coding agent / docs / examples
  • watch: package version / publication date / dist tags

    Corrected 2026-07-27. This contract previously watched @mariozechner/pi-coding-agent, which has been frozen at 0.73.1 since 2026-05-07 while the live package moved to the @earendil-works scope. Watching the abandoned name meant this source read as static for eleven weeks and missed the protobufjs fix. Anyone still installing the old scope is many minors behind.

Accepts as evidence

  • official docs
  • official site
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • package registry release
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
pricing or usage change
observe

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

High-signal patterns

extension / skill / package / prompt template / theme / session tree / branch / share / export / AGENTS.md / SYSTEM.md / compaction / dynamic context / RPC / SDK / json mode / provider / login / permission / sandbox / MCP / subagent / plan mode

Operator questions

  • Does this change how an operator can adapt the harness to their workflow?
  • Does this affect session portability, context shaping, extension design, package distribution, RPC, or SDK embedding?
  • Does Pi intentionally refusing a built-in feature force operators to supply it themselves at the extension or workflow layer?
  • Does this change whether teams should wrap Pi as an agent adapter or borrow an extension-contract idea?
  • Does this make serious agent work easier to start, inspect, or control without hiding authority?

Discovery state

last verified: 2026-05-12 / manual web / high confidence

Canonical repo migrated from badlogic/pi-mono to earendil-works/pi. pi.dev explicitly links to earendil-works/pi as the source. Both repos show identical release v0.74.0 (2026-05-07), confirming the migration is complete. Updated 2026-05-12.

  • Which package-registry or package-index surface should be watched for Pi extension ecosystem movement?

openclaw / active / tier 1 / daily

OpenClaw / OpenClaw

Watch OpenClaw as the accessibility calibration source for the agentic harness frontier. Its most important lesson may be product posture: making autonomous agent work feel reachable to everyday people. Pay special attention to onboarding, gateway surfaces, familiar channels, visual state, permissions, and any design move that hides setup complexity without hiding authority.

Primary surfaces

Accepts as evidence

  • official docs
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary
  • seo clone or mirror

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
accessibility change
study

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

Research lenses

  • accessibility
  • distribution surface
  • everyday use
  • gateway
  • authority visibility

High-signal patterns

onboarding / setup / gateway / visual surface / desktop / mobile / channel / notification / remote access / everyday user / natural language workflow / permission / approval / visibility / handoff / plugin / skill / daemon / background agent / long-running task / memory

Operator questions

  • Does this make agentic work more approachable to ordinary developers or everyday users?
  • Does this reduce setup, terminal fluency, or coordination burden?
  • Does it simplify the surface while preserving visible authority and control?
  • Does this demonstrate a more humane surface over rigorous internals that teams could adopt?
  • Does this suggest a new distribution surface beyond terminal-only agent work?

Discovery state

last verified: 2026-05-07 / manual web / medium confidence

  • Which OpenClaw release surface should be treated as canonical if docs and GitHub move at different speeds?
  • Which user-facing gateway surfaces are official product posture rather than experimental examples?
  • Which security and authority boundaries are visible enough for everyday users to understand?

paperclip / active / tier 1 / daily

Paperclip / Paperclip

Watch Paperclip as the coordination and economic-control-plane source. It poses the control-plane question: can agent work be organized into goals, roles, budgets, accountability, approvals, and operating state without becoming theater?

Primary surfaces

Accepts as evidence

  • official docs
  • official site
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
governance change
study

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

Research lenses

  • coordination control plane
  • economic governance
  • accountability
  • multi agent operations
  • operating state legibility

High-signal patterns

company / org chart / goal / budget / role / manager / employee / approval / governance / accountability / cost / task queue / progress / audit / session / dashboard / multi-agent / agent team

Operator questions

  • Does this make multi-agent labor more legible as operating state?
  • Does this expose goals, budgets, roles, approvals, and accountability in a way a human can govern?
  • Does this produce real control over agent work or only a company-themed dashboard?
  • Does this change how a control plane should model run contracts, allocation, or budgets for agent work?
  • Does this make coordination easier without hiding who approved what and why?

Discovery state

last verified: 2026-05-07 / manual web / medium confidence

  • Which source is canonical for product changes if the public site, docs, and GitHub repository diverge?
  • How much of the company/control-plane metaphor is backed by durable operating state versus UI framing?
  • Which governance and budget primitives are enforceable rather than descriptive?

agent-zero / active / tier 1 / daily

Agent Zero / agent0ai

Watch Agent Zero as the workcell-autonomy source. It raises the question of what happens when an agent gets a real computer environment, can use terminal/browser/files, and can grow tools or subagents inside that environment. Pay special attention to isolation, persistence, cleanup, visibility, and whether power remains governable.

Primary surfaces

Accepts as evidence

  • official docs
  • official site
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
runtime change
test

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

Research lenses

  • workcell autonomy
  • computer use
  • runtime isolation
  • tool creation
  • visible autonomy

High-signal patterns

Linux / terminal / file system / browser / code execution / Docker / container / tool creation / plugin / custom tool / subagent / memory / task / project / remote access / UI / safety / sandbox / persistence / cleanup

Operator questions

  • What does full computer access let the agent do that a narrow tool loop cannot?
  • How is the environment isolated, inspected, persisted, reset, or cleaned up?
  • What state is visible to the human while the agent acts?
  • What happens when the agent creates tools or subagents during a long run?
  • Does this change how operators should model containers, workcells, or full machines?

Discovery state

last verified: 2026-05-07 / manual web / high confidence

  • Which release or docs surface best describes the current runtime isolation model?
  • Which parts of Agent Zero's tool creation are safe to compare against an operator's own tool and receipt boundaries?
  • Which behaviors should operators test locally versus only study as product posture?

openhands / active / tier 1 / daily

OpenHands / OpenHands

Watch OpenHands as the productized software-agent platform source. What it signals is breadth: SDK, CLI, GUI, cloud, enterprise, integrations, sandboxing, collaboration, and evaluation in one system. Study what a full platform makes easier, and where operators are better served by a thin control layer than by adopting the whole platform.

Primary surfaces

Accepts as evidence

  • official docs
  • official site
  • github release
  • tagged release
  • maintainer commit
  • merged pr
  • security advisory
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
enterprise change
study

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

Research lenses

  • productized agent platform
  • sandboxed development
  • cli gui cloud surface
  • enterprise governance
  • evaluation

High-signal patterns

SDK / CLI / GUI / cloud / enterprise / self-hosting / sandbox / runtime / browser / evaluation / benchmark / security / RBAC / permission / collaboration / Slack / Jira / Linear / GitHub / extension / integration / multi-user

Operator questions

  • What happens when an agent harness becomes a full software-development platform?
  • Which surfaces are SDK, CLI, GUI, cloud, enterprise, or integration-specific?
  • How does OpenHands package sandboxing, permissions, collaboration, and evaluation for real teams?
  • Which parts should teams wrap, which parts should they learn from, and which parts should they refuse to adopt wholesale?
  • Does this make agentic software work easier to adopt without hiding evidence or authority?

Discovery state

last verified: 2026-05-07 / manual web / high confidence

  • Which OpenHands surfaces should be treated as one product versus separate SDK, CLI, cloud, and enterprise sources?
  • Which evaluation and sandboxing claims can be probed locally?
  • Which integrations change operator behavior enough to become signals?

heypi / active / tier 1 / daily

heypi / Ronan Berder (hunvreus)

heypi is the governance-shell calibration source: the human-in-the-loop and audit layer that wraps a minimal coding harness (Pi) for team chat-ops. Watch it as the inverse of Pi — where Pi refuses to bake governance into its core, heypi's entire product is that governance shell (approvals with named approvers, an SQLite audit trail, sandboxed tools, encrypted secret handoff, scoped memory). The high-signal question is always enforcement vs surfacing: does an approval block the call, is the audit trail complete, does the sandbox isolate. Pay attention to the durability disclaimer (heypi explicitly does not replay in-flight turns after a crash) — it marks the boundary of what the operator must own. heypi is also a multiplayer/team contrast to OpenClaw's single-user gateway, and a governance-first contrast to eve's durability-first framing. Do not over-weight landing-page feature copy; require a doc or commit before treating a governance feature as real and enforced.

Primary surfaces

  • watch: tags / commits / pull requests / releases / docs / examples / security
  • docs official docs https://heypi.dev/docs/
    watch: getting started / concepts / adapters / approvals / audit / sandboxing / secrets / memory / skills / jobs / admin / cli / security
  • landing official docs https://heypi.dev/
    watch: positioning / feature list / differentiation

Accepts as evidence

  • official docs
  • tagged release
  • github release
  • maintainer commit
  • merged pr
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary
  • marketing claim without doc or code

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
governance change
test

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy
  • accessibility

Research lenses

  • governance
  • accessibility
  • capability
  • coordination control plane
  • productized agent platform

High-signal patterns

approval / named approver / audit trail / SQLite / sandbox / just-bash / Docker / Gondolin / secret / memory scope / channel scope / skill / job / cron / heartbeat / admin panel / adapter / Slack / Discord / Telegram / webhook / Pi extension / durability / replay

Operator questions

  • Does an approval actually block the tool call, or only surface a notification after the fact?
  • Is the audit trail complete enough to reconstruct who approved what, and is it tamper-evident?
  • What does a sandbox runtime (just-bash, Docker, Gondolin) actually isolate, and what leaks across it?
  • How are secrets handed to a tool without passing through chat, and where do they rest?
  • In a multiplayer channel, whose authority binds — the requester's, the approver's, or the agent's config?
  • Does heypi inherit Pi's "no governance in core" posture, or does it genuinely close that gap?
  • What durability does heypi NOT provide (it disclaims crash-replay), and what does an operator have to own themselves?

Discovery state

last verified: 2026-06-24 / manual web / medium confidence

  • heypi publishes git tags but no GitHub "releases"; is the tag the canonical ship signal, or is npm publish the real channel?
  • Which governance surfaces (approvals, audit, sandbox, secrets) are on by default vs opt-in, and which bind at runtime vs surface only?
  • Which version line is "stable" once 0.2.0 leaves beta, and what is the support/upgrade contract?
  • What exactly does the Gondolin sandbox runtime provide relative to just-bash and Docker?

flue / active / tier 2 / weekly

Flue / withastro

Watch Flue as the programmable harness / headless agent calibration source. Its core framing, "Agent = Model + Harness," explicitly separates the model from the harness, filesystem, sandbox, skills, memory, sessions, and deployment surface. That is evidence for the thesis that the valuable layer is the shaped environment around the model, not just the model call itself. Treat it as category evidence and possible integration reference, not stable infrastructure. APIs are self-described as experimental; monitor direction before treating any primitive as architectural precedent.

Primary surfaces

  • watch: commits / releases / tags / pull requests / readme / changelog / examples / docs
  • watch: versions / breaking changes / new features / fixes

    Canonical receipt surface. Flue publishes no GitHub Releases (github.com/withastro/flue/releases is empty), so cite the version-tagged CHANGELOG.md (or a release tag) as the receipt, never the /releases page. Surfaced by the 2026-05-28_2026-06-03 run audit (item 2).

  • homepage official site https://flueframework.com/
    watch: framing / feature surface / deployment targets / skill system / sandbox api

Accepts as evidence

  • github commit
  • github release
  • tagged release
  • merged pr
  • readme change
  • official docs
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary
  • seo clone or mirror

Default actionability

release
observe
docs change
observe
api change
study
breaking change
note
ecosystem package
observe
philosophy change
study

Change types

  • capability
  • api surface
  • runtime
  • sandbox
  • skill system
  • deployment
  • protocol
  • philosophy
  • ecosystem
  • security
  • breaking change

Research lenses

  • agent harness architecture
  • model harness separation
  • programmable runtime
  • headless agent
  • sandbox design
  • skill primitives
  • session memory
  • ci deployment
  • filesystem abstraction

High-signal patterns

model + harness separation / programmable harness / headless agent / sandboxed execution / skill system / markdown skills / AGENTS.md / session management / memory / filesystem abstraction / HTTP server / CLI agent / CI/CD deployment / Cloudflare Workers / API change / breaking change / experimental

Operator questions

  • Does this validate or challenge the "model + harness" framing as a durable public category?
  • Which harness primitives (sandbox, filesystem, skills, memory, sessions, credentials) are stabilizing vs. still experimental?
  • Does this suggest integration surfaces teams should expose, wrap, or treat as precedent?
  • Does any API change affect how operators should think about their own receipt layer or deployment membrane?
  • Is the project gaining enough traction to treat as architectural reference rather than just category evidence?

Discovery state

last verified: 2026-06-03 / harvest run / medium confidence

  • Confirm GitHub repo is github.com/withastro/flue (withastro org is unusual for an agent harness project; verify ownership).
  • Is the Apache-2.0 license confirmed in the repo?
  • What is the actual star count and commit velocity at time of first harvest?
  • Are APIs stable enough to treat individual primitives as architectural precedent, or watch-only for now?

eve / active / tier 2 / weekly

eve / Vercel

Watch Eve as the filesystem-first, durable-execution agent framework on the authority axis. Its model, "an agent is a directory of files" (instructions, tools, skills, channels, schedules, subagents, connections, sandbox, hooks), states a portable, reviewable definition of an agent in public from a major vendor. Two lenses make it frontier-relevant: durable, resumable, crash-safe execution built on the open-source Workflow SDK, and human-in-the-loop approval gates that let an operator approve or deny a tool call before the agent proceeds. Treat it as category evidence and an authority-gate reference, not stable infrastructure: it is a fast-moving public beta under Vercel beta terms, with breaking changes arriving on minor versions. It is a general agent framework, watched as a harness on the coding and agent frontier, not a coding-only tool.

Primary surfaces

Accepts as evidence

  • github release
  • tagged release
  • github commit
  • merged pr
  • readme change
  • official docs
  • maintainer authored post
  • vendor changelog
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary
  • seo clone or mirror

Default actionability

release
observe
docs change
observe
api change
study
breaking change
note
approval gate change
study
sandbox change
study
philosophy change
study

Change types

  • capability
  • api surface
  • runtime
  • sandbox
  • skill system
  • subagents
  • connections
  • durable execution
  • approval gates
  • deployment
  • protocol
  • philosophy
  • ecosystem
  • security
  • breaking change

Research lenses

  • agent harness architecture
  • filesystem first agents
  • durable execution
  • human in the loop approval
  • sandbox design
  • subagent delegation
  • mcp connections
  • skill primitives
  • session recovery
  • gateway credentials

High-signal patterns

filesystem-first agent / agent is a directory of files / durable execution / resumable sessions / crash-safe / human in the loop / approval gate / tool call denied / sandbox backend / Microsandbox / Docker sandbox / subagent delegation / MCP connection / skills / schedules / channels / AI Gateway / OIDC / breaking change / public beta

Operator questions

  • Does Eve's approval-gate model change how teams gate tool calls before an agent proceeds?
  • Which primitives (sandbox backends, durable sessions, subagents, connections, approval gates) are stabilizing versus still moving on a fast beta?
  • Does the filesystem-first "agent is a directory of files" model make agent definitions more portable, reviewable, or version-controllable than code-first frameworks?
  • How does durable, resumable execution on the Workflow SDK change what operators can promise about crash recovery and human pauses mid-run?
  • Is the project stable enough to treat individual primitives as architectural precedent, or watch-only while the beta churns?

Discovery state

last verified: 2026-06-19 / harvest run / medium confidence

  • How stable is the public API across the rapid 0.10..0.11 release cadence (eight releases in the first days)?
  • Which of the three sandbox backends (Vercel, Microsandbox, Docker) are first-class versus best-effort in practice?
  • What does the human-in-the-loop approval surface look like end to end (who approves, where the pause is recorded, what an operator sees)?
  • Does the Workflow SDK dependency constrain where Eve agents can be hosted and recovered?

agent-flywheel / active / tier 2 / weekly

Agent Flywheel / Jeffrey Emanuel (Dicklesworthstone)

Watch Agent Flywheel as both an assembly layer and an operating method. It makes provider agents replaceable inside durable planning, task, coordination, verification, and memory artifacts, then distributes the supporting toolchain through one installer. That is unusually close to Frontier's Bitter Lesson and Amdahl questions. Weekly evidence remains bounded to ACFS releases, tagged docs, and attributed current claims on the official site. A dated ecosystem study may inspect the named related repositories and contextual paper with pinned receipts when their interaction changes the read.

Primary surfaces

  • watch: releases / tags / readme / docs

    Weekly change detection is release- and tag-bounded -- NOT commits or pull requests. README and docs claims are accepted only when pinned to a tag or its dereferenced commit. The repo had 3,440 main-branch commits at the 2026-07-02 intake against 7 tagged releases; the tag is the receipt surface that keeps this source finite. Never cite a moving main URL for a historical posture claim.

  • homepage official site https://agent-flywheel.com/
    watch: methodology / cost transparency / agent lineup / installer claims / security posture

Related surfaces

context only not weekly harvest

A selected set of repositories that implement the operating loop may be studied in explicitly scoped, dated work. This list is not the maintainer's complete portfolio, and these repositories are not silently promoted into weekly Agent Flywheel findings.

Related research

Accepts as evidence

  • tagged release
  • github release
  • tagged readme or docs
  • attributed official site claim for current posture
  • attributed project authored budget
  • contextual primary research paper
  • maintainer authored post with primary receipt
  • reproducible local probe

Refuses to promote

  • unspecified portfolio wide claim
  • untagged main branch commits
  • unsourced social claim
  • third party summary without primary link
  • self reported metric presented as independent fact
  • speculation
  • stale model memory
  • duplicate commentary
  • seo clone or mirror

Default actionability

release
study
docs change
observe
defaults change
note
security
note
cost change
observe
ecosystem package
ignore

Change types

  • defaults
  • security
  • capability
  • cost
  • distribution
  • agent lineup
  • breaking change
  • ecosystem
  • license

Research lenses

  • assembly layer
  • default setting authority
  • installer trust
  • cost of operation
  • multi agent cohabitation
  • individual account repository output
  • durable state outside provider harnesses
  • human liaison work
  • agent native interfaces
  • memory and feedback loops

High-signal patterns

install.sh behavior change / default settings written for Claude Code / Codex / Antigravity / permission or auto-mode posture set by the installer / credential, token, or session handling / version or channel pins for the bundled agents / agent added to or dropped from the lineup / sudo or system-level configuration change / safety-tool enforcement vs. advice / published cost figures updated

Operator questions

  • What defaults does the installer set across the three bundled tier-1 agents, and would their own vendors ship those defaults?
  • Does the installer pin agent versions and channels? An assembly layer inherits the released-is-not-merged problem for every tool it bundles.
  • Where do credentials and session state land on the VPS, and who can read them?
  • Do the bundled "safety tools" enforce anything, or advise? A warning is not a boundary.
  • Are the project's attributed operating-budget examples ($440-656/month at intake, including a two-Claude-account high end) holding as the lineup changes?
  • Does the plan-to-graph-to-coordination loop reduce total human attention across a completed project, or create new maintenance and review queues?
  • Does the OpenAI/Anthropic license rider apply to the operator or intended use? Potentially covered parties should review the tagged LICENSE and obtain their own legal guidance.

Discovery state

last verified: 2026-07-12 / exemplar and receipt audit / high confidence

  • Does the next tag make safe mode gate the dangerous Claude, Codex, and Antigravity shortcuts, remove ACFS-created NOPASSWD state when changing modes, and detect provider-supplied passwordless sudo?
  • Which bundled agent versions and release channels are pinned, and which still float with upstream latest installs?
  • What measured evidence would show that the coordination and memory loop improves verified outcomes per unit of human attention?
  • Will the non-standard OpenAI/Anthropic license rider change, and which operators need to obtain separate permission?

omnigent / active / tier 2 / weekly

Omnigent / omnigent-ai

Added 2026-08-02 as the first meta-harness on the watchlist. Everything else here IS a harness; Omnigent orchestrates them, describing itself as "an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents" and shipping policies, spend caps and access controls on top of harnesses that already have their own. That stacking is why it is worth tracking. This publication's standing argument is that a control existing only as an intention is not a control, and a meta-harness is the hardest test of it: two governance layers now have a claim on the same action, and an operator has to know which one refuses. That question has no public answer yet, which makes it the most interesting open item on the list. Release cadence at intake was roughly weekly, v0.2.0 through v0.7.0 between 2026-06-19 and 2026-07-27, pre-1.0, so the tag-to-tag diff carries more than the release note. Two items in v0.7.0 deserve a second look: optional server-side transcription is a new data path off the operator's machine, and a router that selects the harness means the governance layer an action lands under can change without the operator choosing it. Handling: this is Tier 2 and weekly - do not promote on release velocity alone. A finding observed through Omnigent is a finding about the pair, not about the wrapped harness, and must say so. Adapter lag is a legitimate finding but is not a defect in the harness underneath. Closest comparison is Paperclip, and the contrast is the useful part: Paperclip manages an org of agents it owns, Omnigent orchestrates agents it does not.

Primary surfaces

Related surfaces

watched

Omnigent drives harnesses this publication already tracks separately. Behaviour observed through it is a fact about the pair, never about either component alone, and a finding must say which.

Accepts as evidence

  • official docs
  • tagged release
  • github release
  • maintainer commit
  • merged pr
  • maintainer authored post
  • reproducible local probe

Refuses to promote

  • unsourced social claim
  • third party summary without primary link
  • speculation
  • stale model memory
  • benchmark claim without method
  • duplicate commentary
  • marketing claim without doc or code

Default actionability

release
test
docs change
observe
security change
test
breaking change
adapt
ecosystem package
observe
governance change
test

Change types

  • capability
  • workflow
  • runtime
  • protocol
  • reliability
  • economics
  • security
  • ecosystem
  • evaluation
  • philosophy

Research lenses

  • governance
  • coordination control plane
  • productized agent platform
  • capability
  • economics

High-signal patterns

policy / deny / allow / approval / spend cap / budget / quota / access control / adapter / harness / router / routing / sandbox / isolation / session sync / transcription / scheduled task / automation / project / team / breaking change

Operator questions

  • When an Omnigent policy and the underlying harness's own permission system disagree, which one refuses, and is that documented or only observed?
  • Are spend caps enforced before a call is made, or reconciled after it? A cap that reconciles is a report, not a control.
  • When a wrapped harness changes a permission surface, how long does the adapter lag, and what does an operator's policy mean during the gap?
  • The v0.7.0 router picks both harness and model. Can the governance layer an action lands under change without the operator choosing it?
  • Server-side transcription is a data path off the operator's machine. What is sent, what is retained, and is it opt-in?
  • Is the governance work in the tag an operator installs, or on main?
  • Pre-1.0 and shipping weekly: what is the upgrade contract, and what breaks between minors?

Discovery state

last verified: 2026-08-02 / manual web / medium confidence

  • Which policy decisions are enforced in Omnigent's process versus delegated to the wrapped harness?
  • Does the adapter surface a harness's own refusal to the operator, or swallow it?
  • What is the canonical ship signal - the GitHub release, the tag, or a package publish?
  • Which of the governance features in the landing copy are on by default?