A source contract is the public commitment that defines where the loop looks for evidence, what
it accepts as a finding, and what it refuses before a profile or digest can carry the claim.
The cards expose the reader-facing contract: official surfaces, accepted and refused evidence,
standing questions, and review state. The complete YAML remains linked.
codex / active / tier 1 / daily
Codex / OpenAI
Watch Codex as provider-native frontier capability, not just as an open-source CLI. Pay special attention to features that change how operators run it: long-horizon work, goals, subagents, workflows, sandboxing, permissions, AGENTS.md behavior, skills, plugins, MCP, browser/computer-use surfaces, non-interactive execution, SDKs, cost reporting, and enterprise governance.
Primary surfaces
-
watch: releases / new features / improvements / bug fixes / breaking changes
-
watch: cli / app / ide extension / web / workflows / subagents / sandboxing / memories / commands / agents md / mcp / plugins / skills / authentication / approvals / security / governance / automation
-
watch: releases / tags / commits / pull requests / issues / docs / examples / security
-
watch: package version / publication date / install surface
Accepts as evidence
- official changelog
- official docs
- github release
- tagged release
- maintainer commit
- merged pr
- official blog or developer post
- package registry release
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- pricing or usage change
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
High-signal patterns
goal / long-horizon / subagent / memory / workflow / sandbox / approval / permission / command / hook / MCP / plugin / skill / AGENTS.md / local environment / browser / computer use / automation / non-interactive / SDK / cost / usage
Operator questions
- Does this change how an operator should run, trust, or govern Codex?
- Does this affect long-horizon autonomous software work?
- Does this affect verification, replay, resume, permissions, memory, or receipts?
- Does this change whether teams should wrap, test, adapt, or ignore Codex capability?
- Does this make serious agent work easier to start, inspect, or control without hiding authority?
Discovery state
last verified: 2026-05-06 / manual web / high confidence
- Which GitHub releases, tags, and npm package versions should be treated as canonical when they disagree with the official Codex changelog?
- Which provider-native long-horizon features should be detected through local probes rather than relying on release notes?
claude-code / active / tier 1 / daily
Claude Code / Anthropic
Watch Claude Code as a fast-moving provider-native coding environment with strong session, hook, plugin, skill, permission, and enterprise surfaces. Its changelog is granular; promote findings only when they change how developers should run it, trust it, review its output, or wrap it inside a longer-lived project workflow.
Primary surfaces
-
watch: releases / new features / improvements / bug fixes / breaking changes
-
watch: notable features / examples / operator guidance / weekly rollups
-
watch: memory / hooks / slash commands / plugins / skills / subagents / permissions / settings / sandboxing / mcp / sdk / headless / telemetry / enterprise
-
watch: package version / publication date / install surface
Accepts as evidence
- official changelog
- official docs
- official whats new
- package registry release
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- pricing or usage change
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
High-signal patterns
recap / resume / rewind / plan / subagent / task / hook / permission / managed setting / plugin / skill / slash command / MCP / SDK / headless / telemetry / prompt caching / usage / model picker / enterprise
Operator questions
- Does this change how an operator should govern Claude Code sessions?
- Does this affect agent lifecycle, resume behavior, planning, permissions, or hooks?
- Does this change whether teams should defer to Claude Code-native capability?
- Does this create a new adapter, eval, receipt, or capability-profile need for teams running it in production?
- Does this make serious agent work easier to start, inspect, or control without hiding authority?
Discovery state
last verified: 2026-05-06 / manual web / high confidence
- Which GitHub source backing the published changelog should be captured directly in addition to the rendered official docs?
- Which Claude Code behaviors should be probed locally because the changelog is too granular to imply operator impact by itself?
gemini-cli / active / tier 1 / daily
Gemini CLI / Google
Watch Gemini CLI as a large open-source terminal agent with rapid release channels, explicit context-file behavior, tool and extension surfaces, checkpointing, sandboxing, IDE/GitHub integrations, and Google account or Vertex/enterprise authentication paths. Separate stable operator guidance from preview/nightly churn.
Primary surfaces
-
watch: releases / tags / commits / pull requests / issues / discussions / docs / roadmap / security
-
watch: stable / preview / nightly / breaking changes / security
-
watch: changelog / installation / authentication / configuration / commands / context files / checkpointing / tools / mcp / extensions / headless / ide / sandboxing / trusted folders / enterprise / telemetry
-
watch: package version / publication date / dist tags
Accepts as evidence
- official docs
- github release
- tagged release
- maintainer commit
- merged pr
- security advisory
- package registry release
- official google post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- pricing or usage change
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
High-signal patterns
checkpoint / resume / context file / GEMINI.md / tool call / shell / web fetch / search grounding / MCP / extension / sandbox / trusted folder / permission / IDE / GitHub Action / output format / stream-json / authentication / enterprise / telemetry / preview channel / security
Operator questions
- Does this change how an operator should trust Gemini CLI in a repo?
- Does this affect checkpointing, context files, sandboxing, tools, auth, or output contracts?
- Does this change adapter, receipt, eval, or run-contract assumptions for teams running it in production?
- Does the release-channel cadence change how operators should test stable versus preview behavior?
- Does this make serious agent work easier to start, inspect, or control without hiding authority?
Discovery state
last verified: 2026-05-06 / manual web / high confidence
- Should nightly and preview releases be harvested into findings or only used for adapter-probe canaries?
- Which security advisories should be treated as direct signals even when they do not change public docs?
antigravity / active / tier 1 / daily
Antigravity CLI / Google
Antigravity CLI (the `agy` binary) is Google's closed-source, Go successor to consumer Gemini CLI. Google announced the transition on 2026-05-19 and stopped serving Gemini CLI to consumer tiers (AI Pro/Ultra, free individual Code Assist, new GitHub-org installs) on 2026-06-18; enterprise Code Assist retained access and the open-source gemini-cli repo remains Apache-2.0 and enterprise-serving. Watch Antigravity as two things at once: the market-succession case (a tier-1 vendor retiring an open CLI and force-migrating consumers to a closed one) and the closed-source-governance case (its approval/sandbox model is real and active -- strict "Always Approve" rule matching, project-over-global permission precedence, proceed-in-sandbox, subagent auto-approval -- but you must trust the changelog rather than read the enforcement). The high-signal tension is exactly that: it hardens some gates and auto-opens others in the same release train.
Primary surfaces
-
-
watch: releases / changelog / breaking changes / security
-
watch: approvals / permissions / sandbox / subagents / mcp / plugins / breaking changes
-
watch: cli using / approvals / permissions / sandbox / configuration / migration
-
watch: lifecycle / deprecation / migration / tier access
Accepts as evidence
- official docs
- github release
- official changelog
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- marketing claim without doc or code
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- governance change
- test
- lifecycle change
- adapt
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- security
- ecosystem
- philosophy
- lifecycle
Research lenses
- governance
- capability
- coordination control plane
- market succession
- closed source verifiability
High-signal patterns
approval / Always Approve / permission rule / regex / sandbox / proceed-in-sandbox / --sandbox / --dangerously-skip-permissions / subagent / always proceeds / auto approve / project permissions / dangerous paths / migration / deprecation
Operator questions
- Does a subagent's "always proceeds" auto-approval remove the human gate the parent agent was subject to?
- What does the sandbox actually isolate, and which commands does proceed-in-sandbox auto-run without asking?
- Antigravity is closed source -- how does an operator verify a governance claim they cannot read the code for?
- What is the real migration cost and behavior change moving from consumer Gemini CLI to Antigravity CLI (agy)?
- Which permission scope wins when project and global configs disagree, and can that be exploited?
Discovery state
last verified: 2026-07-01 / manual web / medium confidence
- The code is closed; governance claims rest on docs/changelog, not readable enforcement. What is independently verifiable via local probe?
- Is github.com/google-antigravity/antigravity-cli the canonical ship channel, or is antigravity.google the primary and the repo a mirror for releases/changelog?
- What license, if any, governs the distributed binaries?
grok-build / active / tier 1 / daily
Grok Build / xAI
Added 2026-08-21. xai-org/grok-build, Apache-2.0, Rust, binary `grok`. Open-sourced 2026-07-15. The README states the tree is synced periodically from the SpaceXAI monorepo and that SOURCE_REV records the internal commit. GitHub tags and releases are both empty at intake. Operators install with `curl -fsSL https://x.ai/cli/install.sh | bash`. The changelog is https://x.ai/build/changelog.
It earns a slot as the inspectable frontier-lab CLI that was missing next to Codex, Claude Code, and Gemini CLI. Source is the Codex/Pi floor. Channel is not: public main is a mirror, so a gap in git history proves nothing about the binary, and default install is a moving script. Pin SOURCE_REV for a code claim. Pin the changelog version and, if you can, the installed `grok --version` for a ship claim. Never treat grok.com chat as this product. Never promote an Omnigent ACP wrap into a fact about Grok Build's own gate.
Primary surfaces
-
-
watch: commits / source rev / license / sandbox / permissions / hooks / skills / plugins / acp / worktrees
-
watch: version / permissions / sandbox / hooks / acp / headless / breaking changes / security
-
watch: install / headless / acp / skills / plugins / hooks / mcp / permissions / enterprise
-
watch: resolved version / platform matrix / integrity verification
-
watch: license / scope / sync from monorepo
Related surfaces
watched
The public GitHub tree is a periodic sync from the SpaceXAI monorepo, not the development home. grok.com chat, Grok the model, Grok Bot, and this publication's Lane C use of the grok CLI are different objects. Omnigent can drive Grok Build over ACP; that is a fact about the pair, not about Grok Build's own tool policy.
Accepts as evidence
- official changelog
- official docs
- tagged raw file
- git commit
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- marketing claim without doc or code
- grok com chat as harness
- grok model as harness
- unsynced main as the product
- omnigent acp wrap as grok gate
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- governance change
- test
- channel divergence
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
Research lenses
- governance
- capability
- distribution channel
- closed source verifiability
High-signal patterns
SOURCE_REV / Synced from monorepo / sandbox / permission / auto mode / hook / ACP / grok agent stdio / skill / plugin / worktree / headless / install.sh / GROK_CONFIG
Operator questions
- The public tree is a periodic sync from a private monorepo and GitHub has zero tags: which channel is an operator actually running, SOURCE_REV, the changelog version, or whatever the install script fetched today?
- Does a permission, sandbox, or hook rule read at a public SHA bind in the binary the install script just put on PATH?
- Native ACP is `grok agent stdio`. When Omnigent or another wrapper drives that socket, whose tool policy refuses?
- How far behind the changelog is the last public sync, and does a gap in public git prove anything about what shipped?
- Is grok.com chat, Grok Bot, or the grok-4.6 API being described as if it were this CLI?
Discovery state
last verified: 2026-08-21 / manual web / medium confidence
- Will xAI ever cut GitHub tags or Releases, or is the changelog plus install script the permanent ship channel?
- How often does SOURCE_REV move relative to a changelog version, and can an operator join them without a vendor digest?
- What license, if any, covers the distributed binary as distinct from the Apache-2.0 source tree?
cursor / active / tier 1 / daily
Cursor / Anysphere
Added 2026-08-21 as a closed changelog source, the same floor as Claude Code and Antigravity. Anysphere binaries: the editor, plus the terminal CLI installed with `curl https://cursor.com/install -fsS | bash` (docs command: `agent`). github.com/cursor/cursor is not the product.
It earns a slot because operators already run it and because Omnigent already names it as a harness. A meta-harness finding that says "do not attribute Cursor behavior to Omnigent" is incomplete if this publication has no Cursor surface of its own. Harvest the changelog and the CLI docs. Name the surface. Do not promote a cloud-agent entry into a CLI fact. Do not open the GitHub tracker looking for agent source; it is not there.
Primary surfaces
-
-
watch: editor / cli / cloud agents / origin / permissions / sandbox / subagents / goals / breaking changes
-
watch: install / modes / headless / sandbox / sudo / cloud handoff / sessions
-
watch: resolved version / binary name / platform matrix
-
watch: cloud agent / origin / plugins / permissions / enterprise
Related surfaces
watched
github.com/cursor/cursor is a bug-tracker README. It is not the editor and not the agent. Do not harvest it as source. The changelog mixes editor, CLI, cloud agents, and Origin code hosting; a finding must name which surface it is about. Omnigent can route work onto Cursor; that is a fact about the pair, not about Cursor's own permission system.
Accepts as evidence
- official changelog
- official docs
- package registry release
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- marketing claim without doc or code
- github com cursor cursor as source
- source diff of cursor github
- cloud agent changelog as cli claim
- origin hosting as permission model
- omnigent wrap as cursor gate
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- governance change
- test
- changelog entry
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
Research lenses
- governance
- capability
- distribution channel
- closed source verifiability
High-signal patterns
changelog / agent / cursor-agent / cloud agent / Origin / /goal / subagent / sandbox / sudo / plan mode / headless / permissions
Operator questions
- github.com/cursor/cursor is a bug tracker: did this finding treat it as source?
- Which surface moved, the editor, the `agent` CLI, cloud agents, or Origin hosting, and does the changelog actually say that?
- Can an operator pin a CLI version string or a download digest, or is the dated changelog the entire channel?
- What does `/sandbox` actually isolate, and does sudo prompting mean the model never sees the password, or only that the docs say so?
- When Omnigent routes work onto Cursor, whose permission system refuses?
Discovery state
last verified: 2026-08-21 / manual web / medium confidence
- Is there a versioned CLI release channel with a pinable digest, or only `curl https://cursor.com/install | bash` plus a website changelog?
- Do editor, CLI, and cloud agents share one permission model, or three that the changelog does not separate?
- Does Origin hosting change where an agent writes, and is that described as a control or as a product feature?
github-copilot-cli / active / tier 1 / daily
GitHub Copilot CLI / GitHub
Added 2026-08-21 as a closed changelog source. github/copilot-cli is public and busy: LICENSE.md (proprietary grant to install and run, not OSS), changelog.md, install.sh, GitHub Releases, npm `@github/copilot`. The tree is not the agent.
It is the most pinable of the large closed terminal agents: you can date a GitHub release and an npm publish the way this publication dates Claude Code. At intake, npm `latest` and the newest non-prerelease GitHub tag were both 1.0.80 (2026-08-14); `prerelease` was 1.0.81-7 (2026-08-21). Name the dist-tag.
Do not mix this with VS Code Copilot, Copilot Chat, the cloud coding agent, or `gh copilot`. Do not read the public repo as if it contained enforcement. Do not treat Omnigent's `gh auth login` as Copilot's own gate.
Primary surfaces
-
watch: sandbox / permissions / allow all / enterprise policy / worktree / mcp / plugins / breaking changes / security
-
watch: releases / prerelease flag / changelog / breaking changes
-
watch: dist tags / package version / publication date / latest vs prerelease
-
watch: approval / sandbox / trusted directories / allow all tools / acp / mcp / hooks / skills / enterprise
-
watch: license / install sh / changelog / readme
Related surfaces
watched
This source is the Copilot CLI (`copilot`, npm `@github/copilot`). VS Code Copilot, Copilot Chat, the Copilot cloud coding agent, and the older `gh copilot` extension are different products. github/copilot-cli is license, docs, changelog, and installer -- not the agent source. Omnigent's Copilot path is `gh auth login`; that is wrapper auth.
Accepts as evidence
- official changelog
- github release
- official docs
- package registry release
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- marketing claim without doc or code
- vscode copilot as cli
- copilot cloud agent as cli
- gh copilot extension as cli
- source diff of copilot cli repo
- omnigent gh auth as copilot gate
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- governance change
- test
- prerelease
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
Research lenses
- governance
- capability
- distribution channel
- closed source verifiability
High-signal patterns
sandbox / allow-all / allow-auto-only / --allow-all-tools / trusted director / worktree / enterprise / ACP / MCP / plugin / hook / skill / prerelease
Operator questions
- npm latest and GitHub's non-prerelease tag can lag the prerelease dist-tag: which channel did the operator actually install?
- Does a sandbox or allow-all sentence in changelog.md bind in the binary, or only on the page?
- Is this the Copilot CLI, VS Code Copilot, the cloud coding agent, or the older `gh copilot` extension?
- Enterprise allow-auto-only and sandbox proxy policy: who can turn them off, and does the CLI honor org policy when the operator is local?
- Docs describe Copilot CLI as an ACP server. Is that a shippable surface, and whose permissions apply when a wrapper drives it?
- Omnigent authenticates Copilot with `gh auth login`. Is that Copilot's gate, or the wrapper's?
Discovery state
last verified: 2026-08-21 / manual web / high confidence
- Does github/copilot-cli ever grow agent source, or is the changelog-plus-binary floor permanent?
- How often does npm `latest` lag GitHub prereleases, and is `prerelease` a channel an operator should run?
- Can org-level MCP and sandbox policies that docs say the CLI cannot honor be confirmed by a local probe?
hermes-agent / active / tier 1 / daily
Hermes Agent / Nous Research
Hermes should be watched as a broad self-improving agent platform, not just as a coding CLI. Pay special attention to memory, skills, automations, messaging surfaces, subagents, sandboxing, runtime portability, and research trajectory generation. The operator question around tools like this is the project workflow that surrounds them: permissions, evidence, review, memory, and what the next run should know.
Primary surfaces
-
watch: releases / tags / commits / pull requests / issues / docs / examples / security
-
watch: release notes / breaking changes / migration notes / security
-
watch: installation / configuration / tools / toolsets / memory / skills / mcp / messaging / cron / security / terminal backends / architecture / context files / llms txt
Accepts as evidence
- official docs
- github release
- tagged release
- maintainer commit
- merged pr
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- pricing or usage change
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
High-signal patterns
memory / skill / self-improvement / subagent / delegate / toolset / terminal backend / sandbox / container / SSH / Modal / Daytona / cron / messaging gateway / Telegram / Discord / Slack / MCP / context file / SOUL.md / llms.txt / trajectory / RL
Operator questions
- Does this change how an operator should think about self-improving agents?
- Does this affect long-horizon work, memory, skills, subagents, runtime portability, or messaging gateways?
- Does this change whether teams should wrap Hermes as an agent engine, compare it, or adopt an adapter assumption?
- Does this expose a governance, permission, receipt, or replay gap for teams running it in production?
- Does this make serious agent work easier to start, inspect, or control without hiding authority?
Discovery state
last verified: 2026-05-06 / manual web / high confidence
- Which docs domain should be considered canonical if GitHub README links and deployed docs diverge?
- Which social or Discord announcements are maintainer-authored enough to include, and how should they be cited?
pi-coding-agent / active / tier 1 / daily
Pi Coding Agent / earendil-works / Mario Zechner
Watch Pi as a minimal, extensible terminal coding harness. It is important partly because of what it chooses not to include by default: subagents, plan mode, permission popups, MCP, and other governance features. That deliberate minimalism throws into relief the project workflow operators must build around a coding agent: durable goals, permissions, evidence, verification, and memory.
Primary surfaces
-
watch: positioning / installation / modes / providers / design principles / package ecosystem
-
watch: quickstart / usage / sessions / context files / system prompt files / compaction / skills / extensions / prompt templates / themes / packages / rpc / sdk / providers / settings
-
watch: releases / tags / commits / pull requests / issues / packages coding agent / docs / examples
-
watch: package version / publication date / dist tags
Corrected 2026-07-27. This contract previously watched @mariozechner/pi-coding-agent, which has been frozen at 0.73.1 since 2026-05-07 while the live package moved to the @earendil-works scope. Watching the abandoned name meant this source read as static for eleven weeks and missed the protobufjs fix. Anyone still installing the old scope is many minors behind.
Accepts as evidence
- official docs
- official site
- github release
- tagged release
- maintainer commit
- merged pr
- package registry release
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- pricing or usage change
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
High-signal patterns
extension / skill / package / prompt template / theme / session tree / branch / share / export / AGENTS.md / SYSTEM.md / compaction / dynamic context / RPC / SDK / json mode / provider / login / permission / sandbox / MCP / subagent / plan mode
Operator questions
- Does this change how an operator can adapt the harness to their workflow?
- Does this affect session portability, context shaping, extension design, package distribution, RPC, or SDK embedding?
- Does Pi intentionally refusing a built-in feature force operators to supply it themselves at the extension or workflow layer?
- Does this change whether teams should wrap Pi as an agent adapter or borrow an extension-contract idea?
- Does this make serious agent work easier to start, inspect, or control without hiding authority?
Discovery state
last verified: 2026-05-12 / manual web / high confidence
Canonical repo migrated from badlogic/pi-mono to earendil-works/pi. pi.dev explicitly links to earendil-works/pi as the source. Both repos show identical release v0.74.0 (2026-05-07), confirming the migration is complete. Updated 2026-05-12.
- Which package-registry or package-index surface should be watched for Pi extension ecosystem movement?
openclaw / active / tier 1 / daily
OpenClaw / OpenClaw
Watch OpenClaw as the accessibility calibration source for the agentic harness frontier. Its most important lesson may be product posture: making autonomous agent work feel reachable to everyday people. Pay special attention to onboarding, gateway surfaces, familiar channels, visual state, permissions, and any design move that hides setup complexity without hiding authority.
Primary surfaces
-
watch: releases / tags / commits / pull requests / issues / docs / examples / security
-
watch: getting started / gateway / installation / configuration / channels / plugins / skills / permissions / memory / remote access / security / mobile or desktop surfaces
-
watch: onboarding / setup steps / first run / user workflow
Accepts as evidence
- official docs
- github release
- tagged release
- maintainer commit
- merged pr
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- seo clone or mirror
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- accessibility change
- study
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
Research lenses
- accessibility
- distribution surface
- everyday use
- gateway
- authority visibility
High-signal patterns
onboarding / setup / gateway / visual surface / desktop / mobile / channel / notification / remote access / everyday user / natural language workflow / permission / approval / visibility / handoff / plugin / skill / daemon / background agent / long-running task / memory
Operator questions
- Does this make agentic work more approachable to ordinary developers or everyday users?
- Does this reduce setup, terminal fluency, or coordination burden?
- Does it simplify the surface while preserving visible authority and control?
- Does this demonstrate a more humane surface over rigorous internals that teams could adopt?
- Does this suggest a new distribution surface beyond terminal-only agent work?
Discovery state
last verified: 2026-05-07 / manual web / medium confidence
- Which OpenClaw release surface should be treated as canonical if docs and GitHub move at different speeds?
- Which user-facing gateway surfaces are official product posture rather than experimental examples?
- Which security and authority boundaries are visible enough for everyday users to understand?
paperclip / active / tier 1 / daily
Paperclip / Paperclip
Watch Paperclip as the coordination and economic-control-plane source. It poses the control-plane question: can agent work be organized into goals, roles, budgets, accountability, approvals, and operating state without becoming theater?
Primary surfaces
-
watch: positioning / onboarding / governance / company model / pricing / demos
-
watch: setup / goals / agents / teams / governance / budgets / accountability / approvals / integrations / security
-
watch: releases / tags / commits / pull requests / issues / docs / examples / security
Accepts as evidence
- official docs
- official site
- github release
- tagged release
- maintainer commit
- merged pr
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- governance change
- study
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
Research lenses
- coordination control plane
- economic governance
- accountability
- multi agent operations
- operating state legibility
High-signal patterns
company / org chart / goal / budget / role / manager / employee / approval / governance / accountability / cost / task queue / progress / audit / session / dashboard / multi-agent / agent team
Operator questions
- Does this make multi-agent labor more legible as operating state?
- Does this expose goals, budgets, roles, approvals, and accountability in a way a human can govern?
- Does this produce real control over agent work or only a company-themed dashboard?
- Does this change how a control plane should model run contracts, allocation, or budgets for agent work?
- Does this make coordination easier without hiding who approved what and why?
Discovery state
last verified: 2026-05-07 / manual web / medium confidence
- Which source is canonical for product changes if the public site, docs, and GitHub repository diverge?
- How much of the company/control-plane metaphor is backed by durable operating state versus UI framing?
- Which governance and budget primitives are enforceable rather than descriptive?
agent-zero / active / tier 1 / daily
Agent Zero / agent0ai
Watch Agent Zero as the workcell-autonomy source. It raises the question of what happens when an agent gets a real computer environment, can use terminal/browser/files, and can grow tools or subagents inside that environment. Pay special attention to isolation, persistence, cleanup, visibility, and whether power remains governable.
Primary surfaces
-
watch: positioning / installation / ui / features / pricing / deployment
-
watch: installation / configuration / tools / code execution / browser / memory / subagents / custom tools / docker / security / remote access
-
watch: releases / tags / commits / pull requests / issues / docs / examples / security
Accepts as evidence
- official docs
- official site
- github release
- tagged release
- maintainer commit
- merged pr
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- runtime change
- test
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
Research lenses
- workcell autonomy
- computer use
- runtime isolation
- tool creation
- visible autonomy
High-signal patterns
Linux / terminal / file system / browser / code execution / Docker / container / tool creation / plugin / custom tool / subagent / memory / task / project / remote access / UI / safety / sandbox / persistence / cleanup
Operator questions
- What does full computer access let the agent do that a narrow tool loop cannot?
- How is the environment isolated, inspected, persisted, reset, or cleaned up?
- What state is visible to the human while the agent acts?
- What happens when the agent creates tools or subagents during a long run?
- Does this change how operators should model containers, workcells, or full machines?
Discovery state
last verified: 2026-05-07 / manual web / high confidence
- Which release or docs surface best describes the current runtime isolation model?
- Which parts of Agent Zero's tool creation are safe to compare against an operator's own tool and receipt boundaries?
- Which behaviors should operators test locally versus only study as product posture?
openhands / active / tier 1 / daily
OpenHands / OpenHands
Watch OpenHands as the productized software-agent platform source. What it signals is breadth: SDK, CLI, GUI, cloud, enterprise, integrations, sandboxing, collaboration, and evaluation in one system. Study what a full platform makes easier, and where operators are better served by a thin control layer than by adopting the whole platform.
Primary surfaces
-
watch: positioning / cloud / enterprise / integrations / pricing / deployment
-
watch: installation / sdk / cli / gui / cloud / enterprise / integrations / sandboxing / security / evaluation / configuration / runtime
-
watch: releases / tags / commits / pull requests / issues / docs / examples / security
-
watch: release notes / breaking changes / migration notes / security
Accepts as evidence
- official docs
- official site
- github release
- tagged release
- maintainer commit
- merged pr
- security advisory
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- enterprise change
- study
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
Research lenses
- productized agent platform
- sandboxed development
- cli gui cloud surface
- enterprise governance
- evaluation
High-signal patterns
SDK / CLI / GUI / cloud / enterprise / self-hosting / sandbox / runtime / browser / evaluation / benchmark / security / RBAC / permission / collaboration / Slack / Jira / Linear / GitHub / extension / integration / multi-user
Operator questions
- What happens when an agent harness becomes a full software-development platform?
- Which surfaces are SDK, CLI, GUI, cloud, enterprise, or integration-specific?
- How does OpenHands package sandboxing, permissions, collaboration, and evaluation for real teams?
- Which parts should teams wrap, which parts should they learn from, and which parts should they refuse to adopt wholesale?
- Does this make agentic software work easier to adopt without hiding evidence or authority?
Discovery state
last verified: 2026-05-07 / manual web / high confidence
- Which OpenHands surfaces should be treated as one product versus separate SDK, CLI, cloud, and enterprise sources?
- Which evaluation and sandboxing claims can be probed locally?
- Which integrations change operator behavior enough to become signals?
heypi / active / tier 1 / weekly
heypi / Ronan Berder (hunvreus)
heypi is the governance-shell calibration source: the human-in-the-loop and audit layer that wraps a minimal coding harness (Pi) for team chat-ops. Watch it as the inverse of Pi -- where Pi refuses to bake governance into its core, heypi's entire product is that governance shell (approvals with named approvers, an SQLite audit trail, sandboxed tools, encrypted secret handoff, scoped memory). The high-signal question is always enforcement vs surfacing: does an approval block the call, is the audit trail complete, does the sandbox isolate. Pay attention to the durability disclaimer (heypi explicitly does not replay in-flight turns after a crash) -- it marks the boundary of what the operator must own. heypi is also a multiplayer/team contrast to OpenClaw's single-user gateway, and a governance-first contrast to eve's durability-first framing. Do not over-weight landing-page feature copy; require a doc or commit before treating a governance feature as real and enforced. Handling: stay Tier 1, cadence weekly as of 2026-08-21. Last shippable event at the 2026-08-20 brief was still 0.3.0-beta.2 (2026-07-22). Daily harvest of a frozen beta was the mismatch, not the slot. Return to daily when a tag moves.
Primary surfaces
-
watch: tags / commits / pull requests / releases / docs / examples / security
-
watch: getting started / concepts / adapters / approvals / audit / sandboxing / secrets / memory / skills / jobs / admin / cli / security
-
watch: positioning / feature list / differentiation
Accepts as evidence
- official docs
- tagged release
- github release
- maintainer commit
- merged pr
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- marketing claim without doc or code
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- governance change
- test
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
- accessibility
Research lenses
- governance
- accessibility
- capability
- coordination control plane
- productized agent platform
High-signal patterns
approval / named approver / audit trail / SQLite / sandbox / just-bash / Docker / Gondolin / secret / memory scope / channel scope / skill / job / cron / heartbeat / admin panel / adapter / Slack / Discord / Telegram / webhook / Pi extension / durability / replay
Operator questions
- Does an approval actually block the tool call, or only surface a notification after the fact?
- Is the audit trail complete enough to reconstruct who approved what, and is it tamper-evident?
- What does a sandbox runtime (just-bash, Docker, Gondolin) actually isolate, and what leaks across it?
- How are secrets handed to a tool without passing through chat, and where do they rest?
- In a multiplayer channel, whose authority binds -- the requester's, the approver's, or the agent's config?
- Does heypi inherit Pi's "no governance in core" posture, or does it genuinely close that gap?
- What durability does heypi NOT provide (it disclaims crash-replay), and what does an operator have to own themselves?
Discovery state
last verified: 2026-06-24 / manual web / medium confidence
- heypi publishes git tags but no GitHub "releases"; is the tag the canonical ship signal, or is npm publish the real channel?
- Which governance surfaces (approvals, audit, sandbox, secrets) are on by default vs opt-in, and which bind at runtime vs surface only?
- Which version line is "stable" once 0.2.0 leaves beta, and what is the support/upgrade contract?
- What exactly does the Gondolin sandbox runtime provide relative to just-bash and Docker?
deepseek-harness / active / tier 1 / daily
DeepSeek Harness / deepseek-ai
Added 2026-08-17. deepseek-ai/deepseek-harness, MIT, TypeScript, repo created 2026-08-13. Its own words: "an open-source agent harness developed by DeepSeek AI" using "an architecture where everything is a plugin," powered by Cordis.
It earns a Tier 1 slot on class, not on velocity. This is a frontier model lab shipping its own harness, which puts it in the same bracket as Codex, Claude Code, and Gemini CLI: the lab that trains the model also decides what the agent may do with a machine. Those are the sources where the publication's standing questions land hardest, and the class alone justifies daily cadence before a single finding is written.
The intake condition is the story and should be checked every window, not assumed. Four days after the repo appeared there is exactly one tag, dsh-v0.1.0-rc.7, marked prerelease, and the npm latest dist-tag resolves to that same rc. The README's own install command, npx @deepseek-ai/dsh web, therefore installs a release candidate, and the README warns in capitals that compatibility-breaking changes are coming. There is no stable channel. Under this publication's rule that released is not merged and that a finding must name the channel an operator can actually run, the honest statement for now is that nothing here has shipped to a channel an operator should depend on. Attention is not a channel, and the star count is not evidence of adoption; it is listed under rejected evidence for that reason.
Two structural facts change how this source must be harvested. Issues are disabled, so defects and their resolutions live in Discussions, which is a weaker and less durable receipt surface than an issue tracker; capture the permalink and the text. And a .gitlab-ci.yml is committed, which suggests the public repo may be a mirror of an internal one. If it is, the commit history here is a partial record and absence of a commit is not evidence of absence of work. Establish which it is before treating a gap as a finding.
Handling: everything-is-a-plugin is an architectural claim, and this publication tests claims of that shape against the code rather than repeating them. The useful probe is not whether a component can be swapped but whether the component that enforces a limit can be swapped by the thing it limits. Where behaviour traces to Cordis, say so; that is a fact about the pair, in the same way an Omnigent finding is a fact about the pair.
Primary surfaces
-
watch: releases / tags / commits / pull requests / breaking changes / security / default branch master
-
watch: prerelease flag / first stable tag / new features / breaking changes / plugin api changes / deprecations
-
watch: dist tags / latest points at prerelease / package version / publish cadence
-
watch: architecture / agent lifecycle / capability seams / config catalog / api gateway / plugin contract / sandboxes / permissions / user guide / i18n divergence
-
watch: published advisories / severity / affected ranges
-
watch: maintainer answers / bug reports / breaking change notices / roadmap statements
-
watch: third party plugins / plugin permissions / supply chain
-
watch: positioning / feature list / preview status
Related surfaces
watched
Harness is built on Cordis, a third-party composability framework, so the plugin boundary this harness sells is defined in someone else's repo. A change to Cordis can change what a dsh plugin may do without any commit landing in deepseek-harness. Attribute accordingly: a finding about the plugin contract must say which repo the behaviour comes from.
Accepts as evidence
- official docs
- tagged release
- github release
- maintainer commit
- merged pr
- package registry release
- maintainer authored post
- maintainer discussion answer
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- marketing claim without doc or code
- star count as adoption
Default actionability
- release
- observe
- prerelease
- observe
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- first stable release
- test
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
Research lenses
- provider first party harness
- plugin architecture
- capability
- security
- ecosystem
High-signal patterns
plugin / dsh-plugin / cordis / capability seam / sandbox / permission / approval / api gateway / agent lifecycle / session / storage / scheduling / loop / web ui / rc / developer preview / breaking change / stable
Operator questions
- The only tag is a release candidate and the npm latest dist-tag points at it: is there any channel an operator can install that is not a prerelease?
- The README states in capitals that there will be compatibility-breaking changes. Which surfaces are covered by that warning, and is the plugin API among them?
- If everything is a plugin, what enforces a plugin's limits, and can a plugin replace the component that would have refused it?
- The Web UI binds 127.0.0.1:3080 by default. What authenticates a request to it, and what happens when an operator exposes that port?
- Issues are disabled and a .gitlab-ci.yml is committed. Is the public GitHub repo the development home or a mirror, and where does an operator report and track a defect?
- Docs ship English and Chinese side by side with i18n yaml. When the two disagree, which one is normative?
- Cordis defines the plugin paradigm. Which behaviours does an operator inherit from Cordis rather than from DeepSeek, and who ships a fix when one of them breaks?
- Does the harness run any model other than DeepSeek's own, and is that a supported path or an incidental one?
Discovery state
last verified: 2026-08-17 / manual web / medium confidence
- What is the canonical ship signal, the GitHub release, the tag, or the npm publish? All three currently point at the same rc.
- Is there a published stabilisation target or date for 0.1.0 proper?
- Which parts of the architecture are plugin boundaries in fact rather than in the README's claim?
- Does an internal GitLab pipeline gate what lands here, and does that make the public commit history a partial record?
flue / active / tier 2 / weekly
Flue / withastro
Watch Flue as the programmable harness / headless agent calibration source. Its core framing, "Agent = Model + Harness," explicitly separates the model from the harness, filesystem, sandbox, skills, memory, sessions, and deployment surface. That is evidence for the thesis that the valuable layer is the shaped environment around the model, not just the model call itself. Treat it as category evidence and possible integration reference, not stable infrastructure. APIs are self-described as experimental; monitor direction before treating any primitive as architectural precedent.
Primary surfaces
-
watch: commits / releases / tags / pull requests / readme / changelog / examples / docs
-
watch: versions / breaking changes / new features / fixes
Canonical receipt surface. Flue publishes no GitHub Releases (github.com/withastro/flue/releases is empty), so cite the version-tagged CHANGELOG.md (or a release tag) as the receipt, never the /releases page. Surfaced by the 2026-05-28_2026-06-03 run audit (item 2).
-
watch: framing / feature surface / deployment targets / skill system / sandbox api
Accepts as evidence
- github commit
- github release
- tagged release
- merged pr
- readme change
- official docs
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- seo clone or mirror
Default actionability
- release
- observe
- docs change
- observe
- api change
- study
- breaking change
- note
- ecosystem package
- observe
- philosophy change
- study
Change types
- capability
- api surface
- runtime
- sandbox
- skill system
- deployment
- protocol
- philosophy
- ecosystem
- security
- breaking change
Research lenses
- agent harness architecture
- model harness separation
- programmable runtime
- headless agent
- sandbox design
- skill primitives
- session memory
- ci deployment
- filesystem abstraction
High-signal patterns
model + harness separation / programmable harness / headless agent / sandboxed execution / skill system / markdown skills / AGENTS.md / session management / memory / filesystem abstraction / HTTP server / CLI agent / CI/CD deployment / Cloudflare Workers / API change / breaking change / experimental
Operator questions
- Does this validate or challenge the "model + harness" framing as a durable public category?
- Which harness primitives (sandbox, filesystem, skills, memory, sessions, credentials) are stabilizing vs. still experimental?
- Does this suggest integration surfaces teams should expose, wrap, or treat as precedent?
- Does any API change affect how operators should think about their own receipt layer or deployment membrane?
- Is the project gaining enough traction to treat as architectural reference rather than just category evidence?
Discovery state
last verified: 2026-06-03 / harvest run / medium confidence
- Confirm GitHub repo is github.com/withastro/flue (withastro org is unusual for an agent harness project; verify ownership).
- Is the Apache-2.0 license confirmed in the repo?
- What is the actual star count and commit velocity at time of first harvest?
- Are APIs stable enough to treat individual primitives as architectural precedent, or watch-only for now?
eve / active / tier 2 / weekly
eve / Vercel
Watch Eve as the filesystem-first, durable-execution agent framework on the authority axis. Its model, "an agent is a directory of files" (instructions, tools, skills, channels, schedules, subagents, connections, sandbox, hooks), states a portable, reviewable definition of an agent in public from a major vendor. Two lenses make it frontier-relevant: durable, resumable, crash-safe execution built on the open-source Workflow SDK, and human-in-the-loop approval gates that let an operator approve or deny a tool call before the agent proceeds. Treat it as category evidence and an authority-gate reference, not stable infrastructure: it is a fast-moving public beta under Vercel beta terms, with breaking changes arriving on minor versions. It is a general agent framework, watched as a harness on the coding and agent frontier, not a coding-only tool.
Primary surfaces
-
-
watch: commits / releases / tags / pull requests / readme / changelog / examples
-
watch: versions / breaking changes / new features / fixes
Canonical receipt surface. Eve ships tagged GitHub Releases with per-release notes (e.g. eve@0.10.0, eve@0.11.5). Cite the release tag and its notes as the receipt for a version-level change; cite the docs for architectural claims and the Vercel changelog for the launch itself.
-
watch: framing / project layout / tools / skills / channels / schedules / subagents / connections / sandbox / hooks / durable execution
-
watch: launch framing / positioning
Accepts as evidence
- github release
- tagged release
- github commit
- merged pr
- readme change
- official docs
- maintainer authored post
- vendor changelog
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- seo clone or mirror
Default actionability
- release
- observe
- docs change
- observe
- api change
- study
- breaking change
- note
- approval gate change
- study
- sandbox change
- study
- philosophy change
- study
Change types
- capability
- api surface
- runtime
- sandbox
- skill system
- subagents
- connections
- durable execution
- approval gates
- deployment
- protocol
- philosophy
- ecosystem
- security
- breaking change
Research lenses
- agent harness architecture
- filesystem first agents
- durable execution
- human in the loop approval
- sandbox design
- subagent delegation
- mcp connections
- skill primitives
- session recovery
- gateway credentials
High-signal patterns
filesystem-first agent / agent is a directory of files / durable execution / resumable sessions / crash-safe / human in the loop / approval gate / tool call denied / sandbox backend / Microsandbox / Docker sandbox / subagent delegation / MCP connection / skills / schedules / channels / AI Gateway / OIDC / breaking change / public beta
Operator questions
- Does Eve's approval-gate model change how teams gate tool calls before an agent proceeds?
- Which primitives (sandbox backends, durable sessions, subagents, connections, approval gates) are stabilizing versus still moving on a fast beta?
- Does the filesystem-first "agent is a directory of files" model make agent definitions more portable, reviewable, or version-controllable than code-first frameworks?
- How does durable, resumable execution on the Workflow SDK change what operators can promise about crash recovery and human pauses mid-run?
- Is the project stable enough to treat individual primitives as architectural precedent, or watch-only while the beta churns?
Discovery state
last verified: 2026-06-19 / harvest run / medium confidence
- How stable is the public API across the rapid 0.10..0.11 release cadence (eight releases in the first days)?
- Which of the three sandbox backends (Vercel, Microsandbox, Docker) are first-class versus best-effort in practice?
- What does the human-in-the-loop approval surface look like end to end (who approves, where the pause is recorded, what an operator sees)?
- Does the Workflow SDK dependency constrain where Eve agents can be hosted and recovered?
agent-flywheel / active / tier 2 / weekly
Agent Flywheel / Jeffrey Emanuel (Dicklesworthstone)
Watch Agent Flywheel as both an assembly layer and an operating method. It makes provider agents replaceable inside durable planning, task, coordination, verification, and memory artifacts, then distributes the supporting toolchain through one installer. That is unusually close to Frontier's Bitter Lesson and Amdahl questions. Weekly evidence remains bounded to ACFS releases, tagged docs, and attributed current claims on the official site. A dated ecosystem study may inspect the named related repositories and contextual paper with pinned receipts when their interaction changes the read.
Primary surfaces
-
watch: releases / tags / readme / docs
Weekly change detection is release- and tag-bounded -- NOT commits or pull requests. README and docs claims are accepted only when pinned to a tag or its dereferenced commit. The repo had 3,440 main-branch commits at the 2026-07-02 intake against 7 tagged releases; the tag is the receipt surface that keeps this source finite. Never cite a moving main URL for a historical posture claim.
-
watch: methodology / cost transparency / agent lineup / installer claims / security posture
Related surfaces
context only not weekly harvest
A selected set of repositories that implement the operating loop may be studied in explicitly scoped, dated work. This list is not the maintainer's complete portfolio, and these repositories are not silently promoted into weekly Agent Flywheel findings.
Accepts as evidence
- tagged release
- github release
- tagged readme or docs
- attributed official site claim for current posture
- attributed project authored budget
- contextual primary research paper
- maintainer authored post with primary receipt
- reproducible local probe
Refuses to promote
- unspecified portfolio wide claim
- untagged main branch commits
- unsourced social claim
- third party summary without primary link
- self reported metric presented as independent fact
- speculation
- stale model memory
- duplicate commentary
- seo clone or mirror
Default actionability
- release
- study
- docs change
- observe
- defaults change
- note
- security
- note
- cost change
- observe
- ecosystem package
- ignore
Change types
- defaults
- security
- capability
- cost
- distribution
- agent lineup
- breaking change
- ecosystem
- license
Research lenses
- assembly layer
- default setting authority
- installer trust
- cost of operation
- multi agent cohabitation
- individual account repository output
- durable state outside provider harnesses
- human liaison work
- agent native interfaces
- memory and feedback loops
High-signal patterns
install.sh behavior change / default settings written for Claude Code / Codex / Antigravity / permission or auto-mode posture set by the installer / credential, token, or session handling / version or channel pins for the bundled agents / agent added to or dropped from the lineup / sudo or system-level configuration change / safety-tool enforcement vs. advice / published cost figures updated
Operator questions
- What defaults does the installer set across the three bundled tier-1 agents, and would their own vendors ship those defaults?
- Does the installer pin agent versions and channels? An assembly layer inherits the released-is-not-merged problem for every tool it bundles.
- Where do credentials and session state land on the VPS, and who can read them?
- Do the bundled "safety tools" enforce anything, or advise? A warning is not a boundary.
- Are the project's attributed operating-budget examples ($440-656/month at intake, including a two-Claude-account high end) holding as the lineup changes?
- Does the plan-to-graph-to-coordination loop reduce total human attention across a completed project, or create new maintenance and review queues?
- Does the OpenAI/Anthropic license rider apply to the operator or intended use? Potentially covered parties should review the tagged LICENSE and obtain their own legal guidance.
Discovery state
last verified: 2026-07-12 / exemplar and receipt audit / high confidence
- Does the next tag make safe mode gate the dangerous Claude, Codex, and Antigravity shortcuts, remove ACFS-created NOPASSWD state when changing modes, and detect provider-supplied passwordless sudo?
- Which bundled agent versions and release channels are pinned, and which still float with upstream latest installs?
- What measured evidence would show that the coordination and memory loop improves verified outcomes per unit of human attention?
- Will the non-standard OpenAI/Anthropic license rider change, and which operators need to obtain separate permission?
omnigent / active / tier 2 / weekly
Omnigent / omnigent-ai
Added 2026-08-02 as the first meta-harness on the watchlist. Everything else here IS a harness; Omnigent orchestrates them, describing itself as "an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents" and shipping policies, spend caps and access controls on top of harnesses that already have their own.
That stacking is why it is worth tracking. This publication's standing argument is that a control existing only as an intention is not a control, and a meta-harness is the hardest test of it: two governance layers now have a claim on the same action, and an operator has to know which one refuses. That question has no public answer yet, which makes it the most interesting open item on the list.
Release cadence at intake was roughly weekly, v0.2.0 through v0.7.0 between 2026-06-19 and 2026-07-27, pre-1.0, so the tag-to-tag diff carries more than the release note. Two items in v0.7.0 deserve a second look: optional server-side transcription is a new data path off the operator's machine, and a router that selects the harness means the governance layer an action lands under can change without the operator choosing it.
Handling: this is Tier 2 and weekly - do not promote on release velocity alone. A finding observed through Omnigent is a finding about the pair, not about the wrapped harness, and must say so. Adapter lag is a legitimate finding but is not a defect in the harness underneath. Closest comparison is Paperclip, and the contrast is the useful part: Paperclip manages an org of agents it owns, Omnigent orchestrates agents it does not.
Primary surfaces
-
watch: releases / tags / commits / pull requests / breaking changes / security
-
watch: new features / policy changes / routing changes / breaking changes / deprecations
-
watch: unreleased section / fixes / breaking changes
-
watch: published advisories / severity / affected ranges
-
watch: policies / permissions / sandboxing / spend caps / harness adapters / routing / sessions / teams / deployment
-
watch: positioning / feature list / differentiation
Related surfaces
watched
Omnigent drives harnesses this publication already tracks separately. Behaviour observed through it is a fact about the pair, never about either component alone, and a finding must say which.
Accepts as evidence
- official docs
- tagged release
- github release
- maintainer commit
- merged pr
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- marketing claim without doc or code
Default actionability
- release
- test
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- governance change
- test
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
Research lenses
- governance
- coordination control plane
- productized agent platform
- capability
- economics
High-signal patterns
policy / deny / allow / approval / spend cap / budget / quota / access control / adapter / harness / router / routing / sandbox / isolation / session sync / transcription / scheduled task / automation / project / team / breaking change
Operator questions
- When an Omnigent policy and the underlying harness's own permission system disagree, which one refuses, and is that documented or only observed?
- Are spend caps enforced before a call is made, or reconciled after it? A cap that reconciles is a report, not a control.
- When a wrapped harness changes a permission surface, how long does the adapter lag, and what does an operator's policy mean during the gap?
- The v0.7.0 router picks both harness and model. Can the governance layer an action lands under change without the operator choosing it?
- Server-side transcription is a data path off the operator's machine. What is sent, what is retained, and is it opt-in?
- Is the governance work in the tag an operator installs, or on main?
- Pre-1.0 and shipping weekly: what is the upgrade contract, and what breaks between minors?
Discovery state
last verified: 2026-08-02 / manual web / medium confidence
- Which policy decisions are enforced in Omnigent's process versus delegated to the wrapped harness?
- Does the adapter surface a harness's own refusal to the operator, or swallow it?
- What is the canonical ship signal - the GitHub release, the tag, or a package publish?
- Which of the governance features in the landing copy are on by default?
omp / active / tier 2 / weekly
OMP (Oh My Pi) / can1357
Added 2026-08-17. can1357/oh-my-pi, MIT, TypeScript with a Rust core, repo created 2025-12-31. Its own words: "A coding agent with the IDE wired in," and in the README, "Fork of Pi by @mariozechner." Landing copy claims 60+ providers, 31 built-in tools, 14 LSP operations, 28 DAP operations, and roughly 80k lines of Rust.
It earns a slot for the relationship, not the feature list. Pi is already tracked here, and OMP is what happens when a fork outruns the thing it forked. Upstream Pi publishes 0.84.2; OMP publishes 17.3.5 under a package with the same basename and a different scope. Two projects, one name, version numbers seventeen majors apart, and an operator reading a version string alone cannot tell which one they are running or which is newer in any meaningful sense. This publication spent the 2026-08-03 issue on a release line renumbered below its own predecessor. This is the same failure mode from the other direction, and it is why the fork belongs on the list rather than in the adjacent index.
The channel picture at intake is worth re-checking every window because it was already split on day one. Tags v17.3.6 and v17.3.7 exist with no GitHub release behind them, the newest GitHub release is v17.3.5, and npm latest is 17.3.5. So the tag series runs ahead of every channel an operator installs from, and there are four such channels: an install script piped to a shell, a Homebrew tap, a global Bun install, and a Nix flake. Name the one you tested.
The capability surface is the reason to read it closely. LSP and DAP mean this agent drives a language server and a live debugger rather than only editing text, which is a genuine widening of what a terminal agent operates, and a correspondingly larger question about what confines it. Hindsight memory and time-traveling rules are behaviour loaded from stored state, which is the shape of defect this publication has repeatedly found where a guard consults something the workspace can write.
Handling: Tier 2 and weekly. Do not promote on release velocity, and do not promote a feature merely because upstream Pi lacks it. Single-maintainer projects with 1,498 open issues and PRs open to everyone on a stated trial are a maintainership story only when something concrete turns on it. Attribute carefully in both directions: OMP inherits Pi code, so a defect found here may or may not be Pi's, and the finding must say which was tested.
Primary surfaces
-
-
watch: releases / tags / commits / pull requests / breaking changes / security / contribution policy
-
watch: new features / breaking changes / tag without release / deprecations / upstream divergence
-
watch: unreleased section / fixes / breaking changes
-
watch: dist tags / package version / version lag behind tags / scope collision with upstream
-
watch: published advisories / severity / affected ranges
-
watch: lsp / dap / subagents / plan mode / hindsight memory / hashline edits / rules / permissions / sandboxing / providers / relay sessions / extensions
-
watch: resolved version / platform matrix / integrity verification
-
watch: public framing / divergence from repo changelog
Related surfaces
watched
OMP states in its README that it is a fork of Pi by Mario Zechner, and Pi is already on this watchlist as pi-coding-agent. Findings must not be laundered between them. A behaviour observed in OMP is a fact about OMP unless the code path is shown to be shared, and an upstream Pi change is not an OMP change until it lands in a tag OMP publishes. Note also that the README links badlogic/pi-mono, the pre-migration upstream, while Pi's canonical repo is now earendil-works/pi.
Accepts as evidence
- official docs
- tagged release
- github release
- maintainer commit
- merged pr
- package registry release
- maintainer authored post
- reproducible local probe
Refuses to promote
- unsourced social claim
- third party summary without primary link
- speculation
- stale model memory
- benchmark claim without method
- duplicate commentary
- marketing claim without doc or code
- upstream pi behaviour assumed to hold in omp
Default actionability
- release
- observe
- docs change
- observe
- security change
- test
- breaking change
- adapt
- ecosystem package
- observe
- upstream divergence
- observe
Change types
- capability
- workflow
- runtime
- protocol
- reliability
- economics
- security
- ecosystem
- evaluation
- philosophy
Research lenses
- fork divergence
- ide integration
- capability
- distribution channel
- maintainership
High-signal patterns
lsp / dap / debugger / hashline / hindsight / plan mode / subagent / rules / relay / session share / permission / approval / sandbox / provider / rust core / breaking change / vouch / upstream
Operator questions
- The fork publishes @oh-my-pi/pi-coding-agent at 17.x while upstream Pi publishes @earendil-works/pi-coding-agent at 0.8x: does anything an operator runs resolve the wrong package basename by scope confusion?
- Two tags exist with no GitHub release and npm latest sits at the last released tag. Which of the four install paths, script, Homebrew, Bun, and Nix, resolve to the same version on the same day?
- At better than one release a day, what is the upgrade contract, and where is a breaking change announced other than the changelog?
- LSP and DAP mean the agent drives a language server and a live debugger. What confines those, and does a debugger session widen what the agent can execute past whatever the tool permission layer allows?
- Relay sessions put a live session behind a link and a QR code. Who can join, what does a joiner see, and is the relay operated by the maintainer?
- Hindsight memory and time-traveling rules both change agent behaviour from stored state. What writes that state, and can a repository under review write it?
- The PR vouch requirement is lifted "temporarily as a trial." What is the review gate on a single-maintainer project carrying roughly 80k lines of Rust, and what happens when the trial ends?
- Where does an OMP capability actually come from, its own Rust core or inherited Pi code, and does the answer change who fixes it?
Discovery state
last verified: 2026-08-17 / manual web / medium confidence
- Is there a stated compatibility relationship with upstream Pi, or is the fork now fully independent?
- Which surface is canonical for a shipped version, the GitHub release, the tag, the npm dist-tag, or the install script?
- Does the Homebrew tap or the Nix flake ever lead or lag the npm publish?
- Is the relay a maintainer-operated service, and what does it retain?