Profiles / omnigent-ai

Omnigent

The first meta-harness on this watchlist: a governance layer over the coding agents it drives, whose spend cap is enforced before the call and is a downgrade gate rather than a ceiling.

Edited by Michael Ruescher / reviewed 2026-08-20

Every other project on this watchlist is a harness. Omnigent sits above them: an open-source meta-harness that orchestrates Claude Code, Codex, Cursor and Pi, shipping policies, spend caps and access controls on top of harnesses that already carry their own permission systems.

That stacking is why it is here. This publication’s standing argument is that a control which exists only as an intention is not a control, and a meta-harness is the hardest version of that test, because two governance layers now have a claim on the same action.

Where it stands, 2026-08-20

Channel. v0.10.0 published 2026-08-19T04:34:41Z. CHANGELOG.md at that tag still starts at v0.9.0; the v0.10.0 notes live on the GitHub release body.

What is new. Several sandbox providers at once, Devin as a built-in harness, a Usage page, Copilot via gh auth login. The Usage page is off unless OMNIGENT_FEATURES=usage_page. Unset or empty means every release feature is off (feature_flags.py at v0.10.0).

What did not move. cost.py is blob 5b4ca596 at both v0.9.0 and v0.10.0. max_cost_usd is still a downgrade gate. Omitting expensive_models or setting [] is the hard stop. Shared-editor approval is still any-editor. qwen/goose delegated file I/O still fails open on the result phase; a write-result denial does not undo the write (qwen_executor.py at v0.10.0). Devin is more likely to sit on the generic ACP path, which does not get that content gate. Do not attribute Devin behavior observed through Omnigent to Devin alone.

Tag-push deny is in the tag. Two issues ago deny_tag_push (default true) missed v0.9.0 and lived on nightly. github.py at v0.10.0 has it. git push --tags, --follow-tags, and refs/tags/ refspecs are denied unless you set deny_tag_push: false.

Where it stood, 2026-08-03

Channel. Pre-1.0 and shipping continuously. v0.7.0 was published 2026-07-27T22:40Z and is the newest tag; more than a hundred commits landed on the default branch in the following week. The tag-to-tag diff carries far more than the release note, so anything read here is read at a tag ref rather than on main unless stated.

Its spend cap enforces before the call, and is not a ceiling. Read at v0.7.0, cost_budget gates cumulative session spend at the request phase -- “before the LLM turn, so text-only turns are budgeted too” -- and at the tool-call phase, “the point a native PreToolUse hook can block before the action runs.” That settles the question the source contract opened with: it enforces rather than reconciles.

The catch is what max_cost_usd does when reached. It “forces a model downgrade” rather than stopping the session, denying only while the session runs a model in the operator-supplied expensive_models list, and the module says so directly: “the budget becomes a ‘downgrade gate,’ not a hard stop.” An operator who sets the number expecting spend to end there has configured the point at which the work continues more cheaply.

The gate does close its own worst failure mode. A model with no catalogue pricing never writes a cost to the session, which would score it at zero and let it run unbounded; instead the gate fails closed when token usage is present and priced cost is absent, denying and asking the operator to switch to a priced model. It also notes that a single expensive turn can overshoot between checks.

Its only write confinement for unsandboxed workers did not bind on Windows. worktree_guard reasoned in POSIX terms but normalised with os.path, which is ntpath on Windows and rewrites forward slashes to backslashes, so the absolute-path arm was inert on a Windows runner: /etc/passwd cleared the backslash guard, became \etc\passwd, and returned ALLOW, as did paths into another worker’s tree. Filed 2026-08-01, fixed 2026-08-03 with posixpath normalisation, a drive-letter arm, and four ALLOW-to-DENY cases pinned by tests run on Windows 11. The fix is on main; v0.7.0 predates it.

Its router now picks the harness. v0.7.0 ships an “Auto - smart routing” option that “lets the router pick both harness and model from your task”, and smart routing “activates automatically from your llm:/routing: config (no OMNIGENT_SMART_ROUTING env var)”. On a meta-harness this is the governance question in shipped form: the layer an action lands under can change without the operator choosing it per action.

A vendor claim that checked out. Its official account described stateful policies making dynamic session-context decisions at server, agent and session level, with a Session Risk Score built in. policies/builtins/risk_score.py is present at the v0.7.0 tag. The claim was accurate and shipped.

Operator posture

Use it as the coordination and spend layer it is, and read its policy modules rather than its parameter names. max_cost_usd needs expensive_models beside it to mean anything. If you run unsandboxed implementer worker specs on Windows, confirm the posixpath worktree fix is in the tag you install; it was on main after v0.7.0.

Do not treat a finding observed through Omnigent as a finding about the harness underneath. Adapter lag and policy-layer defects belong to the wrapper.

Open questions

  • When an Omnigent policy and the wrapped harness’s own permission system disagree, which one refuses? Still no public answer, and it remains the most interesting open item on the watchlist.
  • Does the offloadable dictation worker send audio off the operator’s infrastructure? v0.7.0 says audio “never leaves your server” while describing the transcription engine as offloadable to a remote worker.
  • Re-read at v0.10.0: expensive_models None or [] sets block_all_models=True (a hard stop). A non-empty list is the downgrade gate. The remaining question is what happens when the list is non-empty but does not match the running model.
  • Sandboxed Linux agents now trust CA roots under the system capath to reach hosts behind a corporate MITM proxy. The reason is stated; the blast radius is not.

Comparison

Closest to Paperclip, and the contrast is the useful part: Paperclip manages an organisation of agents it owns, Omnigent orchestrates agents it does not. OpenHands and Hermes Agent also position above a single coding loop, but both ship the loop as well. Omnigent is the only source here whose entire product is governance over somebody else’s agent.

Verification

open source commits / evidence floor: official docs / updated 2026-08-20

Source policy: what Frontier watches and accepts as evidence

Edited and maintained by Bitter Frontier.

View source on GitHub