Risk isn't a property of the verb
I pointed an adversarial review at my own agent runtime and it found two holes. One of them is, I suspect, in most agent permission models shipping today.
Two months ago I published a post arguing that most AI agent frameworks have the safety model backwards — that they optimise for autonomy and bolt safety on afterwards, and that the gate belongs at the execution boundary rather than inside a prompt. MachinaOS was built that way: every tool call is classified by risk, LOW and MEDIUM run, HIGH requires human approval, CRITICAL is refused outright, and everything lands in a hash-chained audit log.
Then I did something I'd recommend to anyone shipping a safety claim: I played the sceptic against my own code. Not "are the tests green" — they were, a few thousand of them — but "if I were a hostile security reviewer with the source in front of me, where would I say this fails?"
It found two holes. The first is embarrassing in hindsight and, I suspect, widespread.
Hole one: the policy engine never looked at the arguments
Here is roughly what the gate did. A step arrives naming a tool. The engine looks up that tool's spec, reads its risk_level, checks the capability grants, and decides.
Notice what isn't in that list: the arguments. The engine evaluated which verb was being called and never what it was being called on.
filesystem.write_file is classified MEDIUM. MEDIUM auto-runs — no approval, no banner, no pause. That is a perfectly reasonable classification for writing src/app.py. It is a catastrophic classification for writing:
~/.ssh/authorized_keys— hands out persistent shell access~/.aws/credentials— swaps in someone else's cloud account.github/workflows/ci.yml— arbitrary code execution on every future push/etc/hosts— silently redirects outbound requests
Every one of those is the same MEDIUM verb. Every one of them would have auto-run. My HIGH tier was carefully guarding shell.run and git.push while an agent could achieve strictly worse outcomes without ever touching a HIGH-risk tool.
This matters more than an ordinary bug because of how agents get attacked. Prompt injection is the live threat: the agent reads a file, a web page, an issue comment containing hostile instructions. A runtime gate is genuinely the right defence — it doesn't care what the model was persuaded to want, only what it tries to do. But that only holds if the gate can see the whole action. Mine saw half of it. Injection that produced loud HIGH-risk steps was stopped; injection that produced quiet MEDIUM steps walked straight through.
The generalisable lesson, and the reason I'm writing this up rather than quietly patching it:
Risk is a property of what an action touches, not of which verb names it.
Verb-level permissions are the model most systems reach for first, because it's the model that's easiest to declare. It's also the model Unix escaped decades ago — write isn't a privilege, write to this particular inode is. Agent frameworks are busy reinventing the coarse version.
The fix: rules that read the arguments
Policy packs — the signed, org-distributable policy files MachinaOS uses — gained argument rules. A rule matches on a tool glob, an argument-name glob, and a regex against the argument value, and escalates the effective risk when it hits:
{
"tools": ["filesystem.*"],
"arg": "path",
"pattern": "\\.ssh/|authorized_keys|/etc/|\\.aws/credentials",
"risk": "critical",
"reason": "writes to credential and system paths are never auto-approved"
}
Four design decisions in there matter more than the feature itself.
The strictest match wins. Rules can only escalate, never relax. A pack author cannot accidentally widen what runs unattended by adding a rule.
CRITICAL blocks even for an unknown tool. If a rule matches, the block happens whether or not the runtime recognises the tool being called — so a newly registered or dynamically discovered tool can't slip past on a technicality.
Invalid regexes are rejected when the pack loads, not when it evaluates. This one is the difference between a security control and a decoration. If a malformed pattern were only discovered at evaluation time, the honest options would be to crash mid-execution or to skip the rule — and skipping means failing open: the rule you most needed silently stops protecting you. Validating at load time makes a broken pack a startup failure instead, which is loud, immediate and safe.
No false positives on ordinary work. Writing src/app.py still runs unattended. That isn't politeness. A governance layer that interrupts constantly gets clicked through reflexively within a week, and then you have approval theatre rather than approval. A gate is only as good as the attention it hasn't already exhausted.
Hole two: nothing recorded who asked
The second finding was shorter to state and more awkward. The audit log faithfully recorded who resolved every approval — actor, source IP, timestamp, all hash-chained. It never recorded who requested it.
Which means an engineer could raise a HIGH-risk step and approve it themselves. Segregation of duties — the principle that the person who asks and the person who authorises must differ — is table stakes in any regulated environment, and it was structurally impossible to enforce here, because the data needed to check it didn't exist.
This illustrates something broader about audit design: a log that records outcomes but not provenance can tell you what happened and not whether it was legitimate. The missing field wasn't a feature gap so much as a blind spot in what I had thought was worth writing down.
The fix has three parts. Approvals now carry a requested_by field. That field is populated from the identity of whoever drove the request, propagated through the async call stack in a context variable so concurrent requests can never see each other's identity. A policy pack can then set forbid_self_approval, at which point an actor approving their own request is refused — and the refusal is itself audited, because a blocked attempt is exactly the event a reviewer wants to find later.
Crucially, that check is off by default. Someone running MachinaOS on their own laptop is legitimately both the requester and the approver; enabling segregation of duties globally would lock them out of their own machine. This is the recurring tension in building governance for a local-first tool: the same control that is mandatory for a bank is nonsense for an individual. The answer isn't to pick one — it's to make the control real, make it org-configurable, and leave the single-user path untouched.
Both settings live inside the pack's signed payload, so an attacker cannot quietly disable segregation of duties or delete an argument rule without invalidating the signature.
What else went into 1.1
The same release added the two interfaces organisations ask for first, both built against public standards rather than one company's setup.
OIDC / SSO. OpenID Connect tokens from Entra ID, Okta, Google, Keycloak or Auth0 can now be used as API credentials instead of hand-managed static tokens, with real signature verification against the provider's JWKS and issuer, audience and expiry all enforced. The authenticated subject flows into approvals and the audit log like any other identity. One deliberate decision: configuring SSO forces authentication on. It is otherwise far too easy to wire up an identity provider, forget the separate "require auth" switch, and ship an open API in the belief that SSO is protecting it.
Audit export to SIEM formats. The audit log can be rendered as CEF (ArcSight, QRadar, Sentinel), ECS (Elastic) or Splunk HEC envelopes, so agent activity lands next to everything else a security team already watches instead of sitting in a silo. Every exported record carries its chain hash, so a reviewer can tie any event in their SIEM back to the tamper-evident local log — and notice if the two ever disagree.
And one small change with a large blast radius: the environment flag that exposes HIGH-risk tools to external MCP clients can no longer be silently in effect. It warns loudly at startup, is reported by the status endpoint, and raises a banner in the UI. A dangerous setting that produces no visible signal is a setting you will eventually forget you set.
Why publish the holes
There's an obvious argument against this post: I'm making a governance claim, and this is a public account of my governance having gaps.
I think the opposite holds. Any system claiming to make AI agents safe is asserting something a reader should be sceptical of, and "here are the specific holes we found in our own model, and what we did about them" is far better evidence of a working safety process than a feature list. A governance product that has never published a weakness mostly tells you nobody has looked hard.
The deeper lesson is about where to aim scrutiny. Both findings sat in code that was thoroughly tested. The tests confirmed the engine did what it was designed to do; they could not tell me the design was reasoning about the wrong thing. That gap — between a correct implementation and a correct model — is not one a test suite closes. It takes somebody deliberately trying to get past it.
MachinaOS 1.1.0 is out, and the live demo runs the current build. If you're building agent infrastructure and your permission model classifies by verb, it's worth an hour this week asking what your equivalent of write_file can actually reach.