Code Review Is Not an Authority Boundary

AI can generate the implementation. Your architecture still has to decide what that implementation is allowed to do.

I was reading a good post about the walls JavaScript hits and how WebAssembly gets past them, and I got stuck on one of them: running code you did not write and do not trust.

The framing we normally reach for is “how do I stop this extension from misbehaving?” That question already concedes something. It assumes we have handed the extension access to things worth abusing, and that our remaining defense is its good conduct.

There is a different question underneath it. What is this extension actually entitled to have?

I wanted to know whether that distinction survives contact with running code, so I built a small capability-mediated tool host to find out. It is about 200 lines of Python rather than a Wasm runtime, because the thing worth testing is the host interface, and every environment that mediates untrusted code has one. Everything below that reports behaviour is output I actually got, including the part where my own preferred answer turned out to be wrong.

From Trust to Authority

For most of software’s history we have put enormous weight on code review. Someone writes code, someone else reads it, tests run, static analysis checks it, and eventually we decide it is trustworthy enough to merge, deploy, or execute.

That model made sense when software was expensive to produce and the people writing it were known participants in the process. AI changes the economics. Code is cheap to generate now, and an agent can write a function, produce a plugin, or assemble an entire implementation in minutes. The question is less whether we can produce the code and more what that code is entitled to do once we run it.

Those are different problems. Code review, however good, only solves one of them.

This Has a Name, and It Is Older Than Most of Us

None of the architecture here is my invention. What I am describing is capability-based security, and it has a long, well-developed literature.

Dennis and Van Horn described capabilities in 1966. Mark Miller named the principle of least authority and built the E language around object capabilities. Capsicum brought a capability model to FreeBSD. seL4 has a machine-checked proof of its capability enforcement. Deno shipped a permission model where filesystem and network access are grants rather than defaults. WASI Preview 2 and the Component Model are capability-oriented by design, which is exactly why the Wasm conversation keeps circling this.

The core idea is consistent across all of them. Instead of giving code broad access and expecting it to behave, the execution environment grants specific capabilities: this directory, this socket, this operation and not that one. The implementation does not promise to stay inside its authority. It never receives the authority in the first place.

So the architecture is not new. What is new is that we are about to produce vastly more code whose author we cannot interview.

Sandboxing Constrains Execution, Not Authority

A sandbox sounds reassuring. Put untrusted code in an isolated environment and prevent it from reaching the rest of the system.

Suppose I have a perfectly sandboxed module. It cannot escape its runtime and cannot access anything the host does not explicitly expose. Excellent. Now suppose the host exposes:

read_customer_records()
write_customer_records()
delete_customer_records()
send_email()
issue_refund()
export_database()

The sandbox is working exactly as designed, and the module is wildly over-privileged. We have constrained where the code executes without constraining what it is authorized to affect. The host interface has quietly become the real trust boundary.

The Wasm people say as much themselves. All I/O in a module goes through its imports and exports, which means a module’s view of the outside can be virtualized entirely by controlling what those imports are linked to. That is a precise description of the control surface, and also an admission that the surface is where the decisions live.

That is the eighth wall hiding inside the untrusted-code one. Not “can I run code I do not trust,” but “can I make what that code is allowed to do explicit, minimal, enforceable, and inspectable?”

Sandboxing constrains execution. Capability grants constrain authority. You need both, and only one of them is commonly discussed.

Two Agents, One Test Suite

Here is the experiment. Two implementations reconcile invoices against payments and produce a report. agent_a reads invoices and reads payments. agent_b does the same reconciliation and also decides that a shortfall is worth refunding and that someone ought to be told.

Both produce the identical report. The test suite cannot tell them apart:

The test suite both implementations have to satisfy:

  PASS  agent_a produces the expected reconciliation report
  PASS  agent_b produces the expected reconciliation report

Nothing in agent_b is a bug. A reviewer reading it would find defensible code. The refund is small, the alert is reasonable, the logic is sound. It is simply reaching for a great deal more authority to achieve the same output, and no assertion about the output can detect that, because the difference is not in the output.

Now run both against a least-privilege manifest:

agent_a:
  completed
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments

agent_b:
  stopped: issue_refund requires refunds.issue, which was not granted
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments
  DENY   refunds.issue      issue_refund

The runtime caught it. No reviewer was involved, no test asserted anything about refunds, and that DENY line is a durable record of an authority boundary being enforced rather than a claim that somebody looked and did not see a problem.

One caveat on what that manifest actually bounds, before anyone reads the demo as stronger than it is. Miller and Shapiro distinguish permission from authority: permission is what the access graph grants you directly, while authority is what you can cause, including indirectly through other components you are permitted to talk to. My host controls permission. A component granted invoices.read that can also call something else holding payments.write has more authority than its manifest suggests. Real capability systems address this by making the reference graph itself the thing you reason about. Two hundred lines of Python do not.

That is the argument in fourteen lines of output. Why can an invoice-reconciliation program issue a refund? Perhaps today’s implementation does not. Perhaps review confirms it never calls the refund API. Those are observations about one implementation, and the implementation is the part we just made disposable.

Disposable Implementation, Durable Authority

If implementation is cheap, the code becomes increasingly disposable. We can regenerate it, replace it, or ask a different model for an alternative.

Some things cannot be disposable. The contract describing correct behavior is one. The authority granted to the implementation is another. The manifest above survives every regeneration of the component beneath it, which is precisely why it is worth writing down and the code is not.

This is already live in the agent tooling most of us are wiring up. An MCP server hands an agent a set of tools. That tool list is a capability grant, whether or not anyone has written it down as one, and the people building these servers keep arriving at the same requirements independently: per-tool audit records, tenant isolation, dynamic registration. Those are capability-system features reached from the practical end.

“We Tried This. It Became Permission Fatigue.”

Here is the objection I would raise if someone showed me a capability manifest, and it is a strong one.

We have deployed manifest-based capability declaration at planetary scale already. Android permissions. Browser extension manifests. iOS entitlements. The result was not security. It was users tapping Allow on everything, reviewers rubber-stamping manifests they did not read, and developers requesting broad permissions because narrow ones generated support tickets. Capability systems have a long history of being architecturally correct and operationally ignored.

Two things are different here, and neither is optimism about human diligence.

The grantor is not a person. In the consumer model a human decides at install time, under time pressure, with no context and every incentive to proceed. Here a policy decides. Policies do not get fatigued, do not want the app to work right now, and can be versioned, tested, and audited.

The requester is regenerable. This is the part earlier systems could not use. When a shipped application hits a too-narrow permission, the permission widens, because rewriting the app is expensive and the user is waiting. When a generated component hits a too-narrow capability, regenerating the component is cheap. The pressure that historically pushed grants wider now pushes implementations to fit the grant instead.

That inversion is the actual reason this might work now, and it is a consequence of implementation becoming disposable rather than of anyone becoming more disciplined.

The Model Should Not Grant Its Own Permissions

If the same model generates the implementation, proposes its tests, decides which tools it needs, and determines its own permissions, we have built a very cooperative security model. The agent says: here is the code I wrote, here are the tests demonstrating it works, and here are the permissions I determined I require. All three may be reasonable. All three come from the same information lineage.

I have made a version of this argument before about tests, which I described then as not letting the student grade the exam. Authority is the same problem one layer over, and it is the more consequential one. A wrong test produces a wrong belief. A wrong grant produces a wrong consequence.

A stronger system keeps some decisions outside that lineage. The model proposes that it needs read access to invoices. A policy determines that invoice reconciliation permits invoices.read. The runtime grants that capability while leaving payments.write and refunds.issue unavailable. An audit record shows what authority was actually granted at execution.

The model can participate without owning the boundary.

Where My Own Answer Broke

That leaves the hard part. Somebody has to write the policy, and I wanted to know whether the obvious shortcut works.

The shortcut is observation. Run the component, record which capabilities it actually reaches for, and derive a least-privilege manifest from what you saw. It is the standard move, it is roughly what audit-mode tooling does across the industry, and it is what I assumed I would end up recommending.

So I ran it. Discovery mode on agent_a produces exactly what you would hope:

Discovery mode: allow everything, record what was actually reached for.

  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments

Derived manifest:

component: invoice-reconciler
capabilities:
  invoices: {read: true, write: false}
  payments: {read: true, write: false}
  refunds: {issue: false}
  network: {external: false}

Tight, minimal, derived from real behaviour rather than a guess. Then I ran the same component against an input the discovery run never saw: a credit note, which legitimately requires writing an adjusting payment record.

  stopped: write_payment requires payments.write, which was not granted
  ALLOW  invoices.read      read_invoices
  ALLOW  payments.read      read_payments
  DENY   payments.write     write_payment

The component is correct. The manifest is wrong. It forbids a legitimate path because nothing exercised that path on the day we happened to be watching.

This is not a bug in my implementation, and tightening the derivation would not fix it. Observation tells you what a component did, not what it may need. A derived manifest is a lower bound on required authority presented as an upper bound on granted authority, and those are different claims. In production that failure arrives as an outage on the rare path, which is exactly how capability systems earn their reputation as obstacles and then get widened until they mean nothing.

I do not have a clean answer, and I am not the first to fail to have one. Miller, Tulloh and Shapiro named this in 2004: treating security as a separate concern has not bridged the gap between principle and practice, they argue, because it operates without knowledge of what constitutes least authority, and only when requests are made can we determine how much authority is adequate. My credit note is that sentence with a stack trace attached. Twenty years on the shortcut still does not work, which suggests the diagnosis was right rather than that I picked an unlucky example.

What I have is a demonstration that the answer I expected to give is wrong, which is worth more than the recommendation would have been.

The partial signals that survive: a declared interface implies a floor. Denied-capability telemetry tells you where a grant was too tight, if somebody is reading it. And the credit-note path suggests capability sets belong with the specification rather than the trace, because whoever knew credit notes existed knew it before any code ran.

That is the same conclusion I keep reaching from other directions. Once implementation is cheap and verification is automatable, the scarce work is discovering what the definition of correct should contain. Capability discovery is that problem wearing different clothes.

Correctness and Authority Are Orthogonal

The two-agent run makes this concrete. Code can be correct and over-privileged, as agent_b is. Code can be incorrect and tightly constrained. The authority boundary limits one class of consequence without establishing correctness.

A word about that heading, since the paper linked earlier is called “Why Security Is Not a Separable Concern” and appears to say the opposite. Miller’s argument is that you cannot bolt security on afterwards as its own layer, because knowing what least authority means requires knowing what the program is for. Mine is narrower: evidence that an implementation is correct is silent about what that implementation may affect. He is describing where the work belongs in a design. I am describing what a passing test proves. Both land in the same place, which is that somebody who understands the task has to make the authority decision deliberately.

A perfectly implemented image-resizing plugin should not reach payroll records. A buggy image-resizing plugin with no filesystem, network, or unrelated API access may cause considerably less damage than a correct implementation carrying unnecessary privileges.

Verification asks whether an implementation satisfies its behavioral contract. Authority asks what consequences that implementation is permitted to create. Tests do not establish least privilege, and least privilege does not establish correctness. We need evidence for both, and I now have a fourteen-line proof that one kind of evidence is silent about the other.

Code Review Still Matters

None of this makes code review obsolete. Generated code contains logic errors, security vulnerabilities, bad dependency choices, race conditions, and plenty else worth finding.

But code review is evidence about implementation. It should not be mistaken for enforcement of authority. If a system’s safety depends on a reviewer noticing every dangerous operation an implementation could perform, we have made human attention part of the security boundary.

A capability boundary gives us another layer. The reviewer can miss something, the model can misunderstand something, the implementation can contain behaviour nobody anticipated, and if the component was never granted the capability required to produce a particular consequence, that consequence remains unavailable.

That is a considerably stronger guarantee than “we reviewed the code and did not see it doing that.”

What Survives Regeneration?

Suppose an agent writes a component today. Tomorrow we regenerate it with a different model. Next month we replace the implementation entirely. What should remain stable?

The code does not need to. The tests probably should. The behavioral contract certainly should. The authority boundary absolutely should.

Which suggests a future where the most important artifacts around software are not implementation artifacts at all. They are declarations of what correctness means and what consequences an implementation is entitled to create.

The model can write the code. It can even propose the capabilities it believes the code needs. But the system executing that code should have an independent answer to a much more important question: what is this implementation allowed to do?


The host is about 200 lines of Python with one dependency. Four scenarios: correctness, authority, discover, verify. The last one is the interesting one, because it is where the obvious answer breaks. Gist on GitHub.

Facebooktwitterredditlinkedinmail

Your Metric Is Not Your State

What a refractometer taught me about observability

In July, I wrote about a farmer walking a vineyard row at dawn, crushing a grape onto a refractometer prism, and typing “13.5 Brix” into a chat window. The agent on the other end didn’t care how the number arrived. To the Digital Scribe, I wrote, a number is just a number.

In the vineyard, that’s a defensible position. Fresh grape juice is exactly what a refractometer is built to read.

This fall, I’ve been pointing the same kind of instrument at juice that’s fermenting. It turns out the agent should have cared.

Three batches and a prism

Over the last several weeks I’ve started three small batches: a blackberry wine that’s now bulk aging, a pear wine that began fermenting on September 24, and a fireweed honey mead that got its yeast two days later. They’re one-gallon batches, which matters for this story, because at that size every sample you pull is wine you don’t get to drink and oxygen you’ve let in.

So I track them with a refractometer. A few drops of liquid on a glass prism, close the cover, hold it up to the light, read a number off the scale. The number is Brix, roughly the percentage of dissolved sugar. It’s fast, cheap, and costs a few drops per reading. During the blackberry’s primary fermentation I took readings twice a day, every time I punched down the cap of fruit.

And once fermentation starts, the refractometer is wrong.

The instrument isn’t lying, exactly

A refractometer measures how much light bends as it passes through a liquid. Dissolved sugar bends light, so in fresh juice the amount of bending is a good proxy for the amount of sugar. The scale on the instrument is calibrated on that assumption.

Fermentation breaks the assumption. Yeast converts sugar into alcohol, and alcohol bends light too. A few days into primary, the refractometer is reporting the combined effect of the sugar that remains and the alcohol that’s been produced, and presenting all of it as sugar. The reading comes out higher than the real sugar content. Taken at face value, it says the fermentation is further behind than it actually is.

To get something usable, you run each reading through a correction formula. Most home winemakers use calculators built on formulas originally developed for brewing, and every one of them needs the same extra input: the original reading, taken before the yeast went in.

Here’s what that looked like on the blackberry:

Day Raw refractometer (Brix) Corrected value
08/29/2026 20.7 20.7
08/30/2026 18.4 16.5
08/30/2026 17.0 14.3
08/31/2026 13.2 8.1
08/31/2026 10.8 4.1
09/01/2026 8.7 0.7
09/01/2026 8.3 0.1

Nothing is malfunctioning. The instrument is doing exactly what it was built to do. What changed is the system it’s pointed at, and a number that was a reliable proxy in one state became a misleading one in the next.

Neither instrument measures sugar

The traditional alternative is a hydrometer, a weighted glass float. The deeper it sinks, the less dense the liquid. Sugar makes liquid denser, so falling specific gravity means sugar is being consumed.

Except the hydrometer doesn’t measure sugar either. It measures density, and alcohol is less dense than water, so as fermentation proceeds the alcohol pulls the reading down on its own. A finished dry wine routinely reads below 1.000, lower than plain water, which would be impossible if the hydrometer were really a sugar meter.

So the refractometer measures how light bends and the hydrometer measures how heavy the liquid is. Neither measures what I actually care about: how much sugar is left, whether the yeast is still working, whether this wine is done. I infer those. The instruments give me evidence.

Even a steady reading is ambiguous. The traditional sign that fermentation has finished is the same reading several days running. But a fermentation that has stalled also produces the same reading several days running. Mead is notorious for this: honey has very little natural buffering, the pH drifts down as fermentation proceeds, and the yeast can slow or stop well short of dry. A flat line can mean done or stuck, and the instrument can’t tell you which. You tell them apart with other evidence: where the line flattened, what the pH is doing, what you expected to see.

That distinction between evidence and state sounds pedantic until you notice it’s one we get wrong in software constantly.

Your dashboard has the same problem

CPU utilization isn’t health. A service at 15% CPU can be deadlocked, and one at 90% can be doing exactly what you want. An HTTP 200 isn’t correctness; it tells you the server returned a response, not that the response was right. A green health check tells you the health check endpoint answered. Model confidence isn’t truth; it’s a number the model produced about its own output.

Each of these started as a reasonable proxy under a particular set of conditions. Then the system changed, a new failure mode appeared, or someone started optimizing the number itself, and the proxy drifted away from the state it was meant to represent. The dashboard kept rendering it with exactly the same authority.

That’s the refractometer problem. The metric was calibrated for one state of the system and is now being read in another, and nothing on the display tells you so. And like the flat fermentation line, a steady metric can mean two opposite things: a quiet error rate might mean a healthy service, or one that stopped receiving traffic an hour ago.

The failure isn’t collecting the metric. It’s treating the metric as the state instead of as evidence about the state.

A reading needs its history

Back to that correction formula. To interpret today’s reading, you need the reading from before fermentation began. Without it, the current number is close to uninterpretable. The same Brix value could describe a fermentation that has barely started or one that’s nearly finished, depending entirely on where it began.

The meaning of a measurement depends on its lineage.

This is where I think observability practice most often falls short. We store the value and the timestamp and call it a record. What we usually drop is everything needed to interpret it later: which instrument produced it, under what conditions, what calibration assumptions it carried, what baseline it should be read against, and what state the system was believed to be in at the time.

My fermentation log keeps the raw reading and the corrected value side by side. The corrected value is what I act on. The raw value is what I re-derive from if I later learn the correction was off, find a better formula, or discover I misread the original. A corrected value on its own is a conclusion with its evidence thrown away.

That’s a pattern worth stealing. Store the observation separately from the interpretation. Keep the provenance that makes the observation meaningful. Let state be something you derive, with its evidence attached, rather than something you overwrite. It’s the same idea behind the forensic receipt work I’ve been doing: a claim should carry what it was based on.

Watching isn’t free

One more thing the one-gallon batches have taught me. I use a refractometer because a hydrometer needs a much larger sample, and in a small carboy that sample is a real fraction of the batch, and pulling it lets oxygen in. The act of measuring changes the system being measured.

It also means deciding how often to look. The blackberry had to be punched down twice a day anyway, so twice-daily readings were free. The pear and the mead don’t need that kind of handling, so I’ve settled on a reading every couple of days: often enough to catch a stall, rare enough to leave the batch alone. That’s a measurement policy, and it means my log has deliberate gaps in it. Those gaps are their own story, and I’ll come back to them.

Software has the same trade-off: tracing overhead, probes that add latency, sampling that shifts timing. Usually the effect is small. Sometimes it isn’t, and how often to look is a design decision with costs on both sides.

Evidence, then state

Here’s what I’ve taken from a few weeks of squinting at a prism.

Separate what you observed from what you concluded. “The refractometer read 9 Brix” and “fermentation is about two-thirds done” are different kinds of statement. They belong in different places, and they should be allowed to disagree.

Keep the raw reading. Corrections, normalizations, and aggregations are interpretations. If you keep only the interpreted value, you can never revisit the interpretation.

Record the conditions that make a reading meaningful. Instrument, baseline, calibration assumptions, the believed state of the system. A metric without its context is a number waiting to be misread.

Treat state as a derived claim. “Fermentation is complete” isn’t a reading. It’s a conclusion drawn from several readings over time, plus whatever evidence rules out “stuck,” and it should be traceable back to all of it.

Notice when your proxy stops being a proxy. Every metric was calibrated for some range of system behavior. When the system leaves that range, the metric keeps reporting with exactly the same confidence.

The agent in July was right about one thing: it shouldn’t matter whether a number comes from a clipboard or a five-thousand-dollar probe. Where it was wrong was in believing the number could travel without its history.

The pear wine is still in primary as I write this. The raw refractometer reading says it has a long way to go. The corrected number says otherwise, and the corrected number is only trustworthy because I wrote down what the juice read before the yeast went in.

The number on the instrument isn’t the system. It never was.

Facebooktwitterredditlinkedinmail