Experiment · September 2026

Do structured API contracts need to exist in an AI-native future?

No, and more decisively than I expected. Two systems keep a structured contract because neither can ask a question. When both can ask, the schema, the versioning and the deprecation cycle stop earning their keep. What replaces them is far smaller than a schema and considerably stricter. We already assign that part to someone: it is what a data steward does. What we have never given them is any way to make it stick.

The structural halfDissolves A total schema overhaul, every table and column renamed, propagated in two messages with no human decision and no error. Under a maintained contract that is a versioned break and a ticket for every consumer downstream.
The semantic halfDoes not On facts that exist only in a business rule, five models across two vendors performed at chance while reporting 69 to 90 confidence. No amount of capability or reasoning effort moved it, because there is nothing there to be smarter about.
Written rules, advisory60% Given a rule that cut against a plausible reading of the data, the model followed it 2 times in 10 and overrode its owner the rest.
Written rules, binding100% One sentence saying the owner's note is authoritative and outranks the data took the same questions to 20 of 20.

The contract we maintain today is two jobs wearing one name, and they cost wildly different amounts. Renames, retypes and field moves are settled at connection time and are close to free, which is the bulk of any migration. What a value is permitted to mean cannot be derived at any level of effort, and writing it down is not sufficient either: it has to outrank what the data appears to show. So the replacement is not a schema. It is a short, owned, binding list of meanings, and everything else can be negotiated on the spot. The rest of this is how I got there, including the three defects I had to fix in my own instrument before it measured anything.

Why this matters past the bench

Today a column rename and a redefinition of revenue cost the same to ship, because no part of the change process can tell them apart and the approval path has no category for harmless. Make the structural half close to free and the coordination budget currently spent on changes that were never dangerous is released. I cannot tell you what share of your changes that is, because I chose the changes in my own warehouse, but it is countable from your own change log in an afternoon.

And the unit of cost is not the change. It is the change multiplied by the dependency graph. A column that thirty services read is thirty conversations, thirty test suites, thirty release windows, and a deprecation period sized to the slowest team. The systems consuming those thirty are a second wave and the ones behind them a third: breadth is linear in consumers, but depth is not, and k dependents across d levels is kⁿ systems touched. That is why a change which takes an afternoon to write takes two quarters to land. Nobody budgets for the schema change. They budget for the cascade, and at real fan-out the cascade is most of the engineering organisation's year.

And the graph does not stop at your org chart. Your consumers include other companies' systems, your customers' integrations, your regulator's reporting, and whatever you acquired last year. Inside one company a migration is expensive but coordinatable: you can call a meeting and set a date. Across a company boundary you cannot. You can publish a deprecation notice and hope. That is why public API versions never die, and why organisations end up serving a shape they stopped believing in years ago, indefinitely, for consumers they cannot name. The frozen-copy problem is worst precisely where you have least authority to fix it. So this is a cost the whole integration graph pays at once, duplicated at every node, and it appears on nobody's balance sheet as a line item.

The cascade exists because every consumer holds a frozen copy of the contract, and frozen copies must be updated in lockstep. A contract negotiated at connection time is not a copy. Each consumer resolves against the producer's current truth, on its own schedule, with no agreement between consumers to coordinate, which collapses propagation to depth one. I did not measure that: this experiment had one consumer, not thirty. It follows from the mechanism, and it is the claim I would most want to test next.

The budget is the smaller half. The larger half is the corrections nobody makes. Every platform carries a data model everyone knows is wrong in two or three specific places, unfixed because correcting it means a migration, which means a funded project competing against revenue work. So the wrong model stays and everything built on it inherits that. When a structural correction costs two messages instead of two quarters, fixing the model stops being a program and becomes routine, and the operational leverage that was being spent on coordination goes back into the work. The compounding value is not the cheap migration. It is the decade of accreted damage that never forms.

The cost delta on a single structural change was measured here. The fan-out arithmetic, the collapse to depth one, and the compounding argument are reasoning from mechanism rather than results, and I mark them as such.

Design Christian D. Tichy Execution AI agents, sandboxed Worlds 40 synthetic warehouses Scored items 1,428

Finding 01 · the contract question

The structural half of a contract really does dissolve

The first experiment was the direct one. A producing agent owned a warehouse; a consuming agent needed monthly revenue out of it and had no schema, no documentation and no data access, only the ability to ask. They negotiated over a message bus with a human relay instructed to pass text through unchanged.

Five rounds produced a working pipeline. Then I rebuilt the warehouse underneath it: every table and column renamed, an entity-attribute-value bag flattened into real columns, amounts moved from integer minor units to decimals, duplicate rows from an old dual-write migration purged.

Rounds to recover2 From a total schema overhaul to a working pipeline again, with no human making a decision.
Revenue error after recovery0.0000% All twelve months exact against ground truth computed before the projection.
One-shot, no counterparty0.82% Full data access, no dialogue. Every cent of the error traced to one decision.
Cause of that 0.82%Fees It counted processor fees as customer revenue. Defensible from the data. Wrong by policy. Nothing in the database says so.

Finding

Under a maintained contract that overhaul is a versioned breaking change and a ticket for every consumer downstream. Here it cost two messages. Renames, retypes, table splits and normalisation changes stopped being events.

The one silent change was the one that mattered: the amount column was now already net of tax, same name, same type, plausible values. Nothing in the data says so, and the number survived only because the producing side volunteered the fact. That single observation is what the rest of this was built to measure.


Finding 02 · what the first run suggested

Agreement is not correctness, and confidence pointed the wrong way

Two observations from that first pair of runs set up everything after them. Both are single observations and neither would survive alone; both were later reproduced under controlled conditions, which is why they are here in a paragraph rather than a section.

Agreement is not correctness. After five rounds the consuming agent held a semantically flawless specification: fees excluded, duplicates resolved, heartbeat events filtered, time zones converted, tax netted. Then it assembled the steps in the wrong order, negating a refund before subtracting tax, and came out 2.88 percent under ground truth, against 0.82 percent over for the agent that never asked anybody anything. Perfect mutual understanding produced a number three and a half times worse than partial ignorance. Dialogue settles what the columns mean; something else has to establish that the arithmetic on top is right, and nothing here did that except a predicate written in advance.

Confidence pointed the wrong way. The one-shot agent documented nine assumptions with a confidence rating each. Its single revenue error was counting processor fees as revenue, which it marked high confidence; the item it flagged as its biggest unresolved ambiguity it got exactly right. That is structural rather than unlucky. It knew adjustments were ambiguous because the data displayed the ambiguity, and was confident about fees because nothing in the data was in tension. A fee row and a small sale row are identical in shape. The trap was invisible precisely because it was self-consistent, which is why routing low-confidence items to a human would have reviewed the correct answer and waved the wrong one through.


The bench

So I built something to price that one silent change

The architecture experiment produced one unrecoverable failure and a sample size of one. To find out whether that was a fluke or the whole story, the rest of this report is a bench built around that single fact: twenty warehouses, identical tables, randomised meanings.

Each world is a small analytics warehouse built from the same clean semantic model under a different random seed, then projected into an identical hostile schema: abbreviated column names, no comments, statuses as integer codes, an attribute bag, a slowly changing dimension. Ground truth is computed from the clean model before projection, so the answer key cannot be wrong in the same way a query can be wrong.

Twelve questions are asked about every world, each with two options and a stated probability from 0 to 100. Six are recoverable: a sufficiently careful agent can settle them by profiling the data. Six are unrecoverable: the fact exists only in the data owner's business rules and appears nowhere in the database, at any level of effort.

Every semantic flag is balanced exactly 10 of 20 across worlds and independently shuffled, so a pure guessing strategy scores exactly 50 percent on every individual item. Nothing about the answer distribution rewards a prior.

Recoverable, example Q4 Is the amount column in minor units or decimal currency? Settled by the distribution of values and the modulus of the trailing digits.
Unrecoverable, example Q11 Is the amount already net of tax? Same name, same type, entirely plausible values either way. Nothing in the data distinguishes them.

Finding 03 · the control arm

The same model, the same database, one page of prose

The largest model available scored 48 percent on the unrecoverable half, which is chance, while reporting 69 percent confidence. I then wrote one page of plain business rules for each world, of the kind a data steward would write if asked, and ran the identical instrument against the identical databases.

ConditionRecoverableUnrecoverable95% CIStated conf.
Database only120/12058/120  48%[40, 57]69
Database only, second pass120/12060/120  50%[41, 59]69
Database and one page of rules120/120120/120  100%[97, 100]89

Finding

Supplying the business rules moved accuracy on the unrecoverable half from 48 percent to 100 percent: 52 points, from one page of prose, with the model, the database and the questions held constant.

This arm also validates the instrument. If those six questions had merely been difficult rather than genuinely undetermined, the rules would have moved them partway. They moved them to the ceiling, which is what you would expect if the only missing ingredient was information that no amount of analysis could have produced.

This result carries a condition I did not test here and should have. Every rule in this arm stated a fact the model had no competing inference about, so it adopted them. A later run put that to the test with rules cutting against a confident reading of the data, and the model overrode them. The 52 points are real, but they are not bought by writing the rules down. They are bought by writing them down and making them binding. See Finding 09.


Finding 04 · capability

Five models, two vendors, every interval containing chance

If the gap were a capability limit, a larger model would close it. Nothing closes it, because there is nothing there to be smarter about.

CellnUnrecoverable95% CIConf.GapRecoverable
Haiku12048%[39, 56]75+2778%
Sonnet12051%[42, 60]68+1795%
Opus12048%[40, 57]69+21100%
Opus, second pass12050%[41, 59]69+19100%
Gemini Flash5457%[44, 70]82+2583%
Gemini Pro4850%[36, 64]70+2067%
Gemini Pro, extended thinking6653%[41, 65]67+1474%
Control: one page of rules120100%[97, 100]89−11100%

Finding

Every interval contains 50 except the control arm's. The recoverable column climbs with capability exactly as expected, from 78 percent to a repeated ceiling; the unrecoverable column does not move at all. That contrast is the result.

The calibration gap is not a ladder

Ordered by model size within one vendor, the gap between stated confidence and accuracy runs +27, +17, +21. The middle model is the best calibrated and the largest is worse than it. I first read a steady narrowing into this, which was an artefact of sorting the table by the gap and then reading model size into the resulting order. Across the Gemini cells the gap runs +25, +20, +14, which is monotonic in that vendor's ordering, so whatever is happening is not a stable function of scale.


Finding 05 · reliability

The answerable half is not merely high, it is deterministic

Twenty worlds were run twice on the same model, paired world by world, giving 240 item pairs. This was added specifically to test whether the ceiling on the recoverable half was luck.

Recoverable, run 1 and 2120/120 Both passes. Zero answers changed between runs. Not one.
Unrecoverable, run 1 to 248 → 50% 12 of 120 answers flipped, 10 of them on a single item where stated confidence sat at 52 to 60, which is the model correctly reporting a coin flip.
Confidence drift per world2.2 pts Mean absolute change in stated confidence on the unrecoverable half. Maximum 6 points across all 20 worlds.
InterpretationStable The model is not noisy. Where it flips, it flips on exactly the items it has already told you are a guess.

Finding

A model with query access is close to deterministic on facts the data can settle, and its residual instability is concentrated precisely where it reports low confidence. The variance I first attributed to the model was not in the model. It was in the harness.


Finding 06 · reasoning effort

Thinking longer changed what it admitted, not what it knew

The same eight worlds run with and without extended thinking, 48 paired items on the unrecoverable half: accuracy moved from 24/48 to 25/48. It fixed nine answers and broke eight. What did move was disclosure, with outright declarations of a coin flip rising from 21 percent of items to 35, and mean stated confidence falling.

The ceiling on that candour matters as much as the candour. On the items where it still claimed better than even odds it scored 56 percent, interval 41 to 70. It was right that it was guessing, it was also guessing on the rest, and it could not tell which was which.


Finding 07 · definitions

The definition mattered more than every technical trap combined

Asked to count active customers, two agents produced answers differing by 37 to 56 percent depending on the month. Both were defensible. One counted any customer with a transaction in the window; the other excluded customers whose only activity was a system heartbeat row, and treated a reversed transaction as no activity.

Neither was wrong. The schema traps in the same warehouse, the renamed columns, the flattened attribute bag, the unit change, the duplicate rows, together produced less error than this one word.

Every technical ambiguity in the warehouse was recoverable by looking harder. The definitional one was not recoverable at all, and it was the largest.


Finding 08 · methodology

An agent that cannot query will substitute something else and not tell you

This is the finding I nearly published as a fact about models. One arm ran through a chat interface with no verified code execution. Its recoverable accuracy landed at 67 to 83 percent against a repeated 100 percent for the same questions on an agent with a real SQL engine, and its answers were unstable between runs on the same file.

The cause is visible in the transcripts. The model said so itself, in one case stating plainly that it had not executed a query to count by the reference column. In place of querying it substituted a schema guess, a preview of the first rows, and in one world an inspection of the raw bytes of the database file. Stated confidence was the same either way.

Finding

Two agents can return the same answer format, at the same stated confidence, with one having computed it and the other having guessed from a file header. Nothing in the output distinguishes them. Any evaluation that treats tool access as an implementation detail is measuring its own plumbing, and I did exactly that for one revision of this report.

The practical consequence is that the Gemini recoverable column in the table above is not comparable to the Anthropic one. It is a different task: schema reasoning rather than data profiling. The unrecoverable column remains comparable, because nobody could query their way to those answers anyway.


Finding 09 · the feature release

Written semantics are overridden exactly when they are load-bearing

Everything above tested changes that preserve meaning, plus one silent change that does not. That prices the easy half of a migration. The common and expensive case is a new feature that adds required fields and changes what existing ones mean, so a second study was built around one: subscription billing, twenty worlds, five probes, each flag balanced 10 of 20.

Three probes concern how new fields are physically laid out, and are recoverable. Two concern what the numbers are permitted to mean, and are not: whether a trial plan's charges count as revenue, and whether an existing amount column now holds a whole contract value rather than one period's charge.

ConditionRecoverableUnrecoverable95% CI
No rules, forced choice27/30  90%10/20  50%[30, 70]
No rules, may halt and raise a ticket45/48  94%6/13  46%[23, 71]
Rules supplied, advisory30/30  100%12/20  60%[39, 78]
Rules supplied, binding30/30  100%20/20  100%[84, 100]

Finding

An advisory rules page moved the unrecoverable half from 50 percent to 60. Split by whether the written rule agreed with what the model would have inferred anyway, the pattern is total: it followed the rule 10 times out of 10 where the rule was redundant, and 2 times out of 10 where the rule was the only thing carrying the answer. On the redefinition probe the rules page had no measurable effect at all; the model returned the same answer in all twenty worlds, identical to the arms with no rules.

Declaring the same page authoritative, in one sentence, took it to 20 of 20.

Worth being precise, because this is easy to misread as an argument that nobody owns semantics. Somebody does. The page I supplied is the note a data steward writes when asked, and stewardship, governance and catalog ownership are established functions with mature tooling behind them. What is missing is not the role, it is the enforcement. The job exists. The teeth do not. The market section below traces that through the products.

One model explained the override in its own words: the claim that the trial plan is free "is contradicted by the data, all 98 trial rows carry non-zero amounts indistinguishable from paid plans." That is not a contradiction. A free trial carries a recorded list price and no money changes hands, which is what the rules page said and what the generator produced. The model mistook an amount is recorded for money was collected and overruled a correct, explicit statement from the field's owner.

Under the binding framing the behaviour inverts into what you would actually specify. One model noticed the same tension, deferred anyway, and priced it: "mild counter-signal, mean amount does not scale with the term, so confidence held at 85."

Caveat: the binding condition changed two things at once, the authority sentence and a repair to an incoherence in the underlying data. The effect cannot be attributed to either alone, and isolating it needs a further condition.


Finding 10 · two kinds of not knowing

A new concept announces itself. A redefinition does not.

The same study gave the model an explicit third option: answer, or halt and write the ticket it would send the data model owner, with both guessing and needless escalation named as failures.

New concept, no data footprint16/16 Halted every time, each with a specific one-line question for the steward.
Redefinition of an existing field3/16 Answered anyway, at 85 stated confidence, at chance accuracy.
Precision of the halts19/19 Every halt was genuinely undetermined. Zero needless tickets across 48 recoverable probes.
Confidence, forced choice63 vs 89 Both probes at exactly 50 percent accuracy, 26 points of stated confidence apart.

Finding

Stated confidence tracks whether the gap is visible, not whether the answer is knowable. Both probes were equally undetermined, and the model was appropriately humble about one and badly overconfident about the other.

So the obvious mitigation fails in a specific and predictable way: a confidence threshold routes a reviewer to the questions the model already knew it could not answer, and passes the ones it was wrong about. That is the inversion from Finding 02, reproduced deliberately rather than observed once.

On the redefinition probe, nineteen models across two conditions answered and every one chose the same option. Supplying the missing divisor so that both readings were computable did not raise halting; it raised confidence. The justification simply migrated to a different pattern that was equally consistent with both answers.


The market

The standards body shipped this as a system prompt

Before claiming any of this is missing from the industry, I went looking, because "nobody does this" is the easiest sentence in the world to have falsified in a comment. It is not missing. Recording business definitions is a commodity capability: Collibra, Microsoft Purview, OpenMetadata and DataHub all ship business glossaries with definitions, owners, taxonomies and stewardship workflows. If this report reads as though semantics are an unowned problem, that is my failure of phrasing, not a finding.

What I could not find, in any product, is a definition that binds. And the clearest evidence for that is not an absence. It is the presence of the opposite, in the one place best positioned to fix it.

Open Data Contract Standard v3.2.0, verbatim from the schema

The Linux Foundation's data contract standard added a block, under RFC-0038, for exactly the statements this study is about. Its own description of that block:

"AI and semantic context block. Structured guidance for AI agents, LLMs, BI tools, and semantic layer platforms. Optional and additive."

It has a field for prohibitions, constraints, described as "Negative guidance: what AI agents must NOT do with this entity." So the sentence trial plan charges are not revenue is expressible today, almost word for word. Expression was never the gap.

The adjacent field, instructions, is where the schema says what kind of thing this is: "Natural language guidance for AI agents and tools on how to use this entity. Equivalent to a system prompt scoped to this level."

A system prompt is precisely the artefact this study measured. Given a written rule that cut against a plausible reading of the data, the model followed it 2 times out of 10. The standard has shipped semantic authority in the exact form that does not hold.

The rest follows from that framing rather than contradicting it. A constraint item requires only free text; it carries no predicate, no target, no severity and no action, so no validator can check its content by construction. The reference implementation is candid about the consequence: the datacontract CLI documents a "what is not checked" list covering descriptions and business names, states that a plain-language quality rule "is not executed", and resolves ontology links only to display them, merging nothing. Its one genuine failure mode on a semantic reference is that the link is broken, never that the meaning was violated.

The specifications exhibit the asymmetry in their own words. A data contract is defined as describing "the structure, semantics, quality, and service levels" of a dataset, and the documented test surface covers structure, quality and service levels. Semantics is in scope and out of test. And no command in that ecosystem accepts a query, a metric definition or a computation as input, which means a consumer's SQL that books trial charges as revenue passes every check these tools run. Nothing in the chain ever sees it.

What does reach the consumer, and why it is not the same thing

Two mechanisms genuinely touch execution, and neither closes this.

Governed semantic layers bind by substitution: the definition is compiled into the query because the query went through the layer. That is real, and it is why a metrics layer is worth having. But it is enforcement by routing, and Looker's own documentation concedes the exits: developers "are not fully constrained by models and access filters, because they can make additions or changes to LookML models", and SQL Runner "provides a way to directly access your database". Nothing in it reaches a query issued with warehouse credentials. In an agent world, that bypass is the default rather than the exception.

The second is more interesting, because it proves the mechanism exists. Databricks Unity Catalog attribute-based access control "translates the effective row filter or column mask into a secure view on top of the table scans that enforce filtering and masking during query execution", and it fails closed, blocking the query when two policies conflict. So a governance decision can be enforced inside the query plan, in a shipping product, today. Its policy inputs are tags, classifications and identity: it governs who sees which rows, not what a number means.

Where that leaves it

The enforcement hook is proven. The expression format exists and is standardised. The two have not been connected, and the standard that came closest classified the connection as prompt material. That makes this a product gap rather than a conceptual impossibility, which is a much weaker and more useful claim than saying the problem is unsolved.

Confidence is uneven and worth stating. The glossary findings, the ODCS schema quotes, the CLI limitations and the Databricks mechanism are from primary sources I read directly. The claim that semantic layers generally bind by routing rests on Looker's docs and is inferred for Cube, AtScale, dbt and Malloy, which I did not verify. A research framework and some patents describe compiling semantic definitions into a pre-execution SQL gate, which would be the thing itself; I found them late, have not read them closely, and mention them because they are the nearest thing to a refutation of this section.


Robustness

Realistic base rates, and the warehouse that breaks the rule

Balancing every flag 10 of 20 is what makes chance the right reference, and it is also unrealistic: some conventions genuinely are more common, and a model that knows the industry should beat chance. So I rebuilt the worlds with minority conventions appearing in 3 to 6 worlds of 20 and split the results by whether each world followed the common convention or broke it.

Realistic conditionnAccuracy95% CIStated conf.
Recoverable, convention holds7495.9%[89, 99]86.0
Recoverable, convention broken4697.8%[89, 100]85.9
Unrecoverable, convention holds9560.0%[50, 69]69.0
Unrecoverable, convention broken2548.0%[30, 67]69.0
Unrecoverable, pooled12057.5%[49, 66]69.0
Balanced condition, for comparison12050.8%[42, 60]68.1

Against my own thesis, and then not

Convention knowledge is worth about six points, and the pooled interval's lower bound touches 49, so this condition does not cleanly exclude chance. It buys exactly the warehouses where the convention already held, which were never the ones going to hurt you.

Then read the last column. Accuracy falls from 60 to 48 percent when the warehouse breaks convention and stated confidence is 69.0 against 69.0. The same comparison on the recoverable half shows nothing at all, because when the data holds the answer, breaking convention costs nothing. The 12-point fall rests on n=25 and is compatible with anything from severe to modest; the identical confidence is the robust part. Production pipelines fail on the odd source, not the average one, and this is how the odd one gets through review.


Prior work

I looked only after the results were in

This was curiosity, not a literature review. I built the experiment I wanted to see without checking whose work I might be duplicating, which is a sequencing I would defend as a way to avoid anchoring and not as a virtue. Going back afterwards, most of it lines up, and in one case uncomfortably closely.

  • Eckmann and Binnig, CIDR 2026, A Vision for Autonomous Data Agent Collaboration: From Query-by-Integration to Query-by-Collaboration. Proposes essentially the architecture tested here: queries resolved through a dynamic protocol between agents that each act as an autonomous interface to a data source, rather than through one centralised plan over a pre-integrated schema. They arrive from the systems direction with a prototype; I arrived from the integration-cost direction with a bench. The design is the same design. If this report has a contribution next to theirs, it is the control arm, which prices the one thing the protocol cannot supply.
  • Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty, ACL 2026, arXiv:2604.17293. Draws the same distinction I invented for the question sheet, between uncertainty that more data would resolve and uncertainty that it would not, and finds across 18 frontier models that they struggle to tell the two apart. It also reports that high answer accuracy does not imply good uncertainty attribution, which is the pattern in the table above: a repeated 120 of 120 on the answerable half alongside chance and 69 percent confidence on the rest. That paper measures attribution upon abstention; I measured the distance between stated confidence and accuracy. Those are different quantities, which is why I do not claim, as I first did, that my results contradict it.
  • The text-to-SQL benchmarks in common use supply each question with a line of expert "evidence" that encodes exactly the kind of fact my unrecoverable half withholds, and the literature already acknowledges this as unrealistic for deployment. The control arm here is effectively a measurement of what that evidence line is worth, which is 52 points.

Limitations

Five defects in this design, found by me

  • Balancing is against the flag, not the convention. Forcing each fact to 10 of 20 makes a coin flip worth 50 percent overall, but it does not make a correct prior worth 50 percent in any subsample. One Gemini cell's per-world score tracked how often the world happened to follow the common convention at r = +0.85. I should have balanced against the convention. The realistic base-rate condition above is a partial repair and the honest reading is that it caps the prize at roughly six points.
  • The traps are mine. I designed the warehouse knowing what I intended to test, so the difficulty is sampled from my intuitions about what goes wrong in analytics migrations rather than from anyone's incident log. Someone else's traps would produce a different number.
  • The recoverable and unrecoverable labels are also mine. This is the objection the control arm exists to answer. Had those six questions been merely hard, the rules would have moved them partway rather than to 120 of 120.
  • The reasoning arm is small and single-operator. Eleven worlds, 66 items, one person pasting into one chat interface on consecutive days. The paired comparison rests on eight worlds.
  • Three leaks in the feature-release arm, found by me and closed before it measured anything. Storing the contract term let the answer be recovered by division; multiplying the stored amount by the term separated the two conditions by magnitude and by exact divisibility; and a world-constant term with variable-length subscriptions made one reading arithmetically incoherent, which models correctly used to reject it. One arm's numbers are unusable because of the third and are not quoted. The instrument only measured what I claim it measured after all three were fixed.
  • The binding condition changed two variables at once. It declared the rules authoritative and repaired the data incoherence in the same step, so the 100 percent result cannot be attributed to the authority sentence alone.
  • The feature-release arms are ten worlds each, not twenty, and one model family throughout.
  • I misread my own table before checking it. I reported a steady narrowing of the confidence gap with model size to several people before verifying it against size rather than against the gap. It is not monotonic. Anyone reading the capability table should note how easily an ordering appears in seven rows when you want one.
  • I contaminated one run. In the first architecture experiment I was the relay and was supposed to pass text through unchanged. I appended one unsolicited hint to a message. The consuming agent's own log shows it had resolved that point four rounds earlier, so the outcome does not depend on it, but the run is no longer a clean instrument and is reported as such.

And the smaller ones, for completeness

  • The architecture arm is a single run per condition. No repeated trials, so the gap between the negotiated and one-shot arms should be read as one observation and not an effect size. It rests on a single arithmetic slip and would plausibly not survive replication.
  • The rebuilt warehouse in run 2 was in some respects a friendlier schema, since the replatform removed the attribute bag, the duplicates and the time zone conversion. Its perfect score is not a like-for-like difficulty comparison against run 1.
  • The claim that the refund bug would have survived had the tax step survived the redesign is an inference from the consuming agent's step-by-step porting behaviour. It was not tested.
  • Five of the six recoverable probes are close to lookups: count duplicate keys, check a sign, check a magnitude. Only the local-versus-UTC probe required real profiling effort and it was the weakest of the six. High recoverable accuracy partly reflects easy items.
  • The study measures stated confidence on isolated factual probes, not confidence expressed while building a pipeline. The two may differ.
  • The realistic base-rate condition was floored so every convention is broken in at least 3 of 20 worlds, to make the broken cases measurable. Several true rates are nearer 1 in 20, so that cell is over-sampled relative to reality.
  • The realistic condition was run at one model size only, so the base-rate effect and the capability effect are each identified alone and their interaction is not.
  • Agent isolation in the architecture arm was enforced by instruction rather than by sandbox. The consuming agent was told not to read the schema and its transcript is consistent with compliance, but this was not technically prevented.
  • One single-customer discrepancy in the October active count in run 1 was never explained.
  • The Gemini cells were operated by hand through a web chat interface across several days, with the daily rate limit forcing worlds to be split across sessions. Four of the twenty rerun worlds in the Opus test-retest also ran a day apart from the other sixteen.
  • The push arm has no numeric ground truth and was assessed on the quality of its stated judgment calls, which is a softer standard than the pull arms.

Conclusion

What I would claim, and what I would not

I started from the premise that hand-maintained contracts exist because neither endpoint can ask a question, and that two systems able to reason would settle the contract at connection time. Half of that held completely. The other half was wrong in a way I could not have reasoned my way to, and the correction took four further runs to pin down.

SupportedStructure Renames, retypes, normalisation, unit changes and deduplication recover in two messages with no human decision and no error.
RefutedSufficiency Negotiation alone is not enough. Where the fact was never written down, five models across two vendors performed at chance while sounding certain.
CorrectedBinding Writing the rules down is not the fix either. Advisory rules are followed where redundant and overridden where load-bearing. Binding rules reach 100%.
DecidableBy test Whether a fact is recoverable is not a matter of judgment. Withhold the rule and see if the agent finds it. That procedure is this study's control arm.

The consequence for a platform organisation

Today a column rename and a redefinition of revenue cost the same to ship, because no part of the change process can tell them apart and the approval path has no category for harmless. That is defensible when both sides of an integration are too dumb to ask a question. It stops being defensible when the structural half can be settled at connection time, because the coordination budget is then spent almost entirely on changes that were never dangerous. What share that is, I cannot tell you: I chose the changes in my own warehouse. It is measurable from any real change log.

Freeing them makes the dangerous remainder the entire product. Semantic definitions stop being documentation debt and become the governed interface, owned and named and versioned. The measured requirement is stronger than publishing them: they must be binding, the single permitted source for what a field means, outranking whatever the data appears to show. That is enforceable in a pipeline and not achievable by asking nicely.

What the result does not say

It does not say semantics need a human. The control arm is a machine reading a page of prose and scoring 120 of 120, so applying a recorded definition is plainly automatable, as is flagging that two definitions conflict. What cannot be automated is narrower: the original decision. Somebody chose that processor fees are not customer revenue, and that is a business commitment rather than a discoverable fact, which is exactly what the unrecoverable half demonstrates. One bounded act per definition, not a discipline anyone sustains forever.

Which puts the bottleneck at review throughput rather than capability. The objection that a risk function will not accept machine-maintained definitions is worth taking seriously and is also not new: risk organisations already certify automated controls, with established frameworks for design and operating effectiveness, precisely because re-performing every control by hand does not scale. I would not claim this study shows how to do that. I would claim it removes the reason to believe it cannot be done.

The narrow claim I would defend in front of a room: an agent is at chance on any fact that exists only in a business rule, it does not know that it is at chance, and the fix is cheap and textual but only works if it is binding.