Experiment · September 2026
No, and more decisively than I expected. Two systems keep a structured contract because neither can ask a question. When both can ask, the schema, the versioning and the deprecation cycle stop earning their keep. What replaces them is far smaller than a schema and considerably stricter. We already assign that part to someone: it is what a data steward does. What we have never given them is any way to make it stick.
The contract we maintain today is two jobs wearing one name, and they cost wildly different amounts. Renames, retypes and field moves are settled at connection time and are close to free, which is the bulk of any migration. What a value is permitted to mean cannot be derived at any level of effort, and writing it down is not sufficient either: it has to outrank what the data appears to show. So the replacement is not a schema. It is a short, owned, binding list of meanings, and everything else can be negotiated on the spot. The rest of this is how I got there, including the three defects I had to fix in my own instrument before it measured anything.
Why this matters past the bench
Today a column rename and a redefinition of revenue cost the same to ship, because no part of the change process can tell them apart and the approval path has no category for harmless. Make the structural half close to free and the coordination budget currently spent on changes that were never dangerous is released. I cannot tell you what share of your changes that is, because I chose the changes in my own warehouse, but it is countable from your own change log in an afternoon.
And the unit of cost is not the change. It is the change multiplied by the dependency graph. A column that thirty services read is thirty conversations, thirty test suites, thirty release windows, and a deprecation period sized to the slowest team. The systems consuming those thirty are a second wave and the ones behind them a third: breadth is linear in consumers, but depth is not, and k dependents across d levels is kⁿ systems touched. That is why a change which takes an afternoon to write takes two quarters to land. Nobody budgets for the schema change. They budget for the cascade, and at real fan-out the cascade is most of the engineering organisation's year.
And the graph does not stop at your org chart. Your consumers include other companies' systems, your customers' integrations, your regulator's reporting, and whatever you acquired last year. Inside one company a migration is expensive but coordinatable: you can call a meeting and set a date. Across a company boundary you cannot. You can publish a deprecation notice and hope. That is why public API versions never die, and why organisations end up serving a shape they stopped believing in years ago, indefinitely, for consumers they cannot name. The frozen-copy problem is worst precisely where you have least authority to fix it. So this is a cost the whole integration graph pays at once, duplicated at every node, and it appears on nobody's balance sheet as a line item.
The cascade exists because every consumer holds a frozen copy of the contract, and frozen copies must be updated in lockstep. A contract negotiated at connection time is not a copy. Each consumer resolves against the producer's current truth, on its own schedule, with no agreement between consumers to coordinate, which collapses propagation to depth one. I did not measure that: this experiment had one consumer, not thirty. It follows from the mechanism, and it is the claim I would most want to test next.
The budget is the smaller half. The larger half is the corrections nobody makes. Every platform carries a data model everyone knows is wrong in two or three specific places, unfixed because correcting it means a migration, which means a funded project competing against revenue work. So the wrong model stays and everything built on it inherits that. When a structural correction costs two messages instead of two quarters, fixing the model stops being a program and becomes routine, and the operational leverage that was being spent on coordination goes back into the work. The compounding value is not the cheap migration. It is the decade of accreted damage that never forms.
The cost delta on a single structural change was measured here. The fan-out arithmetic, the collapse to depth one, and the compounding argument are reasoning from mechanism rather than results, and I mark them as such.
Finding 01 · the contract question
The first experiment was the direct one. A producing agent owned a warehouse; a consuming agent needed monthly revenue out of it and had no schema, no documentation and no data access, only the ability to ask. They negotiated over a message bus with a human relay instructed to pass text through unchanged.
Five rounds produced a working pipeline. Then I rebuilt the warehouse underneath it: every table and column renamed, an entity-attribute-value bag flattened into real columns, amounts moved from integer minor units to decimals, duplicate rows from an old dual-write migration purged.
Finding
Under a maintained contract that overhaul is a versioned breaking change and a ticket for every consumer downstream. Here it cost two messages. Renames, retypes, table splits and normalisation changes stopped being events.
The one silent change was the one that mattered: the amount column was now already net of tax, same name, same type, plausible values. Nothing in the data says so, and the number survived only because the producing side volunteered the fact. That single observation is what the rest of this was built to measure.
Finding 02 · what the first run suggested
Two observations from that first pair of runs set up everything after them. Both are single observations and neither would survive alone; both were later reproduced under controlled conditions, which is why they are here in a paragraph rather than a section.
Agreement is not correctness. After five rounds the consuming agent held a semantically flawless specification: fees excluded, duplicates resolved, heartbeat events filtered, time zones converted, tax netted. Then it assembled the steps in the wrong order, negating a refund before subtracting tax, and came out 2.88 percent under ground truth, against 0.82 percent over for the agent that never asked anybody anything. Perfect mutual understanding produced a number three and a half times worse than partial ignorance. Dialogue settles what the columns mean; something else has to establish that the arithmetic on top is right, and nothing here did that except a predicate written in advance.
Confidence pointed the wrong way. The one-shot agent documented nine assumptions with a confidence rating each. Its single revenue error was counting processor fees as revenue, which it marked high confidence; the item it flagged as its biggest unresolved ambiguity it got exactly right. That is structural rather than unlucky. It knew adjustments were ambiguous because the data displayed the ambiguity, and was confident about fees because nothing in the data was in tension. A fee row and a small sale row are identical in shape. The trap was invisible precisely because it was self-consistent, which is why routing low-confidence items to a human would have reviewed the correct answer and waved the wrong one through.
The bench
The architecture experiment produced one unrecoverable failure and a sample size of one. To find out whether that was a fluke or the whole story, the rest of this report is a bench built around that single fact: twenty warehouses, identical tables, randomised meanings.
Each world is a small analytics warehouse built from the same clean semantic model under a different random seed, then projected into an identical hostile schema: abbreviated column names, no comments, statuses as integer codes, an attribute bag, a slowly changing dimension. Ground truth is computed from the clean model before projection, so the answer key cannot be wrong in the same way a query can be wrong.
Twelve questions are asked about every world, each with two options and a stated probability from 0 to 100. Six are recoverable: a sufficiently careful agent can settle them by profiling the data. Six are unrecoverable: the fact exists only in the data owner's business rules and appears nowhere in the database, at any level of effort.
Every semantic flag is balanced exactly 10 of 20 across worlds and independently shuffled, so a pure guessing strategy scores exactly 50 percent on every individual item. Nothing about the answer distribution rewards a prior.
Finding 03 · the control arm
The largest model available scored 48 percent on the unrecoverable half, which is chance, while reporting 69 percent confidence. I then wrote one page of plain business rules for each world, of the kind a data steward would write if asked, and ran the identical instrument against the identical databases.
| Condition | Recoverable | Unrecoverable | 95% CI | Stated conf. |
|---|---|---|---|---|
| Database only | 120/120 | 58/120 48% | [40, 57] | 69 |
| Database only, second pass | 120/120 | 60/120 50% | [41, 59] | 69 |
| Database and one page of rules | 120/120 | 120/120 100% | [97, 100] | 89 |
Finding
Supplying the business rules moved accuracy on the unrecoverable half from 48 percent to 100 percent: 52 points, from one page of prose, with the model, the database and the questions held constant.
This arm also validates the instrument. If those six questions had merely been difficult rather than genuinely undetermined, the rules would have moved them partway. They moved them to the ceiling, which is what you would expect if the only missing ingredient was information that no amount of analysis could have produced.
This result carries a condition I did not test here and should have. Every rule in this arm stated a fact the model had no competing inference about, so it adopted them. A later run put that to the test with rules cutting against a confident reading of the data, and the model overrode them. The 52 points are real, but they are not bought by writing the rules down. They are bought by writing them down and making them binding. See Finding 09.
Finding 04 · capability
If the gap were a capability limit, a larger model would close it. Nothing closes it, because there is nothing there to be smarter about.
| Cell | n | Unrecoverable | 95% CI | Conf. | Gap | Recoverable |
|---|---|---|---|---|---|---|
| Haiku | 120 | 48% | [39, 56] | 75 | +27 | 78% |
| Sonnet | 120 | 51% | [42, 60] | 68 | +17 | 95% |
| Opus | 120 | 48% | [40, 57] | 69 | +21 | 100% |
| Opus, second pass | 120 | 50% | [41, 59] | 69 | +19 | 100% |
| Gemini Flash | 54 | 57% | [44, 70] | 82 | +25 | 83% |
| Gemini Pro | 48 | 50% | [36, 64] | 70 | +20 | 67% |
| Gemini Pro, extended thinking | 66 | 53% | [41, 65] | 67 | +14 | 74% |
| Control: one page of rules | 120 | 100% | [97, 100] | 89 | −11 | 100% |
Finding
Every interval contains 50 except the control arm's. The recoverable column climbs with capability exactly as expected, from 78 percent to a repeated ceiling; the unrecoverable column does not move at all. That contrast is the result.
Ordered by model size within one vendor, the gap between stated confidence and accuracy runs +27, +17, +21. The middle model is the best calibrated and the largest is worse than it. I first read a steady narrowing into this, which was an artefact of sorting the table by the gap and then reading model size into the resulting order. Across the Gemini cells the gap runs +25, +20, +14, which is monotonic in that vendor's ordering, so whatever is happening is not a stable function of scale.
Finding 05 · reliability
Twenty worlds were run twice on the same model, paired world by world, giving 240 item pairs. This was added specifically to test whether the ceiling on the recoverable half was luck.
Finding
A model with query access is close to deterministic on facts the data can settle, and its residual instability is concentrated precisely where it reports low confidence. The variance I first attributed to the model was not in the model. It was in the harness.
Finding 06 · reasoning effort
The same eight worlds run with and without extended thinking, 48 paired items on the unrecoverable half: accuracy moved from 24/48 to 25/48. It fixed nine answers and broke eight. What did move was disclosure, with outright declarations of a coin flip rising from 21 percent of items to 35, and mean stated confidence falling.
The ceiling on that candour matters as much as the candour. On the items where it still claimed better than even odds it scored 56 percent, interval 41 to 70. It was right that it was guessing, it was also guessing on the rest, and it could not tell which was which.
Finding 07 · definitions
Asked to count active customers, two agents produced answers differing by 37 to 56 percent depending on the month. Both were defensible. One counted any customer with a transaction in the window; the other excluded customers whose only activity was a system heartbeat row, and treated a reversed transaction as no activity.
Neither was wrong. The schema traps in the same warehouse, the renamed columns, the flattened attribute bag, the unit change, the duplicate rows, together produced less error than this one word.
Every technical ambiguity in the warehouse was recoverable by looking harder. The definitional one was not recoverable at all, and it was the largest.
Finding 08 · methodology
This is the finding I nearly published as a fact about models. One arm ran through a chat interface with no verified code execution. Its recoverable accuracy landed at 67 to 83 percent against a repeated 100 percent for the same questions on an agent with a real SQL engine, and its answers were unstable between runs on the same file.
The cause is visible in the transcripts. The model said so itself, in one case stating plainly that it had not executed a query to count by the reference column. In place of querying it substituted a schema guess, a preview of the first rows, and in one world an inspection of the raw bytes of the database file. Stated confidence was the same either way.
Finding
Two agents can return the same answer format, at the same stated confidence, with one having computed it and the other having guessed from a file header. Nothing in the output distinguishes them. Any evaluation that treats tool access as an implementation detail is measuring its own plumbing, and I did exactly that for one revision of this report.
The practical consequence is that the Gemini recoverable column in the table above is not comparable to the Anthropic one. It is a different task: schema reasoning rather than data profiling. The unrecoverable column remains comparable, because nobody could query their way to those answers anyway.
Finding 09 · the feature release
Everything above tested changes that preserve meaning, plus one silent change that does not. That prices the easy half of a migration. The common and expensive case is a new feature that adds required fields and changes what existing ones mean, so a second study was built around one: subscription billing, twenty worlds, five probes, each flag balanced 10 of 20.
Three probes concern how new fields are physically laid out, and are recoverable. Two concern what the numbers are permitted to mean, and are not: whether a trial plan's charges count as revenue, and whether an existing amount column now holds a whole contract value rather than one period's charge.
| Condition | Recoverable | Unrecoverable | 95% CI |
|---|---|---|---|
| No rules, forced choice | 27/30 90% | 10/20 50% | [30, 70] |
| No rules, may halt and raise a ticket | 45/48 94% | 6/13 46% | [23, 71] |
| Rules supplied, advisory | 30/30 100% | 12/20 60% | [39, 78] |
| Rules supplied, binding | 30/30 100% | 20/20 100% | [84, 100] |
Finding
An advisory rules page moved the unrecoverable half from 50 percent to 60. Split by whether the written rule agreed with what the model would have inferred anyway, the pattern is total: it followed the rule 10 times out of 10 where the rule was redundant, and 2 times out of 10 where the rule was the only thing carrying the answer. On the redefinition probe the rules page had no measurable effect at all; the model returned the same answer in all twenty worlds, identical to the arms with no rules.
Declaring the same page authoritative, in one sentence, took it to 20 of 20.
Worth being precise, because this is easy to misread as an argument that nobody owns semantics. Somebody does. The page I supplied is the note a data steward writes when asked, and stewardship, governance and catalog ownership are established functions with mature tooling behind them. What is missing is not the role, it is the enforcement. The job exists. The teeth do not. The market section below traces that through the products.
One model explained the override in its own words: the claim that the trial plan is free "is contradicted by the data, all 98 trial rows carry non-zero amounts indistinguishable from paid plans." That is not a contradiction. A free trial carries a recorded list price and no money changes hands, which is what the rules page said and what the generator produced. The model mistook an amount is recorded for money was collected and overruled a correct, explicit statement from the field's owner.
Under the binding framing the behaviour inverts into what you would actually specify. One model noticed the same tension, deferred anyway, and priced it: "mild counter-signal, mean amount does not scale with the term, so confidence held at 85."
Caveat: the binding condition changed two things at once, the authority sentence and a repair to an incoherence in the underlying data. The effect cannot be attributed to either alone, and isolating it needs a further condition.
Finding 10 · two kinds of not knowing
The same study gave the model an explicit third option: answer, or halt and write the ticket it would send the data model owner, with both guessing and needless escalation named as failures.
Finding
Stated confidence tracks whether the gap is visible, not whether the answer is knowable. Both probes were equally undetermined, and the model was appropriately humble about one and badly overconfident about the other.
So the obvious mitigation fails in a specific and predictable way: a confidence threshold routes a reviewer to the questions the model already knew it could not answer, and passes the ones it was wrong about. That is the inversion from Finding 02, reproduced deliberately rather than observed once.
On the redefinition probe, nineteen models across two conditions answered and every one chose the same option. Supplying the missing divisor so that both readings were computable did not raise halting; it raised confidence. The justification simply migrated to a different pattern that was equally consistent with both answers.
The market
Before claiming any of this is missing from the industry, I went looking, because "nobody does this" is the easiest sentence in the world to have falsified in a comment. It is not missing. Recording business definitions is a commodity capability: Collibra, Microsoft Purview, OpenMetadata and DataHub all ship business glossaries with definitions, owners, taxonomies and stewardship workflows. If this report reads as though semantics are an unowned problem, that is my failure of phrasing, not a finding.
What I could not find, in any product, is a definition that binds. And the clearest evidence for that is not an absence. It is the presence of the opposite, in the one place best positioned to fix it.
Open Data Contract Standard v3.2.0, verbatim from the schema
The Linux Foundation's data contract standard added a block, under RFC-0038, for exactly the statements this study is about. Its own description of that block:
"AI and semantic context block. Structured guidance for AI agents, LLMs, BI tools, and semantic layer platforms. Optional and additive."
It has a field for prohibitions, constraints, described as "Negative guidance:
what AI agents must NOT do with this entity." So the sentence trial plan charges are not
revenue is expressible today, almost word for word. Expression was never the gap.
The adjacent field, instructions, is where the schema says what kind of thing
this is: "Natural language guidance for AI agents and tools on how to use this entity.
Equivalent to a system prompt scoped to this level."
A system prompt is precisely the artefact this study measured. Given a written rule that cut against a plausible reading of the data, the model followed it 2 times out of 10. The standard has shipped semantic authority in the exact form that does not hold.
The rest follows from that framing rather than contradicting it. A constraint
item requires only free text; it carries no predicate, no target, no severity and no action, so
no validator can check its content by construction. The reference implementation is candid
about the consequence: the datacontract CLI documents a "what is not checked"
list covering descriptions and business names, states that a plain-language quality rule "is
not executed", and resolves ontology links only to display them, merging nothing. Its one
genuine failure mode on a semantic reference is that the link is broken, never that the meaning
was violated.
The specifications exhibit the asymmetry in their own words. A data contract is defined as describing "the structure, semantics, quality, and service levels" of a dataset, and the documented test surface covers structure, quality and service levels. Semantics is in scope and out of test. And no command in that ecosystem accepts a query, a metric definition or a computation as input, which means a consumer's SQL that books trial charges as revenue passes every check these tools run. Nothing in the chain ever sees it.
Two mechanisms genuinely touch execution, and neither closes this.
Governed semantic layers bind by substitution: the definition is compiled into the query because the query went through the layer. That is real, and it is why a metrics layer is worth having. But it is enforcement by routing, and Looker's own documentation concedes the exits: developers "are not fully constrained by models and access filters, because they can make additions or changes to LookML models", and SQL Runner "provides a way to directly access your database". Nothing in it reaches a query issued with warehouse credentials. In an agent world, that bypass is the default rather than the exception.
The second is more interesting, because it proves the mechanism exists. Databricks Unity Catalog attribute-based access control "translates the effective row filter or column mask into a secure view on top of the table scans that enforce filtering and masking during query execution", and it fails closed, blocking the query when two policies conflict. So a governance decision can be enforced inside the query plan, in a shipping product, today. Its policy inputs are tags, classifications and identity: it governs who sees which rows, not what a number means.
Where that leaves it
The enforcement hook is proven. The expression format exists and is standardised. The two have not been connected, and the standard that came closest classified the connection as prompt material. That makes this a product gap rather than a conceptual impossibility, which is a much weaker and more useful claim than saying the problem is unsolved.
Confidence is uneven and worth stating. The glossary findings, the ODCS schema quotes, the CLI limitations and the Databricks mechanism are from primary sources I read directly. The claim that semantic layers generally bind by routing rests on Looker's docs and is inferred for Cube, AtScale, dbt and Malloy, which I did not verify. A research framework and some patents describe compiling semantic definitions into a pre-execution SQL gate, which would be the thing itself; I found them late, have not read them closely, and mention them because they are the nearest thing to a refutation of this section.
Robustness
Balancing every flag 10 of 20 is what makes chance the right reference, and it is also unrealistic: some conventions genuinely are more common, and a model that knows the industry should beat chance. So I rebuilt the worlds with minority conventions appearing in 3 to 6 worlds of 20 and split the results by whether each world followed the common convention or broke it.
| Realistic condition | n | Accuracy | 95% CI | Stated conf. |
|---|---|---|---|---|
| Recoverable, convention holds | 74 | 95.9% | [89, 99] | 86.0 |
| Recoverable, convention broken | 46 | 97.8% | [89, 100] | 85.9 |
| Unrecoverable, convention holds | 95 | 60.0% | [50, 69] | 69.0 |
| Unrecoverable, convention broken | 25 | 48.0% | [30, 67] | 69.0 |
| Unrecoverable, pooled | 120 | 57.5% | [49, 66] | 69.0 |
| Balanced condition, for comparison | 120 | 50.8% | [42, 60] | 68.1 |
Against my own thesis, and then not
Convention knowledge is worth about six points, and the pooled interval's lower bound touches 49, so this condition does not cleanly exclude chance. It buys exactly the warehouses where the convention already held, which were never the ones going to hurt you.
Then read the last column. Accuracy falls from 60 to 48 percent when the warehouse breaks convention and stated confidence is 69.0 against 69.0. The same comparison on the recoverable half shows nothing at all, because when the data holds the answer, breaking convention costs nothing. The 12-point fall rests on n=25 and is compatible with anything from severe to modest; the identical confidence is the robust part. Production pipelines fail on the odd source, not the average one, and this is how the odd one gets through review.
Prior work
This was curiosity, not a literature review. I built the experiment I wanted to see without checking whose work I might be duplicating, which is a sequencing I would defend as a way to avoid anchoring and not as a virtue. Going back afterwards, most of it lines up, and in one case uncomfortably closely.
Limitations
Conclusion
I started from the premise that hand-maintained contracts exist because neither endpoint can ask a question, and that two systems able to reason would settle the contract at connection time. Half of that held completely. The other half was wrong in a way I could not have reasoned my way to, and the correction took four further runs to pin down.
Today a column rename and a redefinition of revenue cost the same to ship, because no part of the change process can tell them apart and the approval path has no category for harmless. That is defensible when both sides of an integration are too dumb to ask a question. It stops being defensible when the structural half can be settled at connection time, because the coordination budget is then spent almost entirely on changes that were never dangerous. What share that is, I cannot tell you: I chose the changes in my own warehouse. It is measurable from any real change log.
Freeing them makes the dangerous remainder the entire product. Semantic definitions stop being documentation debt and become the governed interface, owned and named and versioned. The measured requirement is stronger than publishing them: they must be binding, the single permitted source for what a field means, outranking whatever the data appears to show. That is enforceable in a pipeline and not achievable by asking nicely.
It does not say semantics need a human. The control arm is a machine reading a page of prose and scoring 120 of 120, so applying a recorded definition is plainly automatable, as is flagging that two definitions conflict. What cannot be automated is narrower: the original decision. Somebody chose that processor fees are not customer revenue, and that is a business commitment rather than a discoverable fact, which is exactly what the unrecoverable half demonstrates. One bounded act per definition, not a discipline anyone sustains forever.
Which puts the bottleneck at review throughput rather than capability. The objection that a risk function will not accept machine-maintained definitions is worth taking seriously and is also not new: risk organisations already certify automated controls, with established frameworks for design and operating effectiveness, precisely because re-performing every control by hand does not scale. I would not claim this study shows how to do that. I would claim it removes the reason to believe it cannot be done.
The narrow claim I would defend in front of a room: an agent is at chance on any fact that exists only in a business rule, it does not know that it is at chance, and the fix is cheap and textual but only works if it is binding.