In 2018 our vendor platform upgrades kept reaching production with problems, and because that is what everybody expects the cause to be, the first several conversations about it were conversations about testing rigor. The evidence did not support them. Teams were testing carefully and signing off honestly, but they were signing off on whatever they had chosen to test, and what they chose was reliably whatever they had built most recently, because nobody anywhere in the organization had a list of what actually existed.

We did not have a quality problem. We had a map problem.

Nobody has the list, and that is the normal condition

It would be easy to read that as an indictment of somebody, and it is not. A platform of any age is built in increments, each of which is owned attentively by the team that built it, and none of which obliges anyone to describe the whole. The increments are deliverables. The census of increments is not a deliverable for anyone, which means it gets produced, if at all, as a by-product of some other effort, and a by-product is exactly the kind of artifact that is accurate on the day it is written and quietly wrong six weeks later.

That is also why the inventories organizations do have tend to rot. They are descriptions, and a description has no natural defender. Nobody is embarrassed when a page listing components falls out of date, because the page was never a claim about anybody's responsibilities, only a snapshot of somebody's understanding. The moment a list carries an accountable name beside every entry, its character changes: it stops being a description and becomes a set of commitments that people will correct when they are wrong, because being listed as the owner of something you do not own is a problem the person listed has an incentive to fix.

The boundaries are the work, and the coverage is the easy part

So I ran the analysis to build the list properly. We enumerated every experience the platform supported, decomposed each of those into the capabilities that composed it, classified those capabilities by criticality, and assigned an accountable owner to every single one. Stated that way it sounds like an afternoon of tabulation. What it actually required was cross functional alignment from all contributing teams, capacity to put it all in place and ensure it was correct, and considerably more patience than the enumeration itself, because the hard part was never the list. The hard part was getting people to agree on where one team's ownership ended and another's began.

That asymmetry is worth sitting with, because it generalizes well beyond upgrades. Test coverage is a technical question with a technical answer, and technical questions with technical answers are the ones organizations are already good at, since you can settle them with evidence, in a room, in an afternoon. Ownership boundaries are not that. An ownership boundary is a question about who gets called when something fails at an hour nobody wants to be awake, and every team in the conversation understands perfectly well that agreeing to a line on a diagram is agreeing to be called. So the discussion moves slowly, and it moves slowest exactly at the seams, where a capability is jointly touched and separately understood and has therefore never been anybody's in particular.

Those arguments were not an obstacle to the work. They were the work. An unowned seam is not a documentation gap, it is where the upgrade defects live, because it is the one part of the system that no team's testing covers and no team's sign-off represents. Every boundary we argued our way to a decision on was a class of production problem we were retiring, and the arguments were slow because they were consequential, not because anyone was being difficult.

What we set out to do, and what we got instead

The upgrade problem went away. Four consecutive vendor upgrades went live clean, which was the entire objective and, at the time, felt like the end of the story.

It was not. Release testing changed shape almost immediately, because teams could finally reason about a component and its dependencies instead of inferring scope from whatever they happened to have touched most recently. Documentation improved across the organization without anyone mandating it or funding it, for a reason I found more interesting than the improvement itself: once a team owned something by name, in a list other people could read, they wanted it described properly. Roadmapping improved through the same mechanism, since a roadmap is far easier to argue about when the things being sequenced have agreed names and agreed owners. And when we added two teams in 2020, and again when we moved from a pure Spotify model to a Spotify and SAFe hybrid, the ownership reassignment took days rather than quarters, because the thing being reassigned already existed as a list and the conversation was therefore about who, not about what.

None of that was in the business case. It could not have been, because I did not know it was coming, and I want to be careful not to reconstruct it afterwards as foresight. But the fact that those second order returns were unforecastable is itself the useful observation, because it means the returns cannot serve as the selection criterion. If you only fund organizational work whose full benefit you can name in advance, you will systematically fund the fixes and systematically decline the things that become infrastructure, since the defining property of infrastructure is that its most valuable uses are the ones nobody has thought of yet.

Why it outlived the structures

That turned out to be the lesson, and it took me years to see it. The artifact outlived every problem it was built for and every org structure it was built under. The reorgs kept coming and the map kept being the thing everyone reached for.

The reason, I think, is that the map described the platform rather than the organization. An org chart is a claim about people, and it is redrawn whenever leadership, strategy or headcount changes, which is to say more or less constantly. A list of capabilities is a claim about what the system does, and a capability does not stop existing because a box moved above it. A reorg cannot invalidate a document of that kind. A reorg can only need it, which is the opposite relationship, and it is the whole reason the thing got more valuable each time the structure changed rather than less.

So I have come to think this is the real test of organizational work: not whether it solves the problem in front of you, but whether it is still load-bearing three structures later, when the people who built it have moved on and nobody remembers why it exists. By that standard the useful question to ask of a proposed investment is not how much it will return, but what it names, and whether those names will still be true after the boxes have been redrawn. Most of what we fund fails that question in the first direction: it changes how a specific group of people work right now, and it becomes meaningless the moment that group is dissolved.

The map passed the test in a way I find slightly humbling. It is used now by people who have no idea it began as a fix for vendor upgrade defects, who would not recognize the argument that produced it, and who reach for it because it is simply how the platform is described. That anonymity is not the artifact being forgotten. It is the only evidence that ever really counted.