We had CI runs that took 30 to 40 minutes once tests and scans finished, and a frustrating share of them ended by surfacing a missing dependency, something the developer could have anticipated before committing and routinely did not. The cost was never really the mistake. Missing dependencies are trivial to understand, trivial to fix, and nobody needed a diagnosis or a training session to avoid them. The cost was finding out about it forty minutes later, after attention had moved somewhere else and the work of reconstructing what you had been thinking had to be paid all over again.
Earlier this year we put AI into the delivery pipeline across a twelve team monorepo, covering code review, security scanning, dependency and package validation, style conformance and test coverage. Dependency validation was one check among several, and on any reasonable ranking of technical interest it was the least impressive of them. It returned the most anyway. The reason had nothing to do with it being the hardest problem in the set. It had the longest feedback loop. So we moved that check left, ahead of the commit, where it runs in seconds rather than costing forty minutes to fail.
That gave us roughly two hours back per developer per release cycle and cut pipeline reruns by about a third. Those numbers came in quickly, because a narrow check with a clear before and after is easy to attribute to one change. Nothing else was moving in the same window that could plausibly explain them, the mechanism was obvious, and the direction of causation was not in dispute. If I had needed a slide, that was the slide.
The habit most engineering organizations fall into is to rank candidate work by the difficulty or the severity of the underlying problem, on the reasonable-sounding theory that the biggest problems deserve the most attention. That ranking is wrong more often than it looks, because the cost a defect actually imposes is not a property of the defect. It is a property of how long the person who caused it takes to find out.
A subtle problem discovered in seconds is cheap. The developer is still inside the context that produced it, the fix is a local edit, and no other person and no other system has been drawn in. A trivial problem discovered in forty minutes is expensive, and it is expensive in a way that compounds, because the delay does not merely add forty minutes of waiting. It forces a context switch out, a context switch back, and a reconstruction of intent that is rarely as complete as the original. The same defect, sitting in the same line of code, costs two entirely different amounts depending only on where in the loop it surfaces.
Which means the question worth asking about any check is not how serious the thing it catches is. It is how long the current loop is between causing that class of problem and learning about it. Rank by that, and the priorities come out in an order that looks almost perverse next to the usual one: dull, well-understood failure modes that take a long time to surface climb above interesting, difficult ones that surface immediately. Dependency validation won on those grounds and on no others.
There is a second-order effect that matters more than the recovered hours. A loop short enough to sit inside a single unit of attention changes behavior, because developers begin to trust it and stop defending against it. Long loops train people to batch. If finding out costs forty minutes, the rational response is to make each attempt carry as much as possible, which produces larger changes, longer-lived branches, more merge risk and less frequent releases. Nobody decides to work that way. The loop decides it for them, and shortening the loop is the only thing that reliably undoes it.
Here is the awkward structure of the whole exercise. The effects I could measure in weeks were real, and they were the smallest thing that happened. The effects I actually cared about were the ones that could not be measured on that schedule: fewer interruptions, less context switching, enough confidence in the pipeline that teams start shipping smaller and more often. Those run on a horizon of quarters rather than sprints.
Pull request quality was still climbing a couple of months later when my role was eliminated in a reorg. I set the pattern and shipped it into production, and somebody else gets to see where it lands. That is a normal enough outcome for platform work, and worth stating plainly because it illustrates the thing rather than being an exception to it. The evidence for the returns that matter arrives long after the decision, and often long after the people who made it have moved on.
The inversion is uncomfortable. Normal prioritization says to favor the work whose value you can demonstrate, because demonstrable value is how budgets get defended. Applied to platform investment, that rule systematically selects the least valuable available work, because the wins that can be cleanly attributed are by construction the narrow ones. Anything broad enough to change how twelve teams behave is also broad enough to be entangled with every other thing that changed in the same year, and entangled effects do not produce clean numbers.
Compounding returns get credited to whatever else changed in that time. If release frequency improves over three quarters, the pipeline is one candidate explanation among a dozen: a reorg, some hiring, a new architecture, a manager who got better at running a team, an unusually quiet period of customer escalations. Every one of those is a genuine candidate, and there is no honest experiment available that separates them. So the credit goes somewhere, and where it goes has more to do with who is telling the story than with what actually caused the movement.
That is why the costs of this work are always legible and the returns rarely are. The costs land inside one sprint and everyone can see them, because they take the form of engineering time not spent on something with a roadmap item attached. The returns compound over a year and dissolve into the general improvement of things.
So the decision gets made in a window where the evidence structurally cannot exist yet, and the person asking you for a number you do not have is not being unreasonable. They are being asked to allocate against alternatives that do come with numbers, and declining to fund something on the grounds that its case is unfalsifiable is a defensible position rather than an obstructive one. Treating that question as bad faith is the most common mistake in these conversations and it loses the argument immediately, because it asks someone to accept a weaker evidentiary standard for your proposal than for everyone else's.
My answer was to go get the one number I could actually get and buy the rest on judgment. The narrow measurable win paid for the decision, and the reasoning about loop length carried the part that could not be measured. An honest case says which is which: here is the piece I can attribute, here is the mechanism I am asserting, here is the horizon on which you should expect to see it, and here is what would tell you I was wrong. That last part is what makes it a case rather than an appeal, and it is the part most often left out.
I do not think that is the only answer, and I am not sure it is the best one. It has an obvious weakness, which is that it works best when a narrow measurable win happens to be available, and there is no guarantee one will be. The general problem it points at is that our funding processes were built to evaluate work whose returns arrive inside the evaluation window, and platform work does not. Every organization I have seen resolves that mismatch the same way, by trusting a person instead of a number, which works exactly as long as that person is around to be trusted. Building something more durable than that is still open, and it is a harder problem than anything in the pipeline was.