A Nobel-winning napkin, and a family in a hospital hallway.
There is a 1970 paper called The Market for Lemons. George Akerlof won a Nobel for it. The whole argument fits on a napkin: when buyers can't tell good from bad, the good gets driven out, because why would you pay for quality nobody can verify? The used-car lot fills with junk. The honest seller goes home. The market doesn't just fail to reward quality — it actively punishes it. Information asymmetry isn't a friction. It's a wrecking ball.
Now look at American home health. Twelve thousand agencies. A family standing in a hospital hallway at 6pm with a discharge packet and a dying phone battery, asked to choose, with nothing but brochures, a star somebody computed, and a Google rating gamed by the agency's own front desk. That is a lemons market in its cruelest form. The honest agency that staffs heavy and gets people walking again has no way to prove it to the one person who needs to know. So it competes on the same axis as the agency that bills Medicare and parks grandma in front of a TV: who has the better pamphlet.
And then you open a public file and find a column that says, in flat affectless government prose, how often patients got better at walking. 88.86%. Measured. Per agency. For the entire country. That column is a receipt. It is the thing Akerlof said markets can't have. And the second a family — or an agent acting for one — can read it, the lemons market inverts.
Watch the market invert
You think it's a number. It's a verb.
Every metric you have ever seen is a noun. Revenue. Headcount. Beds. Latency. Nouns sit there; they describe a state. You can stack them, average them, dashboard them, and at the end of the day you still don't know if anything happened.
An outcome is a verb. Got better. It encodes a change in a human being across time — a before and an after, a person who couldn't walk and then could. The instant your data has a plot — a delta, a transition — you are no longer doing analytics. You are doing physics. You can take derivatives. You can ask what caused the change, because there is now a change to cause. That's why outcome data is rare: measuring a verb means catching the same human twice and differencing them. CMS did exactly that — admission and discharge, scored with OASIS, differenced. The column isn't a snapshot. It's a time machine with the before and after folded into one percentage.
Nouns describe a state and tell you nothing happened. Verbs encode a change and hand you the calculus.
The 2024 staffing rule was a national experiment. Nobody collected the prize.
In April 2024 CMS finalized a minimum-staffing mandate for nursing facilities — the first ever. A policy dropped, on a date, on some facilities harder than others (rural waivers, some states already above the floor, some brutally below). A discontinuity in time, applied unevenly across entities. That is the textbook setup for difference-in-differences: compare the facilities forced to staff up against those already above the line, and watch what happens to the outcomes before vs. after. The change-in-change is a causal estimate of what mandated staffing does to whether residents recover.
Difference-in-differences — drop the policy
And it compounds. Every methodology change, every reimbursement tweak, every star recalibration is a natural experiment with a date stamp — event studies stacking up quarter after quarter. The atlas isn't a directory. It's a standing observatory pointed at every policy shock in American healthcare, holding the one instrument — measured outcomes — that lets you read the dial.
The question behind the question: not whether, but how long.
"How often patients got better" is a scalar squeezed from a richer object: time-to-event. How long until recovery? Who recovered fast, who slow, who never? Think in time-to-event and you've imported the whole arsenal of survival analysis — Kaplan-Meier curves, hazard rates, the machinery oncology uses to ask "how long do people live." A scalar tells you the average ending. A survival curve tells you the shape of the journey — and for a human deciding where to send their parent, the shape is the whole thing. A 5-star with a flat early curve is a slow-motion tragedy dressed as quality.
The outcome is a currency. Now you can arbitrage.
Every domain measures "good" in its own incommensurable unit. Healthcare has stars. Lending has APRs. Insurance has actuarial value. Justice has caseloads. None convert. You cannot ask "is a point of recovery worth more than a point of mortgage approval" because the units don't touch.
An outcome is the universal denominator — the unit every domain secretly shares, because every domain exists to change a human's state. Reduce each to the change in the human and they become commensurable. And the instant things are commensurable, you can arbitrage them: where is the cheapest marginal outcome? A dollar of home-health staffing buys X recoveries; a dollar of legal aid buys Y evictions prevented; a dollar of insurance navigation buys Z covered treatments. An institution that prices everything in outcomes routes capital, attention, and agents to the highest outcome-per-dollar across domains that have never been able to talk. That's a new asset class: the outcome itself, made liquid.
The next economy bills for results, not motion.
Today everything is billed on actions. The agency bills per visit, the lawyer per hour, the doctor per code. The original sin of services: you pay for motion, and motion is not the thing you wanted. The entire fee-for-service catastrophe — over-treatment, churn, the incentive to do more rather than better — is downstream of the fact that nobody could measure the result, so we paid for the motion as a proxy and prayed.
Outcome data kills the proxy. When you can measure whether the patient walked, you can write a contract that pays for the walking. Outcome-based contracts. Pay-for-recovery. And the agent-native twist: an autonomous agent can be the counterparty — representing a family, negotiating, monitoring, settling against a public receipt both sides can re-read. Fee-for-service was a workaround for blindness. We just got eyes.
Routing by outcome doesn't describe the market. It rewrites it.
Today provider behavior has no fast feedback from outcomes, because outcomes are invisible to buyers — an agency can have terrible recovery rates and thrive for years on marketing and inertia. Close the loop. Families (via agents) route by outcome → better-outcome agencies get more patients → worse-outcome agencies feel it in revenue → they staff up and improve → outcomes rise → the panel records it → families route to them. A reinforcing loop where the measurement itself becomes the selection pressure that makes the whole sector better.
This is the deepest thing outcome data does: it's not just a description of quality, it's a cause of it. The systems lens in the repo — stocks, flows, feedback gains, the delays that give the flywheel inertia — lets us model the whole dynamical system and ask not "which agency is good" but "is the system getting better, and what's the constraint slowing the wheel?" Theory of Constraints on a living, outcome-driven market.
The silence is data: what the suppressed 36% are screaming.
Thirty-six percent of home-health agencies have no rating, no measured outcome. The naive move is to drop them. The epistemically alive move is to realize the absence is a signal — CMS suppresses measures when an agency has too few cases to report reliably, so a missing outcome means small, new, or low-volume, itself a yellow flag for a vulnerable patient. This is censoring and selection, exactly where amateurs lie and professionals get careful. Rank only the agencies that have outcomes and your whole leaderboard is biased toward the big and established. The honest surface says three different true things — measured 89% · too few cases · measured but unreliable — and never collapses them into one "N/A." The most honest thing on the page is sometimes the asterisk.
Goodhart is coming for us. The four-status is the armor.
The instant outcomes drive money, outcomes get gamed. Goodhart's law in scrubs: cherry-pick the patients likely to recover, discharge the hard ones early, up-code the admission so the "improvement" looks bigger. Every outcome system in history has been gamed, and a single-number leaderboard accelerates it by pointing a fire-hose of incentive at one figure.
Our architecture is anti-Goodhart by construction. Cross-checks across the four-status: a suspiciously high recovery rate (RECORDED) that doesn't square with the staffing hours (RECORDED) is a tell — gaming leaves fingerprints, and an agent can check internal consistency because every measure declares what it is. The modeled-vs-measured split keeps risk-adjustment honest. And re-runnability means the adjustment is auditable. A single metric is a Goodhart machine; a panel of clothed, cross-checkable, re-runnable outcomes is a Goodhart detector.
Not a citation. A chain of custody.
That 88.86% was not typed by an analyst. It begins with a nurse, in someone's living room, on day one, scoring whether the patient can cross the room — the OASIS instrument, a human judgment about a human body. Then again at discharge. Two keystrokes, weeks apart, about one frightened person, times every patient that agency saw. CMS aggregates, risk-adjusts, publishes; we ingest to bronze, reshape through silver, project to gold, clothe in the four-status, hand to an agent deciding where a stranger's mother recovers.
An agent-native institution should narrate that on demand — not "trust us, 88.86%," but the whole chain from the nurse's keystroke to this pixel, walkable backward by anyone. That's what provenance means when the stakes are a human outcome: not a footnote, a chain of custody. Evidence isn't decoration here. It's the load-bearing wall.
A causal claim you can re-run is the death of authority.
A causal claim about an outcome — "heavier staffing causes better recovery" — is the most powerful and most dangerous sentence an institution can utter. Powerful because it tells you what to do. Dangerous because for all of history you believed it on authority: an expert said it, a journal blessed it, an auditor signed off. Authority is exactly what an agent-native institution is built to abolish.
So the endgame: a causal outcome claim ships not as a conclusion but as a re-runnable computation — the panel, the spec, the controls, the code — so no agent ever trusts "staffing causes recovery." It re-derives it. That's M2, continuous universal auditability, no privileged auditor, applied not to a Lean theorem but to a causal claim about whether human beings recover.
Re-run the claim
An outcome you can re-derive is an outcome no one can lie about. A causal claim you can re-run is the death of authority as the basis for "what's good for people."
This is the whole political philosophy, finally touching ground.
An institution whose atoms are outcomes — verbs not nouns; receipts that un-break markets; a currency that makes every domain commensurable; a flywheel that evolves the providers it measures; an immune system against gaming; a chain of custody from a nurse's keystroke to an agent's decision; and a constitution where even "this is good for you" must be re-runnable by anyone who doubts it.
It was hiding in a 96-column government file, in one column, in flat affectless prose. How often patients got better at walking: 88.86%. The grail had a footnote the whole time. Let's go build the rest of it. gg.
care_gold.home_health_serving_surface · 12,392 agencies · outcome-grade · four-status · CC0