← Atlas Commons
● Methodology & Known Limits

How the data is made — and where it stops.

Every atlas is an aggregation of named public sources, run through the same pipeline: fetch the raw source, clean and type it, shape it for serving. No atlas invents data. This page is the pipeline, what one row means in each atlas, and — stated plainly — the named limit of each one. The limits are not buried; they are the point.

Three sibling surfaces, three jobs: this page = how the data is made, its grain, and its limits · /trust = how to verify one specific number · /02-gap-audit = per-table build status.

● THE FOUR-STATUS LABEL — every figure carries exactly one

RECORDED an authority logged it (it occurred) · MODELED an authority computed it (a score/projection, not a measurement) · DERIVED our arithmetic over the above (a rank/ratio; inputs cited) · INTERPRETATION our reading (lowest warrant). Full standard + how to check each: /trust.

The pipeline — bronze, silver, gold

Each source moves through three layers in a medallion lakehouse. The layer a figure comes from tells you how much has been done to it between the federal record and your screen.

● Bronze

Raw source

The source fetched exactly as published — same rows, same values, nothing changed. The audit trail back to the authority.

143 tables
● Silver

Cleaned & typed

Deduplicated, typed, standardized (state codes, dates, units). Errors in the source are surfaced, never silently patched.

87 tables
● Gold

Serving shape

The shape a page reads: aggregated to its grain, with each figure carrying its four-status label. What you actually see.

59 tables

Catalogued snapshot, 2026-06-19: of 289 catalogued tables, fewer than half (132) are live — 106 are planned and 45 are blocked on a source. The pipeline above describes how a figure is made once it's live; it does not imply the whole catalog is. The full per-table ledger is the coverage manifest.

What one row means — and the limit — per atlas

The single most important methodology fact about any dataset is its grain: what one row actually represents. Get that wrong and every number downstream misleads. Here is the grain, the source, the live status, and the honest limit of each atlas.

AtlasSource & grainStatusNamed limit
CoverageCMS Marketplace PUFone row = a plan in a rating areaLiveThe 30 federally-facilitated-marketplace states only. State-based-exchange states (CA, NY, …) are labeled as such — never shown as zero. Plan/benefit metadata, not claims or loss ratios.
CareCMS Care Compareone row = a provider (CCN)LiveThe overall star is a CMS-weighted composite across measure groups (MODELED). A 3-star hospital can be excellent on the one measure you care about and weak on another — cite the underlying measures, not just the star.
PropertyFEMA NRI, NOAA, USFStwo grains: state-level and county-level rows are separate tablesLiveThe recorded storm history is a single year (2025) — do not read it as a multi-year trend, and a single severe event can swing a state's rank. Cross-state death ranks are normalized per-capita where population is known. Insurance-market signals are sparse.
JusticeState statutes, LSC, FJCtwo grains: SOL table (state×claim type), legal-aid table (provider)PartialStatutes of limitations are the law as written — NOT your personal deadline; consult an attorney. The federal civil-caseload gold has a category-mapping bug and is withheld until fixed (not shipped rather than ship wrong).
FinanceHMDA, CFPB, FDIC, GLEIFone row = a lender or institutionLiveThe cross-registry corporation match is a name heuristic (labeled DERIVED) and is superseded by hard-code joins; treat it as a lead, not proof. Insurance non-renewal data is not yet sourced.
ComplianceGLEIFone row = a legal entity (LEI)LiveA ~1,000-entity demo slice of the multi-million-entity global LEI register — not a census; an entity's absence here means "not in our slice," not "no LEI exists." Registration status and structure only, not a verdict on conduct or solvency.
SingularikiBLS OEWS, O*NET, Anthropic Economic Indexone row = an occupation (SOC)LiveBLS OEWS is an establishment survey that lags ~12 months and suppresses estimates for small-population occupations and areas — a missing wage means suppressed, not zero. The Anthropic Economic Index is non-government, labeled where used; federal BLS/O*NET are the authoritative spine.

What is not built yet

Honesty about absence is part of the method. Of the catalogued tables, 106 are planned (specified, not yet built) and 45 are blocked on a source that is rate-limited, messy, or has no clean API (openFDA, DOL Form 5500, insurance rate filings, NMLS bulk). These are named in the coverage manifest as planned or blocked — never silently presented as done. A figure that isn't built is shown as missing, not estimated.

For agents. The structured version of this page — the pipeline layers, per-atlas grain, and named limits — is at /methodology.json. Pair it with /trust.json (the verification standard) and the per-entity .jsonld alternates to assess data quality before citing.

ATLAS COMMONS · METHODOLOGY v1 · machine-readable at /methodology.json · verification standard · coverage manifest · all domains