ClauseLine logoClauseLine™

Data architecture

Benchmark methodology

Every number published on the benchmark pages is a cell: one statistic (a percentile, mean, or count) for one specialty, one geography, one metric, one source, one year. This page describes where cells come from and the rules every cell must clear before it is published.

Four data tiers

Tier 1 — National benchmarks

National benchmarks

The primary tier: curated national physician compensation benchmarks (2025), multi-source, maintained per specialty. Percentile ladders (p10–p90) cover total compensation, base salary, bonus, annual wRVUs, pay per wRVU, and on-call pay for full-time incumbent physicians, with newly-hired and academic ladders published separately so an employer-type qualifier is always visible.

The public-record and contract-verified tiers below corroborate this tier and will progressively replace it as the contract-verified layer grows.

Tier 2 — Public-record derived

Public record

Cells in this tier are computed from U.S. government bulk files and APIs. Three families are in use:

  • CMS Medicare claims-derived distributions — per-clinician service volume, payment, and productivity proxies from the CMS provider utilization public use files.
  • DOL labor condition filings — offered wages that employers disclosed in public H-1B labor filings, one of the few places a real offered physician salary appears in a public record.
  • BLS Occupational Employment and Wage Statistics — state and national wage percentiles for physician occupations.

Each family measures something different. A Medicare volume proxy is not a salary; an offered wage in a labor filing skews toward specific employer types. Every cell carries its source label so the two are never blended silently.

Tier 3 — Contract-verified

Contract-verified

De-identified data points contributed with explicit consent from real analyzed contracts. This is the highest-signal tier: each point comes from an executed or offered contract, not a survey response. The layer grows as physicians opt in; cells publish only after clearing the record floor below.

Tier 4 — Self-reported

Self-reported

Comp-check submissions typed in by visitors. Lowest tier: unverified, self-selected, and treated accordingly. Self-reported points are never mixed into public-record or contract-verified cells; they form their own cells with their own tier label.

k-anonymity floors

No published statistic rests on fewer than 5 underlying records. Local cells (state or ZIP3 level) publish at 5 or more records; broader public cells publish at 25 or more. Every cell carries its record count (n) next to the number, so a thin cell is visible as a thin cell. Cells below the floor do not exist in the published table at all — they are not hidden, they are never written.

One-way publication

Benchmarks flow one direction: to physicians. Aggregated cells are published on this site. Row-level data — a contributed contract data point, a comp-check submission — is never sold, licensed, or shared with employers, staffing companies, recruiters, or data brokers. There is no employer-side product built on this data.

Derivation versioning

Each cell stores a derivation record: the source file or API, the filter and mapping steps, and the computation that produced the number. When a source file updates or a derivation changes, the cell is re-computed and the derivation record is replaced with it — a published number can always be traced to the exact recipe that produced it.

What this is not

These cells are reference points, not appraisals. An offered-wage percentile from labor filings answers a narrower question than a full compensation survey; a Medicare volume proxy describes productivity, not pay. The cell labels state which question each number answers. Where a number is contested, the source is cited on the cell itself.