Methodology

Last reviewed: September 16, 2026

This page describes how source data becomes a published dataset or figure on NumberResearch.org. The goal of the pipeline is traceability: every published number should be reproducible from an identified source file with a known date.

The pipeline

  1. Daily ingestion. We retrieve the official source files listed on our Data Sources page on their stated check schedules — daily for assignment databases, event-driven for planning letters and reports.
  2. Immutable snapshots. Every retrieved file is stored as an immutable, dated snapshot before any processing. Snapshots are never edited; a new retrieval creates a new snapshot. This is what lets us say exactly what a source said on a given date.
  3. Normalization. Snapshots are parsed into a consistent structure: field names, code formats, date formats, and status vocabularies are standardized across sources. Normalization changes representation, not content; the original snapshot remains the reference.
  4. Diff. Each normalized snapshot is compared against the previous one to produce a change record: new assignments, status changes, reclamations, and corrections in the source itself.
  5. Human review of sensitive changes. Routine changes flow through automatically. Changes flagged as sensitive — unusually large diffs, changes to previously stable records, apparent source errors, or anything feeding a research conclusion — are held for human review before publication, per our AI & Automation Policy.
  6. Publish. Reviewed data is published with the source file date attached. Published pages state which snapshot they were computed from.

How counts are computed

Counts on this site are computed dynamically from source files and are labeled with the source file date. We do not maintain hand-edited totals. If a count on this site differs from a count elsewhere, the first things to compare are the file dates and the counting rules (for example, whether a total includes codes in all statuses or only assigned codes; our dataset pages state the rule used).

Important caveat: assignee is not current carrier

The initial assignee of a numbering resource is not necessarily the current carrier of any individual number, due to number portability. Assignment databases record which carrier a code was assigned to; individual numbers within that code may since have been ported to other carriers. Nothing on this site should be read as identifying the current serving carrier of a specific telephone number.

Known limitations

Questions about method

If a figure looks wrong or a counting rule is unclear, write to research@numberresearch.org. Methodology challenges are welcome and are handled under the Corrections Policy when they identify an error.