OfflistMe Data Broker Research Standard & Evidence Methodology
The open-access technical specification for auditing commercial data brokers, evaluating opt-out friction vectors under a 3-State empirical model, measuring structural evidence saturation without artificial baseline floors, and cross-referencing multi-state regulatory filings.
Catalog Universe
1,027
Total active broker profiles
Verified Catalog Records
1,018
Provenance, identity, and route state recorded
Evidence-Gated Profiles
1,018
Identity + qualifying web, email, or documented-form route gate passed; outcomes not inferred
Avg Completeness
42.3%
Captured evidence completeness average
Open Datasets
5
Machine-readable JSON artifacts
Non-Negotiable First Principles
Traditional data privacy rankings suffer from marketing hyperbole, unverified claim recycling, and fabricated statistical precision. OfflistMe enforces five strict scientific rules across all intelligence gathering:
§1.1 Zero Synthetic or Fabricated Data
Every legal name, corporate relationship, registration number, opt-out workflow, and friction observation must cite a verified primary source URL. No placeholders or simulated observations are permitted.
§1.2 Explicit Unknown Representation (TriState)
Unresearched fields must remain strictly UNKNOWN. We never assume a broker requires zero ID or uses CAPTCHA simply because evidence has not yet been collected.
§1.3 No Baseline Floor Inflation
Completeness scoring does not award free baseline points (e.g. giving 60% for identity or 50% for legal without verified records). Unverified records receive 0 points in their respective dimensions.
§1.4 First-Seen $\neq$ Occurred
In temporal event logs, an observation timestamp reflects when our research engine detected and corroborated a change, not the unverified date the change occurred at the broker.
The Evidence Saturation Model
To eliminate vanity metrics, a data broker profile's completeness is calculated through three orthogonal vectors: Structural Completeness ($S$), Evidence Strength ($E$), and Freshness Coverage ($F$).
Master Formula
Overall Completeness Score = round(0.50 × S + 0.35 × E + 0.15 × F)
S (Structural 50%)
0.15·Identity + 0.15·DataPractices + 0.15·Ownership + 0.20·Removal + 0.20·Friction + 0.15·Legal
E (Evidence 35%)
0.60·PrimarySourceDepth + 0.40·RegistryVerificationTier
F (Freshness 15%)
100% (≤90d) · 75% (≤180d) · 50% (≤365d) · 0% (stale/unverified)
| Dimension | Weight in Vector | Verification Criteria | Zero Floor Condition |
|---|---|---|---|
| 1. Identity & Legal Entity | 15% of $S$ | Verified state corporate registration, DBA names, and operating country jurisdiction. | Unverified corporate registrations award zero bonus points. |
| 2. Data Practices & Collection | 15% of $S$ | CPPA Delete Act sensitive data disclosures (minors, geo, biometrics) and input schemas. | 0% if no collection breakdown or schema is documented. |
| 3. Corporate Ownership | 15% of $S$ | Ultimate parent holding company mapped via SEC 10-Ks, CIKs, or state regulatory filings. | 0% for unresearched holding entities. |
| 4. Consumer Removal Routes | 20% of $S$ | Verified direct privacy email (RFC 5322), interactive web portal, or mail suppression route. | 0% if no verified opt-out destination exists. |
| 5. Empirical Friction Vector | 20% of $S$ | All 13 TriState friction barriers evaluated (Gov ID, CAPTCHA, phone tokens, SLAs). | 0% if barriers remain completely unobserved. |
| 6. Legal & Regulatory Status | 15% of $S$ | Active registration under California SB 362, Vermont 9 V.S.A. § 2446, or SEC. | Baseline 10% statutory jurisdiction applicability. |
| 7. Registry Coverage | 40% of $E$ | Cross-referenced against CPPA, Vermont SOS, and SEC EDGAR registries. | Evaluates presence across 0, 1, 2, or 3 official registries. |
| 8. Primary Source Depth | 60% of $E$ | Citations must be valid HTTPS documents/URLs (P1–P3). Bare domain text is rejected. | 0% if no valid source URLs exist. |
| 9. Freshness Coverage | 15% of Overall | Verified review timestamp within 90, 180, or 365 calendar days. | 0% if unverified or older than 1 year. |
4-Tier Primary Source Hierarchy (P1–P4)
Every factual assertion in our database is ranked according to the legal standing and authority of its source document:
Tier P1: Official State & Statutory Registries
Highest AuthorityCalifornia Privacy Protection Agency (CPPA) Data Broker Registry, Vermont Secretary of State (SOS) Registry, Texas Attorney General Registry, and Oregon Department of Financial Regulation (DFR) records.
Tier P2: First-Party Verified Portals & Policies
Direct OperationalDirect consumer opt-out portals, privacy policy disclosures, DSAR submission workflows, and suppression center forms maintained by the broker itself.
Tier P3: Federal Regulatory & Financial Filings
Enforcement & CorporateSecurities and Exchange Commission (SEC) Annual 10-K filings, Federal Trade Commission (FTC) Section 5 administrative orders, and federal court consent decrees.
Tier P4: Independent Empirical Audits
Reproducible TestingDirect end-to-end verification tests conducted by OfflistMe privacy engineers recording latency, CAPTCHA challenges, phone token demands, and confirmation SLAs.
3-State Removal Friction Methodology
To prevent mischaracterization of data brokers, every friction metric is tracked using strict TriState variables with explicit sample size disclosures:
Positively Verified Barrier
Empirically observed in the live opt-out flow and backed by a primary screenshot or submission test record.
Positively Verified Absence
Opt-out workflow completed end-to-end without encountering the specified barrier.
Unresearched State
Workflow has not been manually audited end-to-end. Excluded from sample denominators to prevent artificial zero rates.
Macro Friction Findings Across Sample Cohorts
Direct Email Support
93.5%
952 / 1018 brokers
Web Form Portals
75.1%
765 / 1018 brokers
Gov ID Mandate (FCRA)
99.6%
Sample N=547
CAPTCHA Anti-Automation
%
Sample N=0
Public Machine-Readable Datasets
In accordance with open research principles, all underlying data structures are published as versioned, machine-readable JSON datasets available for researchers, privacy advocates, and academic institutions:
Entity Research Ledger
.JSON (1.1MB)Catalog broker records with 9-dimension Evidence Completeness Profiles, UNKNOWN field arrays, and primary source citations.
Broker Registry Crosswalk
.JSONCross-registry presence matrix joining California CPPA, Vermont SOS, and SEC EDGAR filings for all catalog entities.
Corporate Ownership Graph
.JSON13 major corporate holding networks, mapped brands, SEC stock tickers, and public EDGAR filing URLs.
Removal Friction Benchmark
.JSONMacro channel distribution, empirical sample denominators, and positive barrier rates.
Empirical Research Findings & Insights
.JSON6 bounded cross-entity observations with evidence boundaries, sample sizes, and consumer implications.
Peer-Reviewable Empirical Findings
Cross-entity analysis of the captured evidence cohort and multi-state registry snapshots yields bounded observations; it does not establish universal industry rates:
Government ID Requirements Are Limited to a Documented Evidence Cohort
Observation: Among 1018 source-backed broker profiles, 545 explicitly mention a government-ID requirement and 471 remain unknown.
Empirical Evidence: 545 of 1018 evidence-ready profiles have an explicit YES observation. The dataset does not infer FCRA status or generalize the result to brokers outside this cohort.
Consumer Implication: Check the current first-party workflow before sending identity documents; UNKNOWN is not a NO result.
Geolocation Disclosure Is Reported Only Where the Registry Evidence Names It
Observation: 0 evidence-ready profiles have an explicit geolocation category in the captured registry evidence; enforcement status is not inferred from that category.
Empirical Evidence: 0 of 0 profiles also include an FTC URL among their recorded sources. This is a source-presence observation, not proof of an active order.
Consumer Implication: Read the linked regulator action and its scope before drawing a legal conclusion about a broker.
B2B Workflow Observations Are Kept Separate From Trend Claims
Observation: 0 of 155 profiles with sales-oriented processing language also have a YES web-form observation.
Empirical Evidence: The cohort does not contain a source-backed average turnaround SLA, so no timing comparison is published.
Consumer Implication: Workflow channel observations do not establish speed, success, or a market-wide trend.
Some People-Search Workflows Document Profile or Record Lookup
Observation: 0 of 2 profiles in this category have source-backed required-field text that mentions a profile, record, listing, or URL.
Empirical Evidence: The result is limited to explicit workflow evidence and does not claim that every people-search site requires URL extraction.
Consumer Implication: When a current source asks for a record identifier, preserve that requirement in the user workflow.
Recorded Parent-Entity Relationships Form a Bounded Ownership Cohort
Observation: 44 configured parent mappings cover 65 catalog brand/domain entries.
Empirical Evidence: This is a bounded internal mapping, not a complete ownership census; relationships require source-level review before publication as corporate fact.
Consumer Implication: A parent relationship does not by itself prove that a rights request cascades across sibling brands.
Captured State Registry Snapshots Do Not Fully Overlap
Observation: The captured California and Vermont registry matches overlap for 79 catalog entities; absence from one snapshot is not proof of non-registration everywhere.
Empirical Evidence: The crosswalk contains 623 California matches and 202 Vermont matches in the captured snapshots.
Consumer Implication: Always cite the jurisdiction, snapshot date, and matching method when comparing registry coverage.
Independent Audit & Verification
All research calculations, validation checks, and dataset generators are checked into the OfflistMe repository and executed on every continuous integration build:
# 1. Run full test suite across 34 test files (222 tests)
npm test
# 2. Rebuild and synchronize all 5 research datasets
npm run generate-entity-research-ledger
# 3. Validate ledger and public research dataset integrity
npm run check-entity-research-ledger
