Skip to main content
Home/Privacy Research/Research Standard
Research Standard · Spec v1.4 · Updated August 26, 2026

OfflistMe Data Broker Research Standard & Evidence Methodology

The open-access technical specification for documenting commercial data-broker evidence, evaluating opt-out friction observations under a 3-State model, measuring snapshot completeness without artificial baseline floors, and cross-referencing captured regulatory sources.

Rahul Kandoriya
Written byRahul Kandoriya·Founder, OfflistMe·Last updated August 26, 2026

Research Universe

1,052

Unique research records; not customer-ready coverage

Verified Catalog Records

1,043

Provenance, identity, and route state recorded

Evidence-Gated Profiles

1,039

Identity + qualifying web, email, or documented-form route gate passed; outcomes not inferred

Avg Completeness

42.1%

Captured evidence completeness average

Open Datasets

5

Machine-readable JSON artifacts

01

Non-Negotiable First Principles

Traditional data privacy rankings suffer from marketing hyperbole, unverified claim recycling, and fabricated statistical precision. OfflistMe enforces five strict scientific rules across all intelligence gathering:

§1.1 Zero Synthetic or Fabricated Data

Every published legal name, corporate relationship, registration number, opt-out workflow, and friction observation should cite a source URL and retain an explicit unknown state when the source is incomplete. Synthetic fixtures are kept separate from live catalog evidence.

§1.2 Explicit Unknown Representation (TriState)

Unresearched fields must remain strictly UNKNOWN. We never assume a broker requires zero ID or uses CAPTCHA simply because evidence has not yet been collected.

§1.3 No Baseline Floor Inflation

Completeness scoring does not award free baseline points (e.g. giving 60% for identity or 50% for legal without verified records). Unverified records receive 0 points in their respective dimensions.

§1.4 First-Seen $\neq$ Occurred

In temporal event logs, an observation timestamp reflects when our research engine detected and corroborated a change, not the unverified date the change occurred at the broker.

02

The Evidence Saturation Model

To eliminate vanity metrics, a data broker profile's completeness is calculated through three orthogonal vectors: Structural Completeness ($S$), Evidence Strength ($E$), and Freshness Coverage ($F$).

Master Formula

Overall Completeness Score = round(0.50 × S + 0.35 × E + 0.15 × F)

S (Structural 50%)

0.15·Identity + 0.15·DataPractices + 0.15·Ownership + 0.20·Removal + 0.20·Friction + 0.15·Legal

E (Evidence 35%)

0.60·PrimarySourceDepth + 0.40·RegistryVerificationTier

F (Freshness 15%)

100% (≤90d) · 75% (≤180d) · 50% (≤365d) · 0% (stale/unverified)

DimensionWeight in VectorVerification CriteriaZero Floor Condition
1. Identity & Legal Entity15% of $S$Source-backed state registration, DBA names, and operating-country context.Unverified corporate registrations award zero bonus points.
2. Data Practices & Collection15% of $S$CPPA Delete Act sensitive data disclosures (minors, geo, biometrics) and input schemas.0% if no collection breakdown or schema is documented.
3. Corporate Ownership15% of $S$Ultimate parent holding company mapped via SEC 10-Ks, CIKs, or state regulatory filings.0% for unresearched holding entities.
4. Consumer Removal Routes20% of $S$Recorded first-party email, interactive web portal, or mail-suppression route, with route semantics kept separate from acceptance.0% if no verified opt-out destination exists.
5. Observed Friction Vector20% of $S$13 tracked TriState friction fields (government ID, CAPTCHA, phone tokens, and other workflow observations).0% if barriers remain completely unobserved.
6. Legal & Regulatory Status15% of $S$Recorded presence in a relevant state registry or SEC filing; an SEC filing is not a data-broker registration or privacy-rights determination.Baseline 10% statutory jurisdiction applicability.
7. Registry Coverage40% of $E$Cross-referenced against captured CPPA, Vermont SOS, Texas SOS, Oregon DCBS/DFR, and SEC EDGAR sources.Evaluates presence across 0, 1, 2, or 3 official registries.
8. Primary Source Depth60% of $E$Citations must be valid HTTPS documents/URLs (P1–P3). Bare domain text is rejected.0% if no valid source URLs exist.
9. Freshness Coverage15% of OverallVerified review timestamp within 90, 180, or 365 calendar days.0% if unverified or older than 1 year.
03

4-Tier Primary Source Hierarchy (P1–P4)

Every factual assertion in our database is ranked according to the legal standing and authority of its source document:

Tier P1: Official State & Statutory Registries

Highest internal source tier

California Privacy Protection Agency (CPPA) Data Broker Registry, Vermont Secretary of State (SOS) registry sources, Texas Secretary of State (SOS) Data Broker Registry, and Oregon Department of Consumer and Business Services (DCBS) / Division of Financial Regulation (DFR) registry records.

Tier P2: First-Party Verified Portals & Policies

Direct Operational

Direct consumer opt-out portals, privacy policy disclosures, DSAR submission workflows, and suppression center forms maintained by the provider; route availability does not prove acceptance or deletion.

Tier P3: Federal Regulatory & Financial Filings

Enforcement & Corporate

Securities and Exchange Commission (SEC) Annual 10-K filings, Federal Trade Commission (FTC) Section 5 administrative orders, and federal court consent decrees.

Tier P4: Independent Evidence Audits

Reproducible Testing

When available, independently documented end-to-end observations of route reachability, barriers, and provider-stated confirmation steps; these do not establish mailbox acceptance or an SLA.

04

3-State Removal Friction Methodology

To prevent mischaracterization of data brokers, every friction metric is tracked using strict TriState variables with explicit sample size disclosures:

YES

Positively Verified Barrier

Recorded from source-backed route or workflow evidence for the documented cohort; this does not establish mailbox acceptance, deletion, or an outcome rate.

NO

Positively Verified Absence

A source or bounded observation records that the specified barrier was not present in the documented route; this is not a universal absence claim.

UNKNOWN

Unresearched State

Workflow has not been manually audited end-to-end. Excluded from sample denominators to prevent artificial zero rates.

Macro Friction Findings Across Sample Cohorts

Direct Email Support

94.1%

981 / 1043 brokers

Web Form Portals

76.1%

794 / 1043 brokers

Government-ID requirement field

99.6%

Sample N=545

CAPTCHA Anti-Automation

%

Sample N=0

05

Public Machine-Readable Datasets

In accordance with open research principles, the underlying research structures are published as machine-readable JSON datasets available for researchers, privacy advocates, and academic institutions:

06

Reviewable Research Findings

Cross-entity analysis of the captured evidence cohort and multi-state registry snapshots yields bounded observations; it does not establish universal industry rates:

Finding #1 · FRICTION PATTERNSSample N = 1039

Government ID Requirements Are Limited to a Documented Evidence Cohort

Observation: Among 1039 source-backed broker profiles, 541 explicitly mention a government-ID requirement and 496 remain unknown.

Recorded Evidence: 541 of 1039 evidence-ready profiles have an explicit YES observation. The dataset does not infer FCRA status or generalize the result to brokers outside this cohort.

Consumer Implication: Check the current first-party workflow before sending identity documents; UNKNOWN is not a NO result.

Finding #2 · REGULATORY ENFORCEMENTSample N = 0

Geolocation Disclosure Is Reported Only Where the Registry Evidence Names It

Observation: 0 evidence-ready profiles have an explicit geolocation category in the captured registry evidence; enforcement status is not inferred from that category.

Recorded Evidence: 0 of 0 profiles also include an FTC URL among their recorded sources. This is a source-presence observation, not proof of an active order.

Consumer Implication: Read the linked regulator action and its scope before drawing a legal conclusion about a broker.

Finding #3 · DATA PRACTICE TRENDSSample N = 155

B2B Workflow Observations Are Kept Separate From Trend Claims

Observation: 0 of 155 profiles with sales-oriented processing language also have a YES web-form observation.

Recorded Evidence: The cohort does not contain a source-backed average turnaround SLA, so no timing comparison is published.

Consumer Implication: Workflow channel observations do not establish speed, success, or a market-wide trend.

Finding #4 · FRICTION PATTERNSSample N = 2

Some People-Search Workflows Document Profile or Record Lookup

Observation: 0 of 2 profiles in this category have source-backed required-field text that mentions a profile, record, listing, or URL.

Recorded Evidence: The result is limited to explicit workflow evidence and does not claim that every people-search site requires URL extraction.

Consumer Implication: When a current source asks for a record identifier, preserve that requirement in the user workflow.

Finding #5 · CORPORATE CONCENTRATIONSample N = 1043

Recorded Parent-Entity Relationships Form a Bounded Ownership Cohort

Observation: 181 configured parent mappings cover 239 catalog brand/domain entries.

Recorded Evidence: This is a bounded internal mapping, not a complete ownership census; relationships require source-level review before publication as corporate fact.

Consumer Implication: A parent relationship does not by itself prove that a rights request cascades across sibling brands.

Finding #6 · REGULATORY ENFORCEMENTSample N = 1043

Captured State Registry Snapshots Do Not Fully Overlap

Observation: The captured California and Vermont registry matches overlap for 78 catalog entities; absence from one snapshot is not proof of non-registration everywhere.

Recorded Evidence: The crosswalk contains 623 California matches and 201 Vermont matches in the captured snapshots.

Consumer Implication: Always cite the jurisdiction, snapshot date, and matching method when comparing registry coverage.

07

Independent Audit & Verification

Research calculations, validation checks, and dataset generators are checked into the OfflistMe repository and can be run through the repository verification workflow:

# 1. Run the repository test suite

npm test

# 2. Rebuild and synchronize all 5 research datasets

npm run generate-entity-research-ledger

# 3. Validate ledger and public research dataset integrity

npm run check-entity-research-ledger