Skip to content

Repository files navigation

Auditing the CFPB Consumer Complaint Database as a compliance-measurement source

An audit of the public CFPB Consumer Complaint Database (17,281,320 complaints, 1,199,531 of them debt collection) as a data source for measuring collections compliance.

Snapshot 2026-07-27. Every number below is produced by src/findings.py and stored in docs/findings.json; this file is generated from it by src/render_readme.py.

Two categories of collection conduct CFPB's sub-issue taxonomy cannot express

Two kinds of conduct with a statute behind them are unreachable from any of the 56 (Issue, Sub-issue) label pairs CFPB uses for debt collection. No dropdown selection a consumer or intake agent can make encodes them.

Continued collection during a pending dispute — §1692g(b)

No label expresses the SEQUENCE 'I disputed, and collection continued before verification arrived'. The dropdown can record that a dispute was made, but not that collection carried on in spite of it.

Detecting this conduct requires only two stored facts: a dispute flag with a timestamp, and any subsequent collection contact. Comparing them is a date comparison, not a reading of the message and not a judgment call — no narrative, no model, no reviewer. It is exactly the kind of violation a deterministic rule layer exists to catch. CFPB's structured fields cannot express it, so it is absent from any compliance measurement built on them.

Harmful or inaccurate credit reporting — FCRA

Exactly one of the 52 Sub-issue values under Debt collection mentions credit at all: 'Threatened or suggested your credit would be damaged' (121,821 rows, 2017-2026), and it sits under the Issue 'Took or threatened to take negative or legal action'. It records a THREAT to report — an action the collector says it may take. No value records the outcome: that the debt WAS reported inaccurately, was not corrected after a dispute, or was not marked as disputed.

That count is checked against the source, not asserted: scratch/credit_subissues.py enumerates every Sub-issue value occurring under Debt collection across all 1,199,531 rows and every vocabulary generation.

Anyone measuring collections compliance from this dataset's structured fields is blind to both by construction. The gap is not a sampling artifact and more data does not close it: the vocabulary has no value that encodes them. Detecting this conduct requires the free-text narrative, present on only 438,328 of 1,199,531 debt-collection complaints (36.5%).

The vocabulary is closed, so the misfit rate is unmeasurable

CFPB's debt-collection list has no residual 'other' option at all. Conduct fitting none of the listed pairs must be filed under a pair that does not fit, and nothing in the record marks that it was a poor fit. The misfit rate is therefore not merely unmeasured but unmeasurable from the structured fields: there is no value to count and no flag to filter on.

This is a different problem from the two gaps above. Those are missing values for known conduct. This one means that when the list fails, the failure leaves no trace: there is no other bucket to count, so the share of complaints filed under an ill-fitting label cannot be recovered from the structured fields at any sample size.

Labels are versioned and collide

43 sub-issue strings across 12 conduct categories are re-wordings of conduct another string already names — CFPB revised the vocabulary and the old and new wordings both persist in the data.

4 sub-issue strings appear under more than one parent Issue. 1 of them denotes different conduct depending on its parent:

You told them to stop contacting you, but they keep trying — 10,022 rows

  • under Communication tactics → contact_after_cease (8,689 rows)
  • under Electronic communications → contact_after_channel_optout (1,333 rows)

A groupby on Sub-issue alone silently merges distinct conduct, and splits single conduct across the vocabulary revision. Aggregations must key on the (Issue, Sub-issue) pair.

The preventable/judgment split is a property of the taxonomy, not of collector behaviour

Mapping all 438,328 narrative-bearing debt-collection complaints through the crosswalk splits them into conduct a deterministic rule layer could have prevented versus conduct requiring human judgment. Note the denominator: this is every year in the snapshot, not the 2015–2024 window used for the two findings below, because CFPB labels are present regardless of whether a narrative was published.

Rows Share
Deterministically preventable 146,121 33.34%
Requires judgment 283,527 64.68%
No category applies 8,680 1.98%

Reclassifying 2 of the 56 label pairs (7,543 rows) from false_threat to time_barred_collection moved the preventable share from 31.62% to 33.34% — a 1.72-point swing from one routing decision.

Treat this number as a description of what CFPB's dropdown can express under a stated mapping, not as a measurement of what collectors did.

Duplicate and template narratives concentrate in recent years

Within the 2015–2024 window, 35,149 of 304,434 narratives (11.55%) are byte-identical duplicates of another narrative. After removing them, a further 3,834 rows (1.42%) share a normalised fingerprint with at least one other row, in 1,667 clusters.

65.10% of those clustered rows are dated 2022–2024. The largest clusters are credit-repair form letters quoting FCRA language, filed under the Debt collection product:

Cluster size Years Distinct companies Opening text
29 2024 3 This serves as an effort to rectify inaccurate information present in my credit report as…
21 2024 3 This serves as an effort to rectify inaccurate information present in my credit report as…
17 2024 3 This serves as an effort to rectify inaccurate information present in my credit report as…
16 2022–2024 9 In accordance with the Fair Credit Reporting act. The List of accounts below has violated…
13 2022–2024 7 In accordance with the Fair Credit Reporting act XXXX Account # XXXX, has violated my righ…

These are not independent observations, and they land in the same years as the volume increase. A trend line over raw counts reads campaign activity as a change in complaint behaviour. Deduplicate on fingerprint before any year-over-year claim; confidence intervals computed on the raw pool are too narrow.

A date-parsing defect silently drops rows

Date received mixes two ISO 8601 shapes: 1,078,703 bare dates and 120,828 full timestamps, interleaved within the same read chunk. Without an explicit format, pandas infers one shape per chunk and coerces the rest to NaT — no error, no warning.

import pandas as pd
s = pd.Series(["2015-03-20", "2016-02-21T21:19:06.000Z"])
pd.to_datetime(s, errors="coerce", utc=True)
#   -> [Timestamp("2015-03-20"), NaT]        <-- silent loss
pd.to_datetime(s, format="ISO8601", errors="coerce", utc=True)
#   -> [Timestamp("2015-03-20"), Timestamp("2016-02-21 21:19:06")]

At the chunk size used here (200,000), that silently drops 4,852 debt-collection rows (1,194,679 seen instead of 1,199,531) and 2,491 rows from the 2015–2024 narrative pool (301,943 instead of 304,434). The count is chunk-size dependent, which is why it is easy to miss. Any analysis of this file that parses Date received without format="ISO8601" is dropping rows.

What is in this repository

  • taxonomy.yaml — 17 conduct categories (10 deterministically preventable, 7 judgment-requiring). 16 carry an FDCPA, Regulation F or FCRA citation; the remaining 1 is the residual other_conduct category, which has no statute by construction.
  • subissue_crosswalk.yaml — all 56 CFPB (Issue, Sub-issue) pairs mapped to a category or explicitly to null. 429,648 rows (98.02%) map; 8,680 (1.98%) describe conduct the taxonomy does not model. Validated by src/crosswalk.py, which fails if any pair is unmapped or any category id is unknown.
  • src/sample.py — streaming sampler. Filters and deduplicates the dump, draws a 4,000-row headline sample, an 8,000-row sample stratified at 800/year, and a 400-row subset, all seeded (20260727) and recorded in data/sample_manifest.json.
  • scratch/scan_products.py — profiler producing docs/data_profile.md and docs/issue_vocabulary.md.
  • src/findings.py — computes every figure above into docs/findings.json.

Method

  • Source — CFPB Consumer Complaint Database full export, snapshot 2026-07-27. SHA-256 begins 1865f3fdcc280fdd (first 16 of 64; the full digest is in docs/findings.json).
  • Scale — 17,281,320 rows; 1,199,531 debt collection; 438,328 of those carry a narrative.
  • Window — 2015-01-01 to 2024-12-31, used for the duplicate and date findings. 2013-2014 contain zero published narratives; 2025–2026 narrative rate is depressed (see docs/limitations.md).
  • Deduplication — exact match on the narrative normalised by strip, whitespace-collapse and casefold, keeping the earliest Complaint ID. Near-duplicates are additionally fingerprinted by removing XXXX redaction masks and punctuation.
  • Streaming — the 1.4 GB zip is read in chunks and never extracted, never loaded whole.
  • What is committed under data/ — only sample_manifest.json, which records the source hash, the seed, pool sizes, per-year counts and the Complaint IDs of each drawn sample. It contains IDs and counts, no narrative text. The export itself and every drawn CSV are gitignored.

Limitations

Full detail in docs/limitations.md. In short: the 2025–2026 narrative rate falls sharply and this study does not assert a mechanism; the snapshot is of a live, mutable database and the exact figures are not reproducible from a fresh download, though the method and the structural findings are; and complaint narratives under-describe omissions relative to threats, so any preventable share derived from narrative text would be a floor rather than a point estimate.

What this does not do

No narrative was classified. A second, narrative-based stream is designed and scaffolded — an axis-blinded labeling CLI (src/label_cli.py), drawn samples, and a two-stream design in docs/ that keeps the crosswalk independent of any narrative reading — but it was not run.

Consequently:

  • No label set in this repository is hand-labeled or human-validated, and nothing here should be described that way.
  • The codebook wording in taxonomy.yaml was iterated against 25 sample complaints. It was not validated against a labeled set.
  • Validating the classification stream would require a blind labeling pass with measured intra-rater agreement. The scaffolding for that exists; the result does not.
  • Finding 3 reports a composition of CFPB's own dropdown labels under a stated mapping. It is not a classification of what consumers wrote, and it is not a measurement of collector behaviour.

Reproduction

Windows PowerShell. The CFPB export is not committed; download it first.

python -m venv .venv
.venv\Scripts\Activate.ps1
pip install -r requirements-dev.txt

# 1. place the CFPB full export at data\complaints.csv.zip
#    https://www.consumerfinance.gov/data-research/consumer-complaints/
curl.exe -L -o data\complaints.csv.zip `
  "https://files.consumerfinance.gov/ccdb/complaints.csv.zip"

# 2. validate the taxonomy and the crosswalk
python src\taxonomy.py
python src\crosswalk.py

# 3. profile the dump (writes docs\data_profile.md, docs\issue_vocabulary.md)
python scratch\scan_products.py
python scratch\scan_pairs.py

# 4. draw the samples (writes data\sample_manifest.json)
python src\sample.py

# 5. compute every figure (writes docs\findings.json)
python src\findings.py

# 6. regenerate this README from findings.json
python src\render_readme.py

Steps 2 and 6 need no data file. Steps 3–5 stream the zip and take a few minutes each.

Run the app locally:

pip install -r requirements.txt
streamlit run app.py

The app reads only docs/findings.json and needs no data file.

Links


Generated from docs/findings.json (computed 2026-08-04T14:41:14Z) by src/render_readme.py. Do not edit by hand.

About

Audit of the CFPB Consumer Complaint Database as a source for measuring collections compliance. Two categories of FDCPA conduct its taxonomy cannot express, a silent date-parsing defect, and template-driven volume inflation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages