10. The Re-identification Case Canon
After this video you can
- Retell the landmark re-identification attacks
- Name each attack's quasi-identifiers
- Name each attack's auxiliary data
- Calibrate lessons without attack folklore
- Convert cases into assessment parameters
Module 5: Attack Literature · Runtime 38:47 · YouTube title: Famous Re-identification Attacks: Weld, Netflix, and the Truth
The most quoted cases in data privacy are also the most misquoted. Each landmark attack is retold with its dataset, its quasi-identifiers, its auxiliary data, and its lesson, followed by the four ways retellings go wrong and what the counter-literature actually measured.
The cases
| Case | Dataset | Quasi-identifiers | Auxiliary data | Lesson |
|---|---|---|---|---|
| Governor Weld, 1997 | Massachusetts GIC hospital discharge records | 5-digit ZIP, full date of birth, sex | Cambridge voter list, bought for $20 | Public context multiplies risk for specific individuals |
| 87 percent or 63 | Population uniqueness on ZIP5 + birth date + sex | Sweeney (1990 census) vs Golle (2000 census) | The headline number depends on the census year and rises sharply with age | |
| AOL, 2006 | 20 million search queries, 650,000 users | The free-text queries themselves | None needed | Pseudonymization is not de-identification |
| Netflix Prize, 2008 | 100 million ratings, ~500,000 subscribers | Movie titles, scores, rating dates | Public IMDb profiles | Sparse longitudinal patterns behave like fingerprints |
| Homer 2008 / GWAS | Pooled allele frequencies in dbGaP | Aggregate SNP frequencies | The target's own genome | Aggregation alone does not remove disclosure risk |
| Gymrek 2013 | 1000 Genomes and Personal Genome Project | Y-chromosome STRs, age, state | Ysearch and Sorenson genealogy databases | Consumer technology creates new quasi-identifiers |
| Unique in the Crowd, 2013 / 2015 | 15 months of mobility traces, 1.5 million users | Spatio-temporal points | None: a uniqueness study | Behavioral traces function like biometrics |
| Washington State, 2013 | Statewide discharge records, sold for $50 | Hospital, diagnosis, procedure, age, sex, ZIP | Newspaper archives searched for "hospitalized" | Public narrative is auxiliary data (35 matches from 81 stories, 43 percent) |
| Rocher 2019 | Generative copula model, 15 attributes | 15 attributes | None | Sampling alone is not a defense (99.98 percent unique) |
| The frontier, 2020 onward | Raw ECG waveforms and wearable telemetry | Signal morphology itself | None | Stripping metadata does not de-identify a signal |
Key takeaways
- Every landmark attack pairs residual quasi-identifiers with a reasonably available auxiliary dataset.
- Pseudonymization, aggregation, sparsity, and sampling have each failed as stand-alone defenses.
- The counter-literature shows properly de-identified data rarely falls, so calibrate rather than catastrophize.
Coming next: Video 11, Becoming the Expert: Pathways and Credentials
No certificate exists, so your career plan is a body of evidence. Where practicing experts actually come from, the five markers that make a CV defensible, which credentials teach law and which teach the math, the practitioner canon, the four market tiers, and a realistic 24-month pathway to a first engagement.
Saved in your browser only — no account, no server.