9. Setting and Defending a Risk Threshold
After this video you can
- Grade threshold precedents by provenance
- Convert between cell size and probability
- Apply agency cell-size suppression rules
- Apply and score the Five Safes
- Write a defensible threshold paragraph
Module 4: Risk Assessment · Runtime 46:10 · YouTube title: Setting and Defending a HIPAA Risk Threshold That Holds Up
The regulator gives no number, so your defense of it is everything. This video grades every commonly quoted threshold by its provenance, converts cell size to probability (with the rounding trap), works CMS complementary suppression, scores the Five Safes against real evidence, and ends by drafting the threshold paragraph of a report.
In this video
- Threshold precedents: CMS cell size 11, 0.09 from clinical trial transparency, 0.05 from the public microdata literature, and the 33 percent trap
- Grading provenance: traced strong (11, 0.09, 0.05), traced moderate (legacy k of 3 and 5), untraced for identity (0.33), practice-derived (0.20)
- From cell size to probability: k = 11 gives 0.0909, a strict 0.09 ceiling needs k = 12
- The CMS Cell Size Suppression Policy and complementary suppression, worked on Springfield × 10
- CalHHS and NCHS/CDC cell-size regimes; ISO/IEC 27559 on why no universal number exists
- Public versus controlled release, the Five Safes, mitigating controls, and inside a secure enclave
- Two releases of Springfield: public is arithmetically impossible, enclave lands one class exactly on 0.20
- The defensibility brief and a model threshold paragraph
Authorities quoted on screen
CMS Cell Size Suppression Policy (via ResDAC); CalHHS Data De-Identification Guidelines v2.2; NCHS / CDC data presentation standards; ISO/IEC 27559; Federal Committee on Statistical Methodology Working Paper 22.
Reference: threshold precedents and their provenance
The course teaches every commonly quoted number with its source and a provenance grade. (Video 9)
| Threshold | Equivalent | Source | Provenance |
|---|---|---|---|
| Cell size k ≥ 11 | Maximum risk 0.0909 | CMS Cell Size Suppression Policy (ResDAC); CalHHS Data De-Identification Guidelines (numerator ≥ 11, denominator ≥ 20,000) | Traced, strong |
| Maximum risk 0.09 | k ≥ 12 for a strict ceiling | EMA Policy 0070 and Health Canada PRCI (clinical trial transparency) | Traced, strong |
| Maximum risk 0.05 | k ≥ 20 | El Emam et al., Canadian public-use microdata literature | Traced, strong |
| k ≥ 3, k ≥ 5 | 0.333, 0.20 | Federal Committee on Statistical Methodology Working Paper 22 (legacy) | Traced, moderate |
| Maximum risk 0.20 | k ≥ 5 | Published practice for tightly controlled enclave access | Practice-derived, not codified |
| 0.33 | WP22's p-percent rule for business magnitude data | Untraced as an identity-risk threshold: a trap |
Estimator rule for population uniqueness (Dankar and El Emam, 2012): sampling fraction below 10 percent, use Pitman; at or above 10 percent, use Zayatz or the negative binomial.
Key takeaways
- Thresholds are defended by graded precedent plus context, never asserted bare.
- Cell size and probability are the same statement, and the conversion has a rounding trap.
- Controls and transformation trade off, and the documented balance expires when the context changes.
Coming next: Video 10, The Re-identification Case Canon
The most quoted cases in data privacy are also the most misquoted. Each landmark attack is retold with its dataset, its quasi-identifiers, its auxiliary data, and its lesson, followed by the four ways retellings go wrong and what the counter-literature actually measured.
Saved in your browser only — no account, no server.