ANONYMIZATION
Authorized use only. Offensive reference for systems you own or are explicitly permitted to test. You are responsible for staying within the law.
Comprehensive guide to privacy-enhancing techniques for protecting personal data while maintaining data utility for analysis.
TERMINOLOGY#
Anonymization: Irreversible process - data cannot be linked back to individual Pseudonymization: Reversible with additional info - data can be re-identified De-identification: General term covering both (US/HIPAA context) PII: Personally Identifiable Information Quasi-identifier: Attributes that combined can identify individuals (age, zip, gender)
K-ANONYMITY#
Concept: Each record is indistinguishable from at least (k-1) other records with respect to quasi-identifiers. Example (k=3): Before: After (k=3): Age | Zip | Disease Age | Zip | Disease 29 | 13053 | Cancer 2* | 130** | Cancer 25 | 13068 | Flu 2* | 130** | Flu 27 | 13068 | Cancer 2* | 130** | Cancer Techniques to achieve: - Generalization: Replace specific values with ranges (age 29 -> 20-30) - Suppression: Remove certain values entirely (replace with *) Limitations: - Homogeneity attack: all k records have same sensitive value - Background knowledge attack: attacker uses external info - Does not protect against attribute disclosure Tools: - ARX Data Anonymization Tool (Java, open-source) - sdcMicro (R package) - Amnesia (web-based)
L-DIVERSITY#
Concept:
Extension of k-anonymity. Each equivalence class must have at least
l "well-represented" values for each sensitive attribute.
Types:
- Distinct l-diversity: at least l distinct sensitive values per group
- Entropy l-diversity: entropy of sensitive values >= log(l)
- Recursive (c,l)-diversity: most frequent value appears < c times
the least frequent value
Example (l=2):
Group must have at least 2 different disease values
Bad: {Cancer, Cancer, Cancer} -- homogeneous, fails l-diversity
Good: {Cancer, Flu, Cancer} -- 2 distinct values, passes l=2
Limitations:
- Skewness attack: if overall distribution is skewed
- Similarity attack: if sensitive values are semantically similar
- Computationally more expensive than k-anonymity
T-CLOSENESS#
Concept: Extension of l-diversity. Distribution of sensitive attribute in each equivalence class must be "close" to overall distribution. Distance measured using Earth Mover's Distance (EMD). Formula: For each equivalence class Q: EMD(distribution of sensitive attr in Q, overall distribution) <= t Parameter selection: - t = 0: perfect privacy (identical distributions) - t = 1: no privacy guarantee - Typical values: 0.1 to 0.3 When to use: - When l-diversity is insufficient due to skewed distributions - When attribute disclosure is a major concern - For numerical sensitive attributes with meaningful ordering
DIFFERENTIAL PRIVACY#
Concept:
Mathematical framework guaranteeing that the output of a query is
approximately the same whether or not any single individual's data
is included in the dataset.
Definition:
A mechanism M satisfies epsilon-differential privacy if for all
datasets D1, D2 differing in one record, and all outputs S:
Pr[M(D1) in S] <= e^epsilon * Pr[M(D2) in S]
Key parameter: epsilon (privacy budget)
- epsilon = 0: perfect privacy (useless output)
- epsilon = 0.1: strong privacy
- epsilon = 1.0: moderate privacy
- epsilon > 10: weak privacy
Mechanisms:
Laplace Mechanism:
- Add noise from Laplace distribution
- noise = Laplace(0, sensitivity/epsilon)
- Good for numerical queries (counts, sums, averages)
Gaussian Mechanism:
- Add noise from Gaussian distribution
- Requires (epsilon, delta)-differential privacy
- Better for high-dimensional queries
Exponential Mechanism:
- For non-numerical outputs
- Selects output with probability proportional to utility score
Randomized Response:
- Each individual randomly lies about their answer
- Classic technique from survey methodology
Tools and Libraries:
- Google's differential-privacy library (C++, Java, Go)
- OpenDP (Rust/Python)
- IBM diffprivlib (Python)
- TensorFlow Privacy (ML-specific)
- Apple's implementation (iOS/macOS analytics)
- Microsoft SmartNoise (SQL queries)
Example (Python with diffprivlib):
from diffprivlib.tools import mean
dp_mean = mean(data, epsilon=1.0, bounds=(0, 100))
DATA MASKING#
Types:
Static Data Masking (SDM):
- Applied to data at rest
- Creates masked copy of database
- Use for: test environments, development, analytics
Dynamic Data Masking (DDM):
- Applied at query time
- Original data remains unchanged
- Use for: role-based access, application layer
Techniques:
Substitution: Replace with realistic fake values
Shuffling: Randomly reorder values within column
Nulling out: Replace with NULL/empty
Number variance: Add random variance (+/- percentage)
Encryption: Format-preserving encryption (FPE)
Character mask: Replace characters (John -> J***)
Example (email masking):
john.doe@company.com -> j***.d**@c******.com
Tools:
- Oracle Data Masking and Subsetting
- IBM InfoSphere Optim
- Delphix
- PostgreSQL: pg_anonymize extension
- MySQL: masking functions (Enterprise)
- Faker (Python library for generating fake data)
TOKENIZATION#
Concept: Replace sensitive data with non-sensitive placeholder (token). Original data stored in secure token vault. Types: Vault-based: Token maps to original in secure database Vaultless: Cryptographic algorithm generates token (no vault needed) Format-preserving: Token matches original format (e.g., 16-digit card number) Use cases: - Payment card data (PCI DSS compliance) - Social Security / National ID numbers - Healthcare identifiers - API tokens for authentication vs Encryption: Tokenization | Encryption Random, no mathematical link | Mathematical relationship exists Format can be preserved easily | Output often different format/length Cannot be reversed without vault| Can be reversed with key Ideal for structured data | Works for all data types Tools: - Vault by HashiCorp (with transit secrets engine) - AWS Payment Cryptography - Protegrity - TokenEx - Thales CipherTrust
HASHING WITH SALT#
Concept:
One-way transformation using cryptographic hash function.
Salt = random value added before hashing to prevent rainbow table attacks.
Process:
1. Generate random salt (at least 16 bytes)
2. Concatenate: salt + plaintext
3. Apply hash function: hash(salt + plaintext)
4. Store: salt + hash_output
Recommended algorithms:
DO USE:
- bcrypt (cost factor >= 12)
- Argon2id (memory-hard, recommended)
- scrypt (memory-hard)
- PBKDF2 (NIST approved, >= 600,000 iterations for SHA-256)
DO NOT USE for passwords:
- MD5 (broken, fast)
- SHA-1 (deprecated, fast)
- SHA-256 alone (too fast without key stretching)
Example (Python):
import bcrypt
salt = bcrypt.gensalt(rounds=12)
hashed = bcrypt.hashpw(password.encode(), salt)
import hashlib, os
salt = os.urandom(16)
hashed = hashlib.pbkdf2_hmac('sha256', data.encode(), salt, 600000)
Limitations:
- One-way: cannot retrieve original data
- Same input always produces same output (deterministic)
- Not suitable when original data must be recovered
GENERALIZATION#
Concept: Replace specific values with less specific but semantically consistent values. Examples: Exact age 29 -> Age range 25-30 Zip code 90210 -> Zip prefix 902** Date 1990-03-15 -> Year 1990 or decade 1990s Job: Data Engineer -> Job: IT Professional City: San Francisco -> Region: West Coast Hierarchy levels: Street address -> City -> State -> Country -> Continent When to use: - Achieving k-anonymity - Reducing re-identification risk while keeping utility - Publishing aggregate statistics Trade-off: more generalization = more privacy but less data utility
SUPPRESSION#
Concept: Remove or hide data values entirely. Types: Record suppression: Remove entire rows (outliers) Value suppression: Replace specific cell values with * Attribute suppression: Remove entire columns When to use: - Outlier records that are easily identifiable - Rare combinations of quasi-identifiers - Attributes that are direct identifiers (name, SSN, email) - When generalization alone cannot achieve desired privacy level Best practice: - Suppress no more than a small percentage of records (typically < 5%) - Document what was suppressed and why
SYNTHETIC DATA GENERATION#
Concept:
Create entirely artificial data that preserves statistical properties
of the original dataset without containing real records.
Approaches:
Statistical models:
- Fit distributions to original data
- Sample from fitted distributions
- Preserves marginal distributions and correlations
Deep Learning:
- GANs (Generative Adversarial Networks)
- VAEs (Variational Autoencoders)
- Captures complex non-linear relationships
Agent-based models:
- Simulate individual behaviors
- Useful for transaction or event data
Tools:
- Gretel.ai (cloud-based, GANs and transformers)
- Mostly AI (enterprise synthetic data)
- Synthetic Data Vault (SDV) - Python library
- Faker (simple rule-based fake data)
- Synthea (healthcare-specific synthetic data)
- DataSynthesizer (differential privacy aware)
Example (SDV - Python):
from sdv.tabular import GaussianCopula
model = GaussianCopula()
model.fit(original_data)
synthetic_data = model.sample(num_rows=1000)
Validation:
- Compare distributions (KS test, chi-squared)
- Check correlation preservation
- Measure re-identification risk
- Evaluate utility for downstream tasks
DECISION GUIDE: WHEN TO USE EACH#
Scenario | Recommended Technique --------------------------------------|-------------------------------- Publishing research datasets | k-anonymity + l-diversity Aggregate statistics / queries | Differential privacy Payment card storage | Tokenization Password storage | Hashing (bcrypt/Argon2id) Test/dev environment data | Data masking or synthetic data Machine learning training | Synthetic data or differential privacy Sharing data with third parties | Anonymization (generalize + suppress) Real-time data access control | Dynamic data masking Compliance with GDPR Article 25 | Pseudonymization (tokenization/encryption) High-dimensional data | Differential privacy Healthcare data sharing | k-anonymity + t-closeness
RISK ASSESSMENT CHECKLIST#
[ ] Identify all quasi-identifiers in the dataset [ ] Assess linkage attack risk (cross-reference with public data) [ ] Evaluate uniqueness of records [ ] Test re-identification with sample attacks [ ] Measure information loss / data utility after anonymization [ ] Document technique chosen and parameters used [ ] Review periodically as new data sources become available [ ] Consider motivated intruder test (UK ICO approach)
REFERENCES#
- EU GDPR Recital 26 (definition of anonymous data) - NIST SP 800-188 (de-identification of government datasets) - Article 29 Working Party Opinion 05/2014 on Anonymisation Techniques - ISO/IEC 20889:2018 (privacy-enhancing data de-identification) - ENISA Pseudonymisation Techniques and Best Practices (2019)