← All cheat sheets

ANONYMIZATION

Authorized use only. Offensive reference for systems you own or are explicitly permitted to test. You are responsible for staying within the law.

Comprehensive guide to privacy-enhancing techniques for protecting
personal data while maintaining data utility for analysis.

TERMINOLOGY#

Anonymization:    Irreversible process - data cannot be linked back to individual
Pseudonymization: Reversible with additional info - data can be re-identified
De-identification: General term covering both (US/HIPAA context)
PII:              Personally Identifiable Information
Quasi-identifier: Attributes that combined can identify individuals (age, zip, gender)

K-ANONYMITY#

Concept:
  Each record is indistinguishable from at least (k-1) other records
  with respect to quasi-identifiers.

Example (k=3):
  Before:                          After (k=3):
  Age | Zip   | Disease            Age  | Zip   | Disease
  29  | 13053 | Cancer             2*   | 130** | Cancer
  25  | 13068 | Flu                2*   | 130** | Flu
  27  | 13068 | Cancer             2*   | 130** | Cancer

Techniques to achieve:
  - Generalization: Replace specific values with ranges (age 29 -> 20-30)
  - Suppression: Remove certain values entirely (replace with *)

Limitations:
  - Homogeneity attack: all k records have same sensitive value
  - Background knowledge attack: attacker uses external info
  - Does not protect against attribute disclosure

Tools:
  - ARX Data Anonymization Tool (Java, open-source)
  - sdcMicro (R package)
  - Amnesia (web-based)

L-DIVERSITY#

Concept:
  Extension of k-anonymity. Each equivalence class must have at least
  l "well-represented" values for each sensitive attribute.

Types:
  - Distinct l-diversity: at least l distinct sensitive values per group
  - Entropy l-diversity: entropy of sensitive values >= log(l)
  - Recursive (c,l)-diversity: most frequent value appears < c times
    the least frequent value

Example (l=2):
  Group must have at least 2 different disease values
  Bad:  {Cancer, Cancer, Cancer} -- homogeneous, fails l-diversity
  Good: {Cancer, Flu, Cancer}   -- 2 distinct values, passes l=2

Limitations:
  - Skewness attack: if overall distribution is skewed
  - Similarity attack: if sensitive values are semantically similar
  - Computationally more expensive than k-anonymity

T-CLOSENESS#

Concept:
  Extension of l-diversity. Distribution of sensitive attribute in each
  equivalence class must be "close" to overall distribution.
  Distance measured using Earth Mover's Distance (EMD).

Formula:
  For each equivalence class Q:
  EMD(distribution of sensitive attr in Q, overall distribution) <= t

Parameter selection:
  - t = 0: perfect privacy (identical distributions)
  - t = 1: no privacy guarantee
  - Typical values: 0.1 to 0.3

When to use:
  - When l-diversity is insufficient due to skewed distributions
  - When attribute disclosure is a major concern
  - For numerical sensitive attributes with meaningful ordering

DIFFERENTIAL PRIVACY#

Concept:
  Mathematical framework guaranteeing that the output of a query is
  approximately the same whether or not any single individual's data
  is included in the dataset.

Definition:
  A mechanism M satisfies epsilon-differential privacy if for all
  datasets D1, D2 differing in one record, and all outputs S:
  Pr[M(D1) in S] <= e^epsilon * Pr[M(D2) in S]

Key parameter: epsilon (privacy budget)
  - epsilon = 0: perfect privacy (useless output)
  - epsilon = 0.1: strong privacy
  - epsilon = 1.0: moderate privacy
  - epsilon > 10: weak privacy

Mechanisms:
  Laplace Mechanism:
    - Add noise from Laplace distribution
    - noise = Laplace(0, sensitivity/epsilon)
    - Good for numerical queries (counts, sums, averages)

  Gaussian Mechanism:
    - Add noise from Gaussian distribution
    - Requires (epsilon, delta)-differential privacy
    - Better for high-dimensional queries

  Exponential Mechanism:
    - For non-numerical outputs
    - Selects output with probability proportional to utility score

  Randomized Response:
    - Each individual randomly lies about their answer
    - Classic technique from survey methodology

Tools and Libraries:
  - Google's differential-privacy library (C++, Java, Go)
  - OpenDP (Rust/Python)
  - IBM diffprivlib (Python)
  - TensorFlow Privacy (ML-specific)
  - Apple's implementation (iOS/macOS analytics)
  - Microsoft SmartNoise (SQL queries)

Example (Python with diffprivlib):
  from diffprivlib.tools import mean
  dp_mean = mean(data, epsilon=1.0, bounds=(0, 100))

DATA MASKING#

Types:
  Static Data Masking (SDM):
    - Applied to data at rest
    - Creates masked copy of database
    - Use for: test environments, development, analytics

  Dynamic Data Masking (DDM):
    - Applied at query time
    - Original data remains unchanged
    - Use for: role-based access, application layer

Techniques:
  Substitution:    Replace with realistic fake values
  Shuffling:       Randomly reorder values within column
  Nulling out:     Replace with NULL/empty
  Number variance: Add random variance (+/- percentage)
  Encryption:      Format-preserving encryption (FPE)
  Character mask:  Replace characters (John -> J***)

Example (email masking):
  john.doe@company.com -> j***.d**@c******.com

Tools:
  - Oracle Data Masking and Subsetting
  - IBM InfoSphere Optim
  - Delphix
  - PostgreSQL: pg_anonymize extension
  - MySQL: masking functions (Enterprise)
  - Faker (Python library for generating fake data)

TOKENIZATION#

Concept:
  Replace sensitive data with non-sensitive placeholder (token).
  Original data stored in secure token vault.

Types:
  Vault-based:       Token maps to original in secure database
  Vaultless:         Cryptographic algorithm generates token (no vault needed)
  Format-preserving: Token matches original format (e.g., 16-digit card number)

Use cases:
  - Payment card data (PCI DSS compliance)
  - Social Security / National ID numbers
  - Healthcare identifiers
  - API tokens for authentication

vs Encryption:
  Tokenization                    | Encryption
  Random, no mathematical link    | Mathematical relationship exists
  Format can be preserved easily  | Output often different format/length
  Cannot be reversed without vault| Can be reversed with key
  Ideal for structured data       | Works for all data types

Tools:
  - Vault by HashiCorp (with transit secrets engine)
  - AWS Payment Cryptography
  - Protegrity
  - TokenEx
  - Thales CipherTrust

HASHING WITH SALT#

Concept:
  One-way transformation using cryptographic hash function.
  Salt = random value added before hashing to prevent rainbow table attacks.

Process:
  1. Generate random salt (at least 16 bytes)
  2. Concatenate: salt + plaintext
  3. Apply hash function: hash(salt + plaintext)
  4. Store: salt + hash_output

Recommended algorithms:
  DO USE:
    - bcrypt (cost factor >= 12)
    - Argon2id (memory-hard, recommended)
    - scrypt (memory-hard)
    - PBKDF2 (NIST approved, >= 600,000 iterations for SHA-256)

  DO NOT USE for passwords:
    - MD5 (broken, fast)
    - SHA-1 (deprecated, fast)
    - SHA-256 alone (too fast without key stretching)

Example (Python):
  import bcrypt
  salt = bcrypt.gensalt(rounds=12)
  hashed = bcrypt.hashpw(password.encode(), salt)

  import hashlib, os
  salt = os.urandom(16)
  hashed = hashlib.pbkdf2_hmac('sha256', data.encode(), salt, 600000)

Limitations:
  - One-way: cannot retrieve original data
  - Same input always produces same output (deterministic)
  - Not suitable when original data must be recovered

GENERALIZATION#

Concept:
  Replace specific values with less specific but semantically
  consistent values.

Examples:
  Exact age 29        -> Age range 25-30
  Zip code 90210      -> Zip prefix 902**
  Date 1990-03-15     -> Year 1990 or decade 1990s
  Job: Data Engineer  -> Job: IT Professional
  City: San Francisco -> Region: West Coast

Hierarchy levels:
  Street address -> City -> State -> Country -> Continent

When to use:
  - Achieving k-anonymity
  - Reducing re-identification risk while keeping utility
  - Publishing aggregate statistics

Trade-off: more generalization = more privacy but less data utility

SUPPRESSION#

Concept:
  Remove or hide data values entirely.

Types:
  Record suppression:    Remove entire rows (outliers)
  Value suppression:     Replace specific cell values with *
  Attribute suppression: Remove entire columns

When to use:
  - Outlier records that are easily identifiable
  - Rare combinations of quasi-identifiers
  - Attributes that are direct identifiers (name, SSN, email)
  - When generalization alone cannot achieve desired privacy level

Best practice:
  - Suppress no more than a small percentage of records (typically < 5%)
  - Document what was suppressed and why

SYNTHETIC DATA GENERATION#

Concept:
  Create entirely artificial data that preserves statistical properties
  of the original dataset without containing real records.

Approaches:
  Statistical models:
    - Fit distributions to original data
    - Sample from fitted distributions
    - Preserves marginal distributions and correlations

  Deep Learning:
    - GANs (Generative Adversarial Networks)
    - VAEs (Variational Autoencoders)
    - Captures complex non-linear relationships

  Agent-based models:
    - Simulate individual behaviors
    - Useful for transaction or event data

Tools:
  - Gretel.ai (cloud-based, GANs and transformers)
  - Mostly AI (enterprise synthetic data)
  - Synthetic Data Vault (SDV) - Python library
  - Faker (simple rule-based fake data)
  - Synthea (healthcare-specific synthetic data)
  - DataSynthesizer (differential privacy aware)

Example (SDV - Python):
  from sdv.tabular import GaussianCopula
  model = GaussianCopula()
  model.fit(original_data)
  synthetic_data = model.sample(num_rows=1000)

Validation:
  - Compare distributions (KS test, chi-squared)
  - Check correlation preservation
  - Measure re-identification risk
  - Evaluate utility for downstream tasks

DECISION GUIDE: WHEN TO USE EACH#

Scenario                              | Recommended Technique
--------------------------------------|--------------------------------
Publishing research datasets          | k-anonymity + l-diversity
Aggregate statistics / queries        | Differential privacy
Payment card storage                  | Tokenization
Password storage                      | Hashing (bcrypt/Argon2id)
Test/dev environment data             | Data masking or synthetic data
Machine learning training             | Synthetic data or differential privacy
Sharing data with third parties       | Anonymization (generalize + suppress)
Real-time data access control         | Dynamic data masking
Compliance with GDPR Article 25       | Pseudonymization (tokenization/encryption)
High-dimensional data                 | Differential privacy
Healthcare data sharing               | k-anonymity + t-closeness

RISK ASSESSMENT CHECKLIST#

[ ] Identify all quasi-identifiers in the dataset
[ ] Assess linkage attack risk (cross-reference with public data)
[ ] Evaluate uniqueness of records
[ ] Test re-identification with sample attacks
[ ] Measure information loss / data utility after anonymization
[ ] Document technique chosen and parameters used
[ ] Review periodically as new data sources become available
[ ] Consider motivated intruder test (UK ICO approach)

REFERENCES#

- EU GDPR Recital 26 (definition of anonymous data)
- NIST SP 800-188 (de-identification of government datasets)
- Article 29 Working Party Opinion 05/2014 on Anonymisation Techniques
- ISO/IEC 20889:2018 (privacy-enhancing data de-identification)
- ENISA Pseudonymisation Techniques and Best Practices (2019)