Data governance and privacy
Analytics platforms collect a copy of everything, which often makes them the largest store of personal data in a company. Governance is the set of habits that keeps that data findable, correctly used, and deleted on time. What follows is engineering practice, not legal advice; your privacy team decides what applies to you.
Know what you hold
Personal data is any information relating to an identifiable person: names, emails, phone numbers, IP addresses, device ids, precise locations, and also combinations such as postcode plus date of birth that identify someone indirectly.
Classify every dataset, and ideally every column, with a small set of levels:
| Level | Examples | Typical handling |
|---|---|---|
| Public | Published prices | No restrictions |
| Internal | Aggregated sales | Staff only |
| Confidential | Customer emails, order history | Need-to-know access, masked by default |
| Restricted | Health data, payment details, government ids | Tightly limited, audited, often not copied at all |
Record the classification in your catalog, so access rules and masking can be driven from it.
Access control and least privilege
Apply least privilege: each person and service gets the minimum access their work needs.
- Grant access to roles, not individuals, and review membership regularly.
- Give each pipeline a service account limited to the tables it uses.
- Keep raw schemas with identifiers apart from the modelled ones analysts use.
- Log queries on sensitive tables, and review the logs.
Masking, pseudonymisation and anonymisation
These three are often confused:
- Masking hides part of a value from a viewer, such as showing
a***@example.com. The underlying data is unchanged. - Pseudonymisation replaces identifiers with tokens, so records can still be joined and counted per person without showing who they are. Someone with the key or lookup table can reverse it. Under GDPR, pseudonymised data is still personal data.
- Anonymisation removes the possibility of identifying anyone, including by combining fields. Truly anonymous data falls outside GDPR, but it is harder to achieve than it looks: rare combinations of attributes can single people out.
A plain hash of an email is not anonymous, since anyone can hash a list of known emails and match them. Use a keyed hash with a secret kept outside the warehouse if you need a stable pseudonym.
A masked view in PostgreSQL, with analysts denied the raw table:
CREATE VIEW analytics.customers_masked AS
SELECT id,
left(email, 1) || '***@' || split_part(email, '@', 2)
AS email_masked,
country,
date_trunc('month', signup_at) AS signup_month
FROM raw.customers;
REVOKE ALL ON raw.customers FROM analyst;
GRANT SELECT ON analytics.customers_masked TO analyst;
Many warehouses also offer built-in column masking policies.
Retention and deletion
Set a retention period per dataset and enforce it with scheduled jobs or table expiry settings, not intentions.
Deletion requests, such as the GDPR right to erasure, must reach every copy: raw and derived tables, exports, caches and search indexes. Backups are the hard part. Keep backup retention short and documented, and re-apply pending deletions after any restore. Some teams encrypt each person’s data with its own key and delete the key, making every copy unreadable at once.
Lineage and catalogs
You cannot delete or protect data you cannot find. A data catalog lists datasets with their owner, description and classification. Lineage records which datasets are built from which, ideally down to columns. It answers “which reports use this email column?” before a deletion or schema change.
Privacy principles in practice
GDPR sets out principles that translate into engineering decisions:
- Purpose limitation: data collected for one purpose should not be reused for an unrelated one. Record each dataset’s purpose.
- Data minimisation: collect and copy only the fields you need. The column you never load cannot leak.
- Storage limitation: keep identifiable data only as long as needed, which is your retention policy.
- Integrity and confidentiality: protect data with access control, encryption and monitoring.
Checklist
- Every dataset has an owner, a classification and a stated purpose.
- Access goes through roles, with least privilege and audit logs.
- Identifiers are masked or pseudonymised for most users by default.
- Retention is enforced automatically, and deletions reach derived tables and backups.
- Lineage shows where each sensitive column flows.