Privacy-Enhancing Technologies (PET)
1. Overview
A. Definition
A collective term for technologies that remove or minimize the risk of identifying individuals while preserving the usefulness (utility) of data. They achieve data-related legal compliance (data laws·GDPR) and data utilization at the same time.
PET is not a single technology but an umbrella concept bundling a family of technologies that share a purpose: reconciling privacy protection with utilization. Therefore, "adopting PET" is less about buying a particular product and closer to the design act of choosing and combining techniques suited to the object of protection, the threat model, and the purpose of utilization. This is the first perspective to grasp when understanding PET.
B. Background and Necessity
Data utilization and privacy protection were traditionally seen as a zero-sum relationship. To utilize data you must share it, but sharing exposes individuals. Enterprises believed "to analyze, you must gather the originals," while regulation blocked movement, saying "if gathered, it leaks." As a result, valuable data was repeatedly locked in each institution's silo, unable to be used.
In particular, multiple cases revealed that mere anonymization cannot prevent re-identification attacks that single out individuals again by combining with other data. A representative example is the incident in which anonymous movie-rating data released by Netflix was cross-combined with publicly available IMDb ratings, re-identifying a substantial number of users; another classic case is where anonymized public medical data, combined with a voter roll, identified the medical records of a particular governor. Research backing this shows that combining just a few quasi-identifiers such as date of birth, sex, and postal code uniquely singles out a substantial proportion of the population.
PET is a family of technologies that emerged to break this zero-sum, reconciling utilization and protection (positive-sum) through approaches like "obtaining only the computed result without showing the data" or "mathematically hiding an individual's contribution." That is, a shift in thinking from "gather then protect" to "utilize without gathering" or "utilize while encrypted." As demand for large-scale data sharing for AI training explodes and regulations such as GDPR and domestic data-protection law strengthen simultaneously, PET is emerging as the only realistic means to "use data while complying with regulation."
C. Characteristics
Three characteristics run through all PET. First, the simultaneous pursuit of utility and protection: unlike traditional approaches that sacrifice one of the two, PET designs the point of reconciliation mathematically or structurally. Second, quantifiability: like differential privacy's ε, it presents "how much is protected" as a measurable indicator to manage risk. Third, being composable: because no single technique is a cure-all, PET presupposes layering several techniques according to the threat model. These characteristics recur throughout the classification and combination discussions that follow.
2. Classification Scheme
PET divides into three branches by its approach principle. The classification diagram below shows three strategies: "transform the data itself, compute while encrypted, or avoid gathering altogether."
flowchart TB
P["PET"] --> A["De-identification tech"]
P --> B["Cryptography-based"]
P --> C["Distributed·collaborative learning"]
A --> A1["Pseudonymization/anonymization"]
A --> A2["Differential privacy"]
B --> B1["Homomorphic encryption"]
B --> B2["Zero-knowledge proof ZKP"]
B --> B3["Multi-party computation MPC"]
C --> C1["Federated learning FL"]
De-identification technology is the family that transforms the data itself to lower identifiability; it is simple because the processed data can be used as-is, but the re-identification risk does not entirely disappear. It is the most traditional area, most tightly tied to law and institutions, and forms the foundation of the pseudonymized-information regime.
The cryptography-based family computes while the data is encrypted (homomorphic encryption), proves only a fact without revealing the original (ZKP), or lets multiple participants jointly compute while each hides their data (MPC). It guarantees mathematically strong security but shares the common weakness of large computation·communication costs, so performance optimization becomes the key to practical adoption.
The distributed·collaborative learning family takes a different approach with federated learning, which does not gather data in one place but learns in each location and shares only the model, thereby eliminating data movement itself. It is an approach that removes risk not by "protecting" but by "not moving in the first place."
The three families differ in the "location" of protection. De-identification focuses protection on the data itself, cryptography-based on the computation process, and distributed learning on the data's location. Understanding this difference lets one design combinations suited to the situation.
3. Major Technologies
Each technology has a different trade-off of "what it hides and what it gains." The architecture diagram below takes federated learning (FL) as an example, showing the flow in which the original data stays in each place and only the learned model parameters are gathered centrally.
flowchart LR
subgraph Site1["Hospital A"]
D1["Local data"] --> M1["Local training"]
end
subgraph Site2["Hospital B"]
D2["Local data"] --> M2["Local training"]
end
M1 -->|"Parameters only"| G["Central aggregation(global model)"]
M2 -->|"Parameters only"| G
G -->|"Distribute updated model"| M1
G -->|"Distribute updated model"| M2
Pseudonymization/anonymization is the most basic method, replacing identifiers such as names and resident numbers with pseudonyms or masks; its legal basis is clear and processing is simple. However, as in the re-identification cases above, even erasing direct identifiers can let individuals be singled out again through combinations of quasi-identifiers (age, region, occupation, etc.), so additional models such as k-anonymity·l-diversity control the combination risk. Simple though it is, the essential limitation is the difficulty of judging "how much must be erased to be safe."
Differential privacy (DP) adds precisely calculated noise to a statistical result to mathematically limit "the effect that whether a particular individual is included in the data has on the result." It quantifies the upper bound of this effect as ε (epsilon); the smaller ε, the stronger the protection but the lower the accuracy of the result. It is the principle of protecting individual respondents while preserving the utility of aggregate statistics, and its provability distinguishes it from other de-identification techniques.
Homomorphic encryption (HE) performs addition·multiplication operations in the ciphertext state without decryption, so that sensitive data entrusted to the cloud for computation while encrypted does not expose the original. In the sense of "computing without opening," it is the most ideal, but computation can be thousands to tens of thousands of times slower than plaintext, making performance the biggest obstacle. Zero-knowledge proof (ZKP) proves only a fact—"I know that value / the condition is satisfied"—without disclosing the secret value, used to prove one is an adult without revealing one's age, or to prove a transaction's validity on a blockchain while hiding the balance.
Multi-party computation (MPC) is a technique in which multiple institutions jointly compute a function without disclosing their data; no participant can learn another's input, yet all obtain the final result together. Federated learning (FL) gathers and integrates only the model parameters obtained from local training instead of the originals, so there is no data movement, but the possibility of inferring parts of the original from the shared parameters remains, so it is usually reinforced by combining with DP, etc.
| Technology | Principle | Strength | Limitation (trade-off) |
|---|---|---|---|
| Pseudonymization/anonymization | Replacing·removing identifiers | Simple·legal basis | Re-identification risk remains |
| Differential privacy (DP) | Adding noise to statistics | Mathematical guarantee (ε) | Accuracy loss |
| Homomorphic encryption (HE) | Computing in encrypted state | Strong confidentiality | High computation cost |
| Zero-knowledge proof (ZKP) | Proving facts without exposure | Authentication·blockchain use | Complex proof design |
| Multi-party computation (MPC) | Joint computation without sharing | Multi-institution collaborative analysis | Communication overhead |
| Federated learning (FL) | Sharing only parameters | No data movement | Model inversion risk |
4. Application Cases
PET especially shows its worth in domains where data movement is difficult for legal or competitive reasons. In finance, several banks run a joint fraud-detection model via MPC without handing customer data to one another, detecting money laundering and anomalous transactions spanning multiple banks that could not be caught alone. The key is that banks—competitors and regulated entities—can "keep data each in place and share only the result," a collaboration that was impossible with traditional data sharing.
In healthcare, strong regulatory and ethical constraints prevent taking patient data outside the hospital, so each hospital does only local training and gathers model parameters through federated learning to jointly develop disease-diagnosis·image-reading models. For example, multiple hospitals use their own rare-disease images only for model training and do not share the originals, thereby securing the training volume that a single hospital's data lacked while protecting patient privacy.
In statistics·the public sector, when releasing demographic statistics, differential privacy adds noise to secure both individual-respondent protection and statistical utility at once. A representative example is the US Census Bureau actually adopting differential privacy in the 2020 Census, demonstrating that DP can be applied practically at the enormous scale of national statistics. Furthermore, in the data-combination area, pseudonymized information is safely combined through a combination-specialist institution and analyzed in a secure processing environment (data safety zone) that controls external export, thereby running institution and technology in parallel.
Advertising·marketing is also a representative application area. In line with the trend toward abolishing third-party cookies, as demand grows to measure advertising effectiveness without tracking individual users, PET such as differential privacy and aggregation-based measurement is being introduced. One example is the method whereby advertisers and publishers confirm only aggregate results inside a data clean room without directly exchanging their respective customer data, showing a typical positive-sum of obtaining only campaign performance without individual-level identification.
| Field | Use |
|---|---|
| Finance | Inter-institution joint fraud-detection analysis (MPC) |
| Healthcare | Inter-hospital federated learning (FL) |
| Statistics·public | Differential-privacy-based statistics release |
| Data combination | Pseudonymized-information combination·data safety zone |
5. Deep Dive: Regulatory·Standards·Market Trends
PET is now settling in as a means actively recommended by regulators, going beyond an academic concept. The UK ICO, the UN, and others have published PET-utilization guides, and the US and UK operate a joint challenge competing on PET-based innovation, among other policy drives. It is a trend in which PET is singled out as the de facto standard means to technically implement the Data Protection by Design that GDPR requires.
Technical standardization is also advancing. For homomorphic encryption, parameter·security-level standardization is discussed centered on the international HomomorphicEncryption.org community, and standards related to de-identification and privacy technology are also being developed at ISO/IEC (e.g., ISO/IEC 20889, de-identification terminology·classification of techniques). Standardization lets "safety each institution used to claim on its own" be compared and verified on a common scale, forming the basis for PET to settle as an institutional technology that becomes an object of procurement and audit beyond the laboratory.
From a market perspective, PET is being reinterpreted from a "cost center (regulatory response)" into a "value-creation means (data collaboration·monetization)." In the past it was regarded as defensive spending to avoid regulation, but now recognition is shifting toward it as a precondition for collaboration that safely combines data with competitors and other industries to create new analytical value. This shift in recognition is the fundamental driver propelling the growth of the PET market.
From the perspective of expected exam directions, PET regularly appears as the go-to solution in questions about the point where "AI·data utilization" and "privacy protection" collide. Therefore, in an answer, rather than stopping at listing individual techniques, the high-scoring strategy is to connect "why they must be combined (the limits of a single technique)" and "how they interlock with institutions (pseudonymized information·safety zone)."
6. Considerations and Implications
- Limits of a single technology and combination: Since none is complete on its own, techniques are combined. For example, federated learning (FL) alone can allow the original to be inferred from parameters, so FL + DP adds noise, and MPC and HE are combined to secure both collaboration and confidentiality—composing a layered defense suited to the threat model.
- Practical tuning of performance-accuracy: Since the computation·communication costs of HE·MPC and the accuracy loss of DP are obstacles to practical adoption, the operational capability to adjust the protection level (ε, etc.) against performance·accuracy to match business requirements is key. Excessive protection renders data useless, while insufficient protection breeds regulatory violation.
- Running institution·governance in parallel: Technology alone is insufficient; one must apply privacy law·the pseudonymized-information regime together with Privacy by Design (PbD) principles, embedding privacy from the design stage to be sustainable. Technology is the means, governance the framework.
- Securing verification·transparency: Rather than the organization itself merely asserting whether PET application is "truly safe," it must be proven through re-identification risk assessment·privacy impact assessment (PIA) and external verification to gain trust. The basis for selecting protection parameters such as the ε value must also be documented.
- Preparing for future threats (crypto-agility): Cryptography-based PET such as homomorphic encryption and ZKP may be affected by future changes in the computing environment such as quantum computing, so they must be designed with a replaceable structure (crypto-agility) that is not locked to a particular algorithm, to secure long-term safety.
References
- ICO (UK Information Commissioner's Office), "Privacy-enhancing technologies (PETs)", https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/privacy-enhancing-technologies/
- US Census Bureau, "Differential Privacy and the 2020 Census", https://www.census.gov/programs-surveys/decennial-census/decade/2020/planning-management/process/disclosure-avoidance.html
- ISO/IEC 20889:2018, "Privacy enhancing data de-identification terminology and classification of techniques", https://www.iso.org/standard/69373.html
In one line: PET is a family of technologies that reconcile data utilization and privacy protection through pseudonymization/anonymization, differential privacy, homomorphic encryption, ZKP, MPC, and federated learning; the key is to understand the three families with different protection foci (data·computation·location), combine them to fit the threat model (FL+DP, etc.), and run them in parallel with the pseudonymized-information regime·PbD.