← Back to list
Security & Privacy
#데이터마스킹#토큰화#FPE#PCI-DSS#가명처리
Last updated · 2026-10-11

Data Masking & Tokenization

1. Overview

Data masking transforms and hides actual values into de-identified fake values while preserving the format and referential integrity of the original data, so the data can be used without exposing sensitive information. Tokenization replaces a sensitive value with a substitute value (a token) that has no mathematical relationship to the original, and keeps the mapping between original and token recoverable only through a separate secure store (a token vault) or an algorithm.

Sensitive data does not stay inside the operational database. The practice of copying production data into development, test, training, and analytics environments; the business need to share data with outsourced developers, data analysts, and marketing teams; and migration to the cloud all constantly create situations where "data leaves the controlled operational boundary." If values such as national ID numbers, card numbers, accounts, and contact details are copied verbatim, a single breach spills millions of real records to the outside. The fact that many large breaches, both at home and abroad, occurred not in the production tier but in loosely secured development, backup, and analytics environments shows the heart of this problem.

Data masking and tokenization emerged as protection techniques that resolve exactly this gap — the contradiction that "data must be used, but the original cannot be exposed." The traditional control of access management limits "who can see," but once data is copied or extracted it is powerless; encryption is strong at rest and in transit but must eventually be decrypted into plaintext for an application to compute or display, so it has limits for "in use" protection. Masking and tokenization change the value of the data itself into an irreversible (masking) or separately recoverable (tokenization) form, so that wherever the data flows, it carries no meaning on its own — complementing access control and encryption.

Regulatory pressure also drove the spread of these technologies. The card industry's PCI-DSS requires minimizing the scope of systems that store, process, and transmit card numbers (PANs); replacing a PAN with a token removes the systems handling tokens from the audit scope, greatly reducing compliance burden and cost. Korea's Personal Information Protection Act and its standards for safety-assurance measures require encryption and access control for unique identifying information, and permit and encourage processing that uses pseudonymization and masking. In this way, masking and tokenization have established themselves not as mere security techniques but as de-risking measures that simultaneously lower compliance costs and the damage from incidents.

2. The Spectrum and Classification of Data Protection Techniques

To understand masking and tokenization precisely, one must first grasp where they sit within the overall spectrum of data protection techniques. Protection techniques divide along the axes of "is restoration of the original required," "how much of the analytical usefulness of the value is preserved," and "is the protection embedded in the data or applied at the moment of access." The diagram below shows the relationships among the main techniques.

graph TD
    ROOT["Sensitive data protection techniques"]
    ROOT --> ENC["Encryption<br/>reversible, key-dependent"]
    ROOT --> TOK["Tokenization<br/>reversible, vault/algorithm-dependent"]
    ROOT --> MASK["Masking<br/>in principle irreversible"]
    ROOT --> DEID["De-identification/pseudonymization<br/>statistical anonymity"]
    MASK --> SDM["Static masking (SDM)<br/>at copy-creation time"]
    MASK --> DDM["Dynamic masking (DDM)<br/>real time at query time"]
    TOK --> VLT["Vaulted"]
    TOK --> VLS["Vaultless (FPE)"]

Encryption is a reversible technique in which a party holding the key can restore the original at any time, and its protection strength depends entirely on key management. If the key leaks, everything is exposed; the ciphertext changes form so its length and type differ from the original, making it prone to conflict with existing schemas and validation logic. Tokenization is also reversible, but the token itself contains no mathematical information from which the original could be derived, so restoration is impossible unless the token store is stolen — this is the decisive difference from encryption. That is, if the safety of encryption lies in "the secrecy of the key," the safety of tokenization lies in "the physical separation of the mapping."

Masking is, in principle, an irreversible transformation that does not presuppose restoration. It suits cases where it is sufficient for a developer to test with data statistically similar to production data, with no need to revert to the original. De-identification and pseudonymization are statistical protections that prevent identifying an individual, linked to k-anonymity, differential privacy, and the like, and masking is used as one of their implementation means. The four techniques are not mutually exclusive but combine in layers — for example, card numbers in the production DB are protected by tokenization, the development environment that copies that DB applies static masking, and the transmission segment is wrapped in encryption.

Category Encryption Tokenization Masking (irreversible)
Reversibility reversible (key) reversible (vault/FPE) in principle irreversible
Format preservation usually not possible (FPE) possible
Basis of safety key secrecy mapping separation loss of original
Main use protection at rest/in transit production PAN/identifiers dev/test/analytics
Regulatory effect evidence of control PCI scope reduction environment separation

3. Types and Techniques of Data Masking

Masking divides broadly into static and dynamic masking depending on "when the value is changed," and the time of application determines the architecture and threat model. The detailed flow below shows the difference in processing paths between the two approaches.

flowchart LR
    subgraph SDM["Static masking"]
        P1["Production DB"] --> P2["Extract & transform (ETL)"]
        P2 --> P3["Masked copy"]
        P3 --> P4["Dev/test environment"]
    end
    subgraph DDM["Dynamic masking"]
        Q1["Original DB (unchanged)"] --> Q2{"Requester privilege check"}
        Q2 -->|"unauthorized"| Q3["Return masked value"]
        Q2 -->|"authorized"| Q4["Return original value"]
    end

A. Static Data Masking (SDM)

Static masking extracts production data and creates a permanently transformed copy to supply to non-production environments. Its core value lies in making the original "not exist at all" in development, test, and analytics environments. Since the copy contains not a single real record, even if that environment is leaked wholesale, no breach of personal information occurs. This fundamentally turns the many non-production environments — which inevitably have looser security — into safe zones.

The hardest challenge in static masking is preserving referential integrity. If the customer ID in the customer table is masked, the foreign key in the order table must be transformed by the same rule to keep joins intact. It must also be deterministic masking that always produces the same output for the same input, so that consistency is maintained across multiple tables and multiple batches. For example, the statistical distribution of the gender and birth-year digits of a national ID must be preserved to keep tests realistic, so the transformation requires consideration not of simple substitution but of format, distribution, and constraints. For this reason, mature SDM tools provide the ability to analyze the schema to automatically discover sensitive columns and trace their relationships.

As a practical example, a typical case is a financial firm copying production data to development monthly, applying deterministic masking to roughly 20 million customer records while maintaining the integrity of dozens of foreign keys. In this case the performance of the masking batch (transforming hundreds of millions of cells) and its reproducibility become the crux of quality.

B. Dynamic Data Masking (DDM)

Dynamic masking leaves the original data as is and, the instant a query request arrives, hides the result value in real time according to the requester's privileges. The representative scene is a card number showing on a call-center agent's screen only as **** **** **** 1234. Even for the same original column, depending on privilege, one user receives the original and another receives the masked value.

Dynamic masking divides by implementation location into (1) DB-engine-embedded (e.g., a commercial DB's Dynamic Data Masking, a cloud DBaaS policy) and (2) application/proxy (gateway) types. The DB-embedded type is simple to apply but carries the risk of having masked values reverse-inferred by bypass queries (e.g., inference attacks through WHERE clauses), so in sensitive environments it must run alongside access control and row-level security. The proxy type intercepts and rewrites SQL, so it can apply consistent policy across diverse DBs, but it must manage performance overhead and single-point-of-failure risk. Dynamic masking preserves the original, so it is strong for "minimizing exposure during operation," but unlike static masking which turns an entire environment into a safe zone, one must not confuse the fact that the fundamental risk of the original remaining in place persists.

C. Masking Transformation Techniques

Concrete transformation techniques include substitution (replacing with realistic-looking fake dictionary values, e.g., name → random full name), shuffling (scrambling values within the same column to preserve the statistical distribution while breaking row linkage), nulling/deletion, variance (applying a ± error to salaries and dates), and format-preserving substitution (keeping the first 6 and last 4 digits of a card number and transforming only the middle). The choice of technique depends on the data's purpose of use — for statistical analysis, distribution-preserving types (shuffle/variance) are suitable, while for functional testing, format- and constraint-preserving types are suitable.

D. Sensitive Data Discovery and Policy

The first step of masking is data discovery to find "what is sensitive." Using regular expressions, dictionaries, and machine learning, patterns of national ID and card numbers are identified in columns and unstructured data, and masking rules are mapped according to the data classification grade. If specified manually without discovery, omissions occur every time the schema changes, so automatic discovery and policy centralization decide success or failure in large environments.

Data discovery should be operated not as a one-off task but as a continuous governance activity. It is desirable to periodically scan for situations where new tables/columns are added or sensitive information flows into unstructured logs/documents, and to build a pipeline that links discovered sensitive assets with the data catalog and data classification scheme so that masking policy is applied automatically. For example, at financial firms and telecoms with thousands of tables, it is common to discover hundreds of new sensitive columns per quarter through a full scan, and managing this by hand will inevitably produce omissions. Therefore, automating the cycle of discovery → classification → policy mapping → application → verification governs the effectiveness of masking in large environments.

4. Types and Architecture of Tokenization

Tokenization replaces a sensitive value with a token, while allowing only authorized processes to restore the original. The core design variable is "where to keep the mapping," and this splits it into vaulted and vaultless types.

sequenceDiagram
    participant A as "Application"
    participant T as "Tokenization service"
    participant V as "Token vault (isolated store)"
    A->>T: "Pass original PAN (tokenize request)"
    T->>V: "Create & store token-original mapping"
    V-->>T: "Return token"
    T-->>A: "Return token (only token stored in production DB)"
    A->>T: "Restore request (only if authorized)"
    T->>V: "Look up original by token"
    V-->>T: "Return original"
    T-->>A: "Return original"

A. Vaulted Tokenization

The vaulted type keeps the mapping table of tokens and originals in an isolated secure store (a vault). The token itself is close to a random number, so no analysis can derive the original, and restoration is possible only through a path with vault-access privileges. Safety is very high, but the vault becomes a single point and availability, scalability, and synchronization become bottlenecks. In large transaction environments, vault lookup latency and the consistency of replication/backup become the core design challenges.

Another design issue for the vaulted type is collision avoidance and determinism of token generation. Depending on whether the same original always receives the same token (deterministic) or a different token each time (random), analytical usefulness and safety differ. Deterministic tokens enable joins and deduplication across different systems, which favors analytics but is relatively exposed to frequency-analysis attacks; random tokens are safe but make identifying the same individual impossible. In an environment processing tens of thousands of payments per second, vault lookups become a major cause of transaction latency, so a design that caches frequently used tokens or replicates the vault across multiple regions while maintaining strong consistency is required.

B. Vaultless Tokenization and Format-Preserving Encryption (Vaultless/FPE)

The vaultless approach generates and restores tokens with a cryptographic algorithm, without a mapping table. Representatively, Format-Preserving Encryption (FPE) is used, which is an encryption mode that makes a 16-digit card number still output as a 16-digit number when encrypted. In SP 800-38G, the US NIST standardized FF1 and FF3-1 as FPE modes (note, however, that an early vulnerability was reported in FF3 and it was revised to FF3-1, so the latest specification must be checked at implementation). FPE removes the bottleneck of the vault, so it scales excellently and does not require changing the existing schema, but since its essence is encryption, key management again becomes important, and given its nature as a "token that is nonetheless reversible ciphertext," the extent to which it is recognized as tokenization under regulations must be confirmed in advance.

C. Scope of Application and PCI-DSS Scope Reduction

The greatest practical value of tokenization is regulatory scope reduction. If a card number is tokenized immediately after capture, the subsequent order, settlement, and CRM systems handle only tokens and are thus excluded from PCI-DSS audit scope. For example, if the segment handling original card numbers among dozens of internal systems is compressed into two places — the payment gateway and the token vault — security controls, auditing, and cost can be concentrated on those two. This is a typical case of simultaneously strengthening security and lowering TCO, and it is why tokenization is an architectural strategy that goes beyond a mere substitution technique.

5. Comparison: Encryption vs. Tokenization vs. Masking

The three techniques are often used interchangeably, but their selection criteria clearly differ, and one must understand the fundamental reason the differences arise to design correctly. Encryption is advantageous when "the original must be restored frequently and protection is needed across the entire at-rest/in-transit range"; tokenization when "restoration of the original is rare but format compatibility and regulatory scope reduction are important"; and masking when "restoration of the original is unnecessary and only usefulness is needed."

The root of the difference is where the basis of safety lies. Encryption has an "all-or-nothing" risk structure where everything is exposed if the key leaks, while tokenization has a "double-barrier" structure requiring additional theft of the vault. Therefore, even if an attacker dumps the production DB, encryption leaves residual exposure risk depending on key management, but with tokenization the attacker obtains only useless tokens. The practical implication is: card numbers and identifiers that are continuously used in operation and become leak targets are best served by tokenization; development/analytics copies that pass wholesale to external environments by static masking; and segments whose purpose is at-rest/in-transit protection such as communications and backups by encryption. In short, the three techniques are not competitors but complements, each covering a different point in the data lifecycle.

To gauge the effect with concrete figures, if an e-commerce firm with 30 PCI-DSS-in-scope systems introduces tokenization immediately after capture and reduces the systems holding original card numbers to two — the payment gateway and the token vault — the annual targets for audit, vulnerability assessment, and penetration testing shrink by roughly 93%, so compliance cost and incident exposure surface plummet simultaneously. Conversely, if the same data must be copied to a dozen or so development environments, rather than opening a token-restoration path in each environment, it is far lower in risk-versus-cost to remove the original itself through static masking. In this way, the answers to three questions — "is restoration needed, is the environment controlled, is it within regulatory scope" — determine the choice of technique.

6. Deep Dive: Cloud, Pseudonymization, and Exam Trends

A recent trend is masking and tokenization being embedded as native features of cloud data platforms. Cloud DBaaS and data warehouses provide dynamic data masking, column/row-level security, and policy-tag-based governance declaratively, universalizing the approach of "keeping data in one place but making it look different depending on who views it" in an environment where data sharing explodes. This dovetails with the data-mesh and lakehouse trend of sharing data without copying it (zero-copy).

On the regulatory/governance side, tokenization and masking are closely tied to Korea's pseudonymization. Pseudonymization is processing that makes an individual unidentifiable without additional information, and masking and tokenization become its concrete means. However, pseudonymization must be accompanied by a "statistical assessment of re-identification risk," so merely masking does not automatically make something pseudonymized information; the purpose of processing and the possibility of linkage must be judged together. This point is a deep argument frequently demanded in professional-engineer answers, and connecting masking not only as a "technique" but to "legal processing requirements" is a high-scoring point.

As for expected exam directions, the following are repeatedly covered: (1) the comparison and selection criteria of encryption, tokenization, and masking; (2) the mechanism of PCI-DSS scope reduction; (3) the architecture and risks of static and dynamic masking; and (4) linkage with pseudonymization under the Personal Information Protection Act. An effective answer-composition strategy is to present a framework that "maps techniques to the data lifecycle (collection-storage-use-sharing-disposal)" and to develop, with grounds, why a particular technique suits each stage.

7. Considerations and Implications

When designing and adopting data masking and tokenization from a professional engineer's perspective, the following must be considered in balance.

  • Application strategy and data-lifecycle mapping: No single technique solves every problem. A layered strategy should be established in connection with data classification grades — tokenizing immediately after capture to minimize original storage, static masking for non-production copies, dynamic masking for operational queries, and encryption for storage/transmission.
  • Referential-integrity vs. usefulness trade-off: The higher the protection strength, the lower the analytical and functional usefulness of the data. Preserve joins, validation, and statistical distribution with deterministic masking and format-preserving techniques, but reverse-verify whether excessive preservation increases re-identification risk.
  • Performance and availability design: The vault in vaulted tokenization and the gateway in proxy-type dynamic masking become bottlenecks and single points of failure. Non-functional requirements such as vault replication/caching, a shift to vaultless (FPE), and the performance of batch masking at the scale of hundreds of millions of cells must be reflected in the design from the outset.
  • Key/vault management and defense against bypass attacks: For FPE and encryption tokens, key management again becomes the vital point, and DB-embedded dynamic masking can be reverse-inferred by inference/bypass queries. Combination with HSM-based key management, least privilege, row-level security, and audit logs is essential.
  • Regulatory alignment and pseudonymization linkage: Confirm in advance, from a legal/regulatory perspective, the PCI-DSS scope-reduction effect and the fit with pseudonymization and safety-assurance measures under the Personal Information Protection Act, and distinguish simple masking from legal pseudonymized information while concurrently assessing re-identification risk.
  • Outlook and related technologies: The future direction is combination with the declarative masking of cloud DBaaS, the zero-copy sharing of data mesh, and "in-use protection" technologies such as differential privacy, homomorphic encryption, and confidential computing. The design capability to combine masking and tokenization with these in layers is required.

References


In one line: Data masking transforms the original into irreversible fake values (static/dynamic) to protect non-production environments and query exposure, while tokenization replaces sensitive values with vault/FPE-based tokens to keep the original separated yet recoverable and to reduce PCI scope; together with encryption, they should be layered and combined across the data lifecycle.