← Back to list
Database
#공공데이터#데이터품질#표준화#예방적품질관리#진단#132회
Last updated · 2026-10-01

Public Database Standardization Management (Preventive Quality Management, Apr. 2023)

1. Overview

A. Definition and Background

It is a preventive (ex-ante) quality-management standard that prevents errors from the information-system build stage onward, presented by the Ministry of the Interior and Safety through the "Public Database Standardization Management Manual" (Apr. 2023) for the provision, opening, and utilization of high-quality public data. Rather than the after-the-fact correction of finding and fixing errors once data has accumulated, it places weight on blocking the occurrence of errors itself by embedding standards at the design and build points.

Because public data is the source of truth for public services, policy decisions, and private-sector use (MyData, AI training data), an error in a single piece of data spreads and amplifies to the many agencies and services that link to or reference that data. For example, if the address system deviates from the standard, welfare-benefit eligibility selection, disaster text messaging, and statistical aggregation all go wrong in a chain. Whereas conventional quality management leaned toward the after-the-fact (corrective) approach of finding and fixing errors once data had already accumulated, the core of this manual is that it shifts the center of gravity toward embedding standards at the analysis and design points to stop errors before they occur.

The ideological root of this shift touches the old empirical rule of software engineering—the 1:10:100 rule—that "the later a defect is found, the more exponentially its correction cost grows." Fixing in the operation stage a defect that could have been blocked at a cost of 1 in the analysis stage costs tens to hundreds of times more. The same principle applies to data quality, and the manual is the product of systematically institutionalizing this principle in the public-data domain.

B. Necessity

After-the-fact cleansing (data cleansing) requires tracing and reverting data that has already been wrongly entered or linked, so its cost is large, and errors in data that has already been opened and distributed are even hard to recall. If data that the private sector has already downloaded and is using in services is corrected after the fact, that correction would have to propagate out to the private services again, which is practically close to uncontrollable.

Concretely, if, without standard terms, each department writes "resident registration number / resident number / resident reg. No." in its own way, then mapping and cleansing work occurs repeatedly each time systems are later linked or integrated. With 10 agencies, the linkage combinations grow to dozens, and a conversion rule must be created and verified for each combination. By contrast, if standard words, standard domains, and standard terms are fixed in advance at the analysis stage, such repeated costs and consistency errors can be blocked at the source. It is precisely here that the justification for the preventive approach of "embedding quality from the build stage" arises.

In the same vein, data quality is often divided into several dimensions such as completeness, accuracy, consistency, validity, and timeliness; whereas after-the-fact correction mainly stops at fixing "values already wrong (accuracy)," preventive management is qualitatively different in that it structurally guarantees consistency and validity at the design point. That is, the preventive approach is not about "erasing wrong values" but about "building a structure where wrong values cannot enter," and this difference creates a gap in both cost and sustainability.

2. Preventive Quality-Management Activities by Build Stage (A)

A. Overall Structure — Placing Quality Management in the SDLC

The overall picture of preventive quality management is to place a quality line of defense at each stage of the information-system life cycle (SDLC). The overall structure diagram below shows the per-stage activities together with the error types they block.

flowchart LR
  A["Analysis<br/>define standards & requirements"] --> D["Design<br/>standard-compliant modeling"]
  D --> I["Implementation<br/>apply constraints"]
  I --> T["Test & transition<br/>diagnosis & migration verification"]
  T --> O["Operation<br/>monitoring & improvement"]
  O -. periodic diagnosis & feedback .-> A
  style A fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px

The reason for placing quality-management activities across all SDLC stages is the 1:10:100 rule seen earlier. If requirements and standards are not captured at the analysis stage, errors harden into the code and data structure as they pass through design and implementation, and in the operation stage one pays the costly price of cleansing and rebuilding. Therefore each stage is designed to serve as a gate that "does not pass errors to the next stage."

B. The Principle of Per-Stage Activities

The analysis stage is the starting point of all prevention. It defines data requirements and establishes a dictionary of standard words and standard domains. If terms and meanings are not unified here, every subsequent stage wobbles, so the standardization of the analysis stage corresponds to the upstream headwater of the whole quality.

The design stage reflects the standards fixed in analysis into the data model. It keeps naming rules with standard words and designs integrity and constraint rules on entities and relationships to secure referential integrity and model consistency in advance. If normalization or identifier design goes wrong at this stage, duplication and anomalies become structurally built in.

The implementation stage actually applies the designed constraints to the DBMS. In particular, the constraints of the implementation stage are the representative case of preventive control. Placing NOT NULL, FK, CHECK, and domain constraints on columns means out-of-range values cannot be stored in the first place, so the work of finding and deleting anomalies after the fact itself becomes unnecessary. For example, placing a CHECK (gender IN ('1','2','9')) constraint on a gender column rejects the entry of typos and non-standard values at the input point.

The test and transition stage diagnoses and verifies quality and, in particular, verifies the consistency of the migration that moves legacy data into the new DB. Migration is a high-risk segment where data moves in bulk and loss or distortion easily occurs, so count reconciliation and value verification are performed thoroughly.

The operation stage manages the drift by which quality degrades over time even after the build is finished. As business change, new code additions, and exceptional entries accumulate, the initial standard-compliance rate gradually drops, so quality level is maintained and improved through periodic diagnosis and continuous monitoring. The heart of the operation stage lies in treating quality as a process, not a state, on the premise that "even data once made clean gets dirty again if neglected."

In this way, the activities of each stage are not independent but form a chained structure in which the output of the preceding stage becomes the input of the next stage. The standard words defined in analysis become the naming rules of design, the constraint rules of design become the CHECK constraints of implementation, and those constraints become the criteria of diagnosis in the test and operation stages. Therefore, if any one stage neglects the quality gate, its defect propagates to the entire downstream, so threading all stages with consistent standards is the manual's design intent.

Stage Preventive quality-management activity Purpose (what it prevents)
Analysis Define data requirements & standards, establish standard-word/domain dictionary Prevent naming/semantic confusion from the absence of standards
Design Standard-compliant data modeling, design of integrity/constraint rules Secure model consistency & referential integrity in advance
Implementation Build DB reflecting standards, apply constraints (NOT NULL·FK·CHECK) Block the entry of wrong values itself
Test & transition Quality diagnosis & verification, data-cleansing & migration-consistency verification Prevent loss/distortion in the migration process
Operation Continuous quality monitoring & improvement, periodic diagnosis Manage quality degradation (drift) during operation

3. The Four Diagnostic Areas and Nine Diagnostic Items of Preventive Quality Management (B)

A. The Flow Structure of the Diagnostic System

Diagnosis is organized into four areas following the upstream-to-downstream flow (standard → model → value/structure) of "Are the standards properly defined? → Are those standards reflected in the model? → Do the actual values and structures keep the standards and rules?" The process detail diagram below shows this dependency.

flowchart TD
  S["① Standardization diagnosis<br/>standard word·domain·term"] --> M["② Model-quality diagnosis<br/>model consistency·naming rules"]
  M --> V["③ Value-quality diagnosis<br/>mandatory values·valid values"]
  M --> C["④ Structure·integrity diagnosis<br/>referential integrity·code consistency"]
  V --> R{"Diagnosis result<br/>detect standard violations"}
  C --> R
  R -. improve·re-diagnose .-> S
  style S fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px

The heart of this flow is inter-area dependency. If the upstream (standardization) collapses, downstream (value/structure) diagnosis, however dense, cannot block the root error. This is because inspecting only values while the standard terms are wrong falls into the contradiction of "measuring correctness against a wrong criterion." That is why diagnosis must always start from the standardization area.

B. The Roles of the Four Areas

① Standardization diagnosis checks whether standard words, standard domains, and standard terms are correctly defined in advance. It is the most upstream diagnosis guaranteeing the consistency of data names and meanings. ② Model-quality diagnosis checks whether those standards are properly reflected in the data model and whether naming rules and model consistency are kept. ③ Value-quality diagnosis checks whether the actually stored values keep the mandatory-value (NOT NULL) and valid-value (domain) rules, confirming the accuracy and completeness of the data. ④ Structure·integrity diagnosis checks referential integrity and code consistency, confirming the consistency of inter-table relationships and code values.

Diagnostic area Representative diagnostic items What it guarantees
Standardization Standard word, standard domain, standard term Consistency of terms/meanings
Model quality Data-model consistency, naming-rule compliance Standard's reflection in the model & structural consistency
Value quality Mandatory-value (NOT NULL) · valid-value (domain) compliance Accuracy·completeness of actually stored values
Structure·integrity Referential integrity, code consistency Consistency of relationships·code values

C. The Three Building Blocks of Standardization

The standardization area in turn consists of three axes—standard word, standard domain, and standard term—and these three have a combinatorial relationship of word → domain → term.

  • Standard Word: the minimum unit of meaning composing a data name. For example, it defines the standard notation, English abbreviation, and forbidden words of words such as "customer," "number," and "date." If "customer" is written by one person as "顧客," by another as "customer," and by another as "cust," the names become disparate, so they are first unified at the word level.
  • Standard Domain: it defines the format, range, and type of the values a column may hold. For example, a "date" domain is prescribed as DATE·YYYYMMDD 8 digits, and an "amount" domain as NUMBER(15). By preventing data of the same nature from being stored with different types or lengths per table, it structurally guarantees the validity of values.
  • Standard Term: a business term made by combining standard words, with a standard domain connected to each term. It works as "customer number = customer (word) + number (word)" with a "number" domain (e.g., VARCHAR(10)) mapped to it. Since terms lead to actual column and attribute names, model and value quality follow only when standard terms stand right.

Because these three axes cross-reference one another, when a standard word changes, every standard term using that word is affected. That is why standardization diagnosis sits at the most upstream of all diagnoses.

Under the four areas above, advance diagnosis is performed with a total of nine diagnostic items (standard word·standard domain·standard term, data-model consistency·naming rules, mandatory values·valid values, referential integrity·code consistency, etc.).

As a concrete cross-verification case, code-consistency diagnosis checks "whether the gender-code column has any value other than the standard code set (1=male/2=female/9=unknown)," and inspects this by linking it to the standard-term and standard-domain definitions. That is, by cross-verifying so that they do not conflict with one another across value (is there a non-standard value such as 3?), structure (is it linked to the code table via FK?), and standard (does it match the standard code-set definition?), it catches errors that one area alone would miss through inter-area linkage.

As another case, in mandatory-value diagnosis, one does not simply check "whether there is a NULL" but whether a value that must exist by business rule is empty. For example, if the "receipt date" of a civil-complaint receipt table is empty, computing the processing deadline and statistics is impossible, so this column is a mandatory value in business terms. Thus diagnosis guarantees substantive quality only when it is combined with business meaning beyond a technical NULL check, and that is why designing diagnostic rules requires collaboration between data staff and the business side. Diagnosis results are usually produced as indicators such as per-item error rate and compliance rate, becoming the basis for setting improvement priorities and comparing before-and-after effects.

4. Comparison and Expected Effects — After-the-Fact Correction vs. Preventive Quality Management

To understand the value of the preventive approach, one must examine, in comparison with the existing after-the-fact correction method, "where and why the cost differs." The after-the-fact method finds errors once data has accumulated, so the discovery point is late, and accordingly the cost explodes because even the already-linked and derived data must be corrected in a chain. By contrast, the preventive method blocks errors at the input and build points, so the blocking point converges to one upstream place, and as a result the total volume of correction cost drops dramatically. The table below summarizes that difference and its reasons.

Effect Content Rationale (why so)
Quality improvement Pre-removal of errors/duplication, securing consistency Minimize error inflow by blocking at the input point
Cost reduction Reduced after-the-fact cleansing/rework cost Avoid correction cost that grows the later it is fixed (1:10:100)
Enhanced usability Higher reliability/reusability of opened·linked data Standardization eases inter-agency linkage·integration

Another essential difference this comparison reveals is whether responsibility is dispersed or concentrated. After-the-fact correction discovers errors after they have spread to many systems, so "who created the error and when" is hard to trace and responsibility is blurred. By contrast, preventive management places a quality gate and a responsible party at each build stage, so when an error arises it is clear at which stage the gate let it through, and improvement feedback is fast.

The practical implication is that "preventive investment is an upfront cost but insurance against downstream cost." Spending extra effort on standard establishment and modeling in the analysis and design stages is a burden right now, but it is an upfront investment that avoids the far larger cost of cleansing, error spread, and damage to public-service trust in the operation stage. In particular, public services carry a reputation risk in which a single data error spreads into complaints, media, and audits, so considering even the preventive effect not convertible to money, its value grows further.

5. In Depth — Linkage with Data Governance, MyData, and AI Training Data

A. Data Standardization and the Governance System

The manual's standard-word and domain dictionary has effect only when maintained and updated atop enterprise-wide data governance (standard-management organization, process, and policy). Even if a standard is made once, without a party to manage it and a change-control procedure, departmental exceptions and temporary codes increase over time and it comes apart again. That is, the manual provides only the technical criteria, and what keeps them alive is the organizational device called governance. This is why quality and governance are always treated as a pair in data-management maturity models (e.g., the data-quality and data-governance knowledge areas of DAMA-DMBOK).

B. The Quality Foundation of MyData and AI Training Data

As public data is reused in private services (MyData) and in AI training, source quality directly governs derived service and model quality. The principle "Garbage In, Garbage Out (GIGO)" is especially fatal in the AI era. This is because a model trained on biased or heavily missing public data reproduces and amplifies those defects. From this perspective, preventive management that secures standard and value quality upstream becomes not mere DB management but a strategic activity that lays the data foundation of national AI competitiveness.

C. Automation and Organizational Operation

The nine diagnostic items must be run periodically with automatic diagnostic tools (profiling, rule engines) to reduce reliance on manual human work, for it to be sustainable. Manual diagnosis makes a full inspection practically impossible once the target tables reach hundreds to thousands, but a rule engine applies the defined diagnostic rules uniformly to all data and automatically extracts error cases. At the same time, a dedicated quality-management organization and role (data steward, etc.) is placed to institutionalize the cycle of prevention → diagnosis → improvement. Automation handles the efficiency of repeated inspection, and the organization handles the responsibility of rule design, judgment, and improvement, in a mutually complementary relationship. With only tools and no organization, diagnosis results are neglected; with only an organization and no tools, inspection stops at sampling.

D. Linkage with Related Institutions and Standards

Preventive quality management does not operate alone but runs in mesh with domestic data-quality institutions and standards. The main linkage points are as follows.

  • Data Quality Certification (DQC): a certification system by which the public and private sectors have their data-quality level officially recognized, verifying quality criteria such as value, structure, and standard through external review. Because its aims touch the manual's diagnostic areas and items, an agency that steadily performs preventive management finds certification easier to handle.
  • Public-data opening·quality-management level evaluation: the government periodically evaluates each agency's public-data quality and opening level, and the more an agency has systematized preventive quality management, the better the result it obtains. That is, compliance with the manual forms a virtuous cycle with evaluation response.
  • Metadata·standard-code sharing: aligning government-wide common standard terms and common standard codes with an agency's internal standards reduces conversion cost in inter-agency data linkage and raises interoperability.

In this way, the manual's preventive management maximizes its value not merely in the internal activity of an individual system but within the institutional ecosystem of national data-quality governance.

6. Considerations and Implications (Professional Engineer's Perspective)

  • A standard without governance cannot last: Technical outputs such as standard dictionaries and diagnostic items must be operated atop a standard-management organization and change-control process. Standardization lacking governance ends as a one-off project and quality degrades again.
  • Institutionalizing the virtuous cycle of prevention-diagnosis-improvement: Do not end with a single diagnosis; build a standing system with periodic diagnosis and automation tools to continuously manage the quality drift that arises during operation.
  • Expanded quality responsibility for opening·linkage·AI reuse: The more public data is reused by the private sector and AI, the wider the source agency's quality responsibility becomes. Escaping the closed view of "data only our agency uses," one must preemptively equip quality criteria (interoperability·metadata·standard codes) premised on linkage and opening.
  • Trade-off management: Though initial standard establishment and modeling take more time and manpower, this is a far cheaper upfront investment than the cost of after-the-fact cleansing and error spread. However, so that excessive standardization does not harm the business side's flexibility, a priority strategy of applying standards stepwise starting from core common data is needed.
  • Quantifying quality indicators: Indexing data-quality dimensions such as completeness, accuracy, consistency, and validity and managing diagnosis results numerically allows objectively proving improvement effects and making them the basis for investment decisions.
  • Institutionalizing business·data-staff collaboration: Diagnostic rules (especially mandatory values and business rules) cannot be set by data staff alone and require collaboration with the business side that knows the business meaning. Rather than leaving quality management as the IT department's job alone, a governance design that grants data ownership to the business side governs long-term quality.

References


In one line: The Public DB Standardization Management Manual places preventive quality-management activities at each build stage (analysis to operation) and performs upstream→downstream advance diagnosis with the four diagnostic areas and nine diagnostic items of standardization·model·value·structure-integrity, avoiding the after-the-fact cost that grows per the 1:10:100 rule and securing the quality foundation of opened·linked·AI-training data.