← Back to list
Database
#데이터품질#품질성숙도#데이터거버넌스#정형비정형#131회#129회
Last updated · 2026-09-27

Data Quality Management

1. Overview

a. Definition

Data Quality Management (DQM) is the management activity that continuously measures, improves, and assures multidimensional quality characteristics of data—accuracy, completeness, consistency, validity, uniqueness, timeliness, and so on—through a system of policy, organization, process, and technology, and it is the foundation for making data a reliable asset for decision-making and AI.

The decisive reason data quality management matters is that 'poor-quality data leads directly to wrong decisions'. A wrong customer address causes a failed delivery, duplicate sales data distorts performance, and an AI trained on data mixed with bias and errors makes wrong predictions. The old adage "Garbage In, Garbage Out" returns as an even greater cost in the age of data and AI. Marketing sent on wrong data, wrong inventory forecasts, and regulatory reporting errors translate immediately into monetary loss and damaged trust.

The starting point for understanding quality management is defining quality by dividing it into multiple dimensions. Data quality is not a single yardstick but the aggregate of several axes: whether values match reality (accuracy), whether values that should exist are not missing (completeness), whether there is no contradiction across systems (consistency), whether it keeps the defined format and range (validity), whether there is no duplication (uniqueness), and whether it is current at the moment needed (timeliness). These dimensions are not independent; for example, forcing real-time processing to raise timeliness can make validation shallow and lower accuracy. Therefore quality management is also an activity of adjusting the trade-offs among these dimensions according to business importance.

Another important point is that data quality does not improve by chance. Quality is secured continuously only through a management system that combines architecture, process, organization, and technology. It does not end with a single spring-cleaning cleanse; continuously inflowing data must be managed at all times, and so quality management must settle in as an 'operating process', not a 'project'.

b. Need

When data is scattered across multiple systems and accumulates by different standards, inconsistency, duplication, and errors pile up. Leaving this unattended collapses trust in data utilization, so a management system equipped with standards, quality, and governance is needed. In particular, regulations such as the Data 3 Acts and MyData require accurate and reliable management of personal information and data, and as data trading and openness spread, 'data with assured quality' has come to determine asset value itself.

Quality degradation quietly accumulates and then bursts as a large cost—this too is important. Individual errors may look trivial, but as duplicate customers and inconsistent codes pile up, integrated analysis itself becomes impossible, and decisions made and AI trained in this state lose the trust of the whole organization. So quality management should be understood as a preventive investment that regularly measures and fixes before problems grow, and it is far cheaper than the cost of responding after an incident.

c. Characteristics

  • Multidimensionality: Quality is not a single measure but the aggregate of several dimensions such as accuracy, completeness, and consistency.
  • Continuity: It is not a one-off cleanse but an operating activity that manages continuously inflowing data at all times.
  • Accountability-based: Quality is maintained only when the data owner and steward are clear.
  • Measurability: It must be quantified with metrics (KPIs) to enable goal setting and improvement tracking.

2. Data Quality Dimensions and Management Architecture

a. Data Quality Dimensions

Improvement activity begins by first defining what to measure. Each dimension in the table below is converted into a measurement metric and managed. For example, completeness is quantified as the 'missing rate of required fields', uniqueness as the 'ratio of duplicate records', and timeliness as the 'delay in the latest data update'. Only once metricized can goals be set and improvement be tracked.

Dimension Meaning Measurement example
Accuracy Values match actual facts Error rate against a validated sample
Completeness No missing needed values Missing rate of required columns
Consistency No contradiction across systems/times Number of cross-validation mismatches
Validity Compliance with defined format/range Rule-violation ratio
Uniqueness No duplication Ratio of duplicate records
Timeliness Current at the needed time Update delay time

b. Management Architecture

flowchart TB
  A["Data Quality Management(DQM)"] --> V["Quality criteria·policy"]
  A --> O["Organization·governance"]
  A --> P["Process·procedure"]
  A --> T["Tools·technology"]
  V --> P
  O --> P
  P --> T
  style A fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px

The quality management architecture interlocks four layers organically. At the top is policy and criteria (quality metrics and standards) that decide what to regard as quality, and there is an organization and governance (data owners and stewards) to take responsibility for and execute it. Below that runs the process (profiling, cleansing, validation, monitoring) that actually secures quality, and this is underpinned by tools and technology (quality diagnostic tools, MDM, metadata management).

Of these four layers, the one most commonly missing in practice is organization and governance. No matter how good a diagnostic tool is adopted, if 'who is responsible for this data's quality' is not defined, then even when a problem is found no one fixes it. So designating a data owner (business lead) and a data steward (working-level manager) to clarify accountability becomes the premise of quality management success. Policy provides direction, organization provides accountability, process provides execution, and tools provide efficiency; only when the four axes run together is sustainable quality secured.

Layer Composition Key question
Policy·criteria Define quality criteria, metrics, standards What do we regard as good quality
Organization·governance Data owners·stewards, accountability Who is responsible
Process Profiling·cleansing·validation·monitoring How to secure and maintain
Tools·technology Quality diagnostics·MDM·metadata management What to make it efficient with

3. Quality Management Process and Maturity

a. Quality Management Process

flowchart LR
  D["Define quality criteria"] --> M["Measure·profile"] --> A["Root-cause analysis"] --> C["Cleanse·improve"] --> Mo["Monitor"]
  Mo -. "Continuous cycle(PDCA)" .-> M
  style Mo fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px

Quality management is not a one-off cleanse but a cyclical process. First, define (Define) quality criteria from a business perspective, and profile the data to measure (Measure) current quality. Errors found from measurement must not be fixed only at the symptom level; one must analyze the root cause (Analyze). For example, if customer address errors are frequent, one should not stop at fixing individual values but find and fix the source problem that the input screen lacks validation logic, to prevent recurrence. Then cleanse and improve (Correct), and again continuously monitor (Monitor) the metrics to alarm on deviation. This cycle is the data version of PDCA.

A particularly emphasized point is the empirical rule that prevention is overwhelmingly cheaper than after-the-fact cleansing. In the data quality field, the '1-10-100 rule' is often cited: if the cost of blocking an error at the source is 1, the cost of cleansing and recovering it after it flows downstream is far larger. The exact multiples vary by situation, but the direction alone is clear: input-stage validation and standards are more fundamental and economical than an after-the-fact spring-cleaning.

A common mistake people make in this cycle is 'cleansing first without measuring'. If you fix only conspicuous errors without measuring what is how bad, you cannot prove the improvement effect and the root cause is also left unattended. Conversely, if you make lots of metrics but do not connect them to improvement and accountability, it becomes a 'quality theater' with only a flashy dashboard. The key is to design each stage of the process to lead to the next and produce actual action.

b. Quality Management Maturity

The level of quality management differs by organization, and a maturity model is used to diagnose the current position and set the direction of improvement. The initial stage relies on individual competence for quality management, the formalization stage develops partial procedures, the standardization stage settles enterprise-wide standards and processes, and the optimization stage achieves quantitative measurement, continuous improvement, and automation. The purpose of maturity diagnosis is not the grade itself but deriving concrete improvement tasks for reaching the next stage.

Stage Characteristics Improvement focus
1 Initial Poor quality management, individual-dependent Problem awareness·criteria setting
2 Formalization Partial procedures·criteria exist Procedure documentation·expansion
3 Standardization Enterprise-wide standards·processes settled Automation·metricization
4 Optimization Quantitative measurement·continuous improvement, automation Prediction·autonomous improvement

4. Quality Criteria for Structured/Unstructured Data and Comparison

Quality criteria differ fundamentally by data type. The reason lies in 'whether it can be expressed as rules'. Structured data such as tables and code values have clear schemas and value rules, so it can be automatically validated by quantitative criteria such as accuracy, completeness, consistency, validity, and uniqueness. For example, rule violations such as 'a resident number is 13 digits' or 'the gender code allows only M/F' are caught mechanically.

The starting point for securing structured data quality is profiling. Profiling automatically scans each column's value distribution, minimum, maximum, null ratio, number of distinct values, and patterns, revealing anomalous patterns that people failed to define (e.g., 'emails mixed into the phone number column'). Grasping the current state through profiling, then defining validation rules and continuously measuring rule violations, is the standard flow of structured data quality management.

By contrast, unstructured data such as documents, images, and logs are hard to apply formalized rules to. This is because the 'accuracy' of a sentence is hard to judge by simple rules. So it uses relatively qualitative criteria such as reliability, suitability, understandability, and usability, and instead shifts the management point to metadata and labeling quality. For example, for AI training image data, rather than the image's own resolution, 'whether the label (ground truth) is attached accurately and consistently' becomes the core of quality. Here, a method of managing labeling quality quantitatively by taking the agreement (concordance) of labels among multiple reviewers as a metric is widely used.

Category Structured data Unstructured data
Criteria Accuracy·completeness·consistency·validity·uniqueness Reliability·suitability·understandability·usability
Object Tables·code values Documents·images·logs
Method Rule-based profiling Metadata·labeling quality
Automation Relatively easy Sampling inspection·human review in parallel

The reason this difference matters in practice is that a large share of AI data is unstructured. Structured data quality tools alone cannot guarantee the quality of AI training data; separate processes such as labeling guidelines, review systems, and sampling inspection are needed. That is, 'whether data is structured or unstructured' is a strategic branch point that determines what quality tools, organization, and cost structure to equip.

5. Data Quality Management Strategy and Cases

Effective quality management strategy is summarized in four directions. First, standardization of data lays the foundation of quality. Since consistency itself cannot hold unless terms, codes, and formats are standardized, "no standards, no quality." Second, a prevention-centered approach that controls quality at the source (input) stage rather than after-the-fact cleansing is far more efficient (the earlier 1-10-100 logic). Third, measuring and monitoring with quantitative metrics (KPIs) makes improvement visible. Fourth, establishing an accountability system through data ownership and stewardship secures continuity.

These four strategies have an order. Without standardization laying the foundation, prevention rules cannot be defined; without measurement metrics, one cannot know whether prevention is effective; and without an accountability system, no one moves even when metrics worsen. That is, standardization → prevention → measurement → accountability should be seen as one cycle supporting each other, and emphasizing only one of them will not produce sustainable quality.

As a concrete case, when a financial institution integrates customer data from several affiliates and the same customer is registered redundantly due to notation differences, it builds a 'single customer view' with master data management (MDM) and applies standard cleansing rules to secure uniqueness and consistency. If it then sets the post-integration duplication rate as a KPI and manages it with a target (e.g., under 1%), improvement can be tracked quantitatively.

As another case, since large-volume real-time data such as manufacturing IoT sensor logs cannot be validated one by one by people, automatic quality validation is embedded in the pipeline to filter out missing values and outliers immediately upon inflow. Automatically detecting anomalies where a sensor fault fixes a value at 0 or makes it swing, and excluding them from downstream analysis, can prevent wrong equipment predictive-maintenance judgments. In this way, the fact that the appropriate way to secure quality differs by the nature of the data (whether large-volume/real-time, whether structured/unstructured) is the core of practice.

In Korea, data quality certification schemes have operated mainly in the public and financial sectors, and the practice of having a database's quality diagnosed and certified by external criteria has taken root. Such schemes verify quality by objective criteria rather than leaving it to internal judgment alone, and thereby act as a catalyst that elevates quality management into a standing organization-wide activity.

6. Deep Dive: Data-centric AI and the Evolution of Quality Management

The most notable trend in quality management recently is data-centric AI. Whereas past AI research focused on improving model architecture, now the approach of "fixing the model and raising data quality to improve performance" is emphasized. As cases are reported in which merely cleansing label errors and raising consistency greatly improves model performance, data quality management is being re-spotlighted as a core lever of AI performance.

Second, the move to observability-based standing quality monitoring. Beyond batch validation, quality monitoring is embedded in the data pipeline to monitor freshness, distribution, and schema changes in real time and automatically detect and alarm on anomalies. This connects with the observability concept explained earlier in DataOps and shows that quality management is evolving from static diagnosis to dynamic monitoring.

Third, data contract and governance automation. Data producers and consumers agree in advance on schema and quality expectations as a contract and automatically validate it, thereby bringing quality responsibility forward to the inflow point. It is an attempt to prevent the chronic problem of producers arbitrarily changing schemas and breaking downstream, by automatically blocking contract violations at the deployment stage.

Fourth, AI-powered quality management automation is also expanding. Learning past history to automatically detect outliers, correcting missing values with statistics or models, and finding duplicate candidates by similarity—attempts to apply AI to quality management itself are increasing. However, since AI-based automatic correction risks distorting the original, a control device that records the correction history and has people review it must be put in place together.

Even when adopting such latest approaches, the basics do not change. Without the foundation of standards, organization, and process, merely switching tools to the newest does not secure quality. One must remember that the latest techniques take effect only when layered on top of a well-established quality management system.

7. Considerations and Implications

From a professional engineer's perspective, data quality management should be designed not as a single department's technical task but as part of enterprise-wide data governance. Consider the following five in balance.

  1. Data quality is the premise of AI and analytics reliability. From a data-centric AI perspective, quality management often takes precedence over model performance improvement, and managing bias and label errors is directly tied to AI fairness and safety.
  2. Prevention-centered, source control is more economical than after-the-fact cleansing. Blocking errors at the source through input-stage validation and standardization is more fundamental and cost-effective than downstream cleansing and recovery (the 1-10-100 logic).
  3. Organization and governance take precedence over tools. Unless accountability is clarified through data owners and stewards, no diagnostic tool sustains its effect. Quality management is a matter of operating culture, not technology.
  4. For real-time, large-volume data, embed automatic quality monitoring in the pipeline. Automatically validate quality and detect anomalies every time data arrives, to block problems before they spread downstream.
  5. It is directly tied to regulation and compliance. Since the Data 3 Acts and MyData require accurate and reliable data management, quality management becomes the common foundation for regulatory response and for enhancing data asset value.
  6. Set priorities on a cost-versus-effect basis. Since managing all data at the same level is unrealistic, first identify the 'critical data elements (CDE)' with the greatest impact on decision-making, regulation, and revenue and manage them intensively, while lowering the management intensity for less-important data to allocate resources efficiently.

References


In one line: Data quality management is a standing activity that measures and improves multidimensional quality such as accuracy and completeness through a policy·organization·process·tools architecture; it applies criteria by structured/unstructured type and, with standardization·prevention·measurement·accountability strategies and data-centric AI·observability, turns data into a reliable asset for decision-making and AI.