← Back to list
AI & Data
#어노테이션#라벨링#컴퓨터비전#데이터#AI학습#134회#126회
Last updated · 2026-09-28

Image Data Annotation (Image Data Annotation)

1. Overview

A. Definition

The task of attaching ground-truth labels for objects, regions, attributes, etc. (Ground Truth) to images for the supervised learning of computer vision models.

Annotation is the process of making the textbook from which a model learns "what the correct answer is." A supervised-learning model reduces prediction error against the answers people assign, so the labels become the very reference point of learning. Therefore the accuracy and consistency of annotation govern model performance no less than the algorithm itself. This is because no matter how sophisticated the neural network architecture, if it learns from wrong answers it learns wrong rules.

Annotation is easily seen as simple labor, but in reality it is an engineering activity that combines domain knowledge, judgment, and quality control. For example, deciding how far the boundary of a lesion extends in a medical image requires a specialist's judgment, and how to label a "partially occluded pedestrian" in autonomous driving differs by worker unless there is a clear rule. Thus the quality of annotation ultimately depends on the rule design and quality control of "what to label and how."

Moreover, annotation is also the activity of creating the ground truth for validation and test data that evaluate the model, not just training data. If the evaluation answers are inaccurate, the model's actual performance is mismeasured, misjudging a bad model as good or committing the opposite error. Therefore annotation quality can be called the foundation of an AI system, affecting both training performance and evaluation reliability at once.

B. Background and Necessity

As the saying "garbage in, garbage out" goes, model performance is directly governed by label quality. Even with the same algorithm, if the answers are inaccurate or vary from worker to worker with differing criteria, the model learns wrong patterns. This is also why the center of gravity of recent AI development has shifted from improving model architecture to improving data quality (Data-centric AI). Supplying cleaner, more consistent labels to the same model often contributes more to performance improvement than fine-tuning the model architecture.

In particular, as fields requiring precise recognition—such as autonomous driving (misrecognizing a pedestrian links directly to accidents), medical imaging (missing a lesion links directly to misdiagnosis), and defect inspection—have expanded, systematically securing large volumes of high-quality labels has become a core bottleneck of AI development. In fact, there are many reports that data collection, cleaning, and labeling consume a substantial portion of the entire development period in commercial AI projects, which shows that annotation is not a one-off task but an asset-building process that must be managed continuously. Furthermore, since data must keep being augmented for new situations (snowy roads, night, new road signs) even after the model is deployed, annotation becomes not a one-time job but a cyclical pipeline.

2. Classification of Annotation Types

Understanding annotation types means understanding "in what form the answer will be expressed." Even for the same object, expression forms are diverse—rectangle, polygon, point, line, 3D box, etc.—and each form differs in the kind and precision of information it can hold. Because the expression form determines the upper bound of information a model can learn, type classification is not mere terminology tidying but the design of a learnable capability.

Image annotation is classified by expression dimension (2D/3D) and expression method (box, polygon, point, line, etc.). The concept diagram below shows the whole classification scheme at a glance.

flowchart TB
  A["Image Annotation"] --> B["2D"]
  A --> C["3D"]
  B --> B1["Bounding Box"]
  B --> B2["Polygon / Segmentation"]
  B --> B3["Keypoint / Landmark"]
  B --> B4["Polyline"]
  C --> C1["3D Cuboid"]
  C --> C2["Point Cloud"]
  style A fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px

The key is to understand annotation types as a trade-off between expression precision and cost. The higher up (bounding box), the faster and cheaper the work but the coarser the shape representation; the lower down (segmentation, 3D), the more precise at the pixel/spatial level but the more sharply the work time and cost rise. For example, drawing one bounding box on an image takes only a few seconds, but precisely tracing the same object with pixel-level segmentation takes dozens of times longer. Therefore the principle is to choose the type according to the precision the target task requires.

The important principle here is not "as precise as possible" but "only as precise as needed." Labeling to a precision the task does not require only raises cost without leading to performance improvement. Conversely, if precision is insufficient, the model cannot learn the information it needs. Ultimately, type selection must start from analysis of the task requirements, which shows that annotation is an activity requiring design judgment, not mechanical work.

The whole flow by which annotation leads to actual model training can be organized as a pipeline below. This flow does not end in one direction but has a cyclical structure in which errors filtered out in review and validation return to labeling. An especially important stage in this cycle is the frontmost label guide definition. If the rules for which objects to target for labeling and how to handle occluded objects and boundary cases are not clear, all subsequent work wavers. In other words, the pipeline's quality is already largely determined by upstream rule design.

flowchart LR
  R["Raw image collection"] --> G["Label guide definition"]
  G --> L["Labeling(manual/semi-auto)"]
  L --> Q["Quality review(IAA/sampling)"]
  Q -->|error rejected| L
  Q -->|passed| D["Training dataset"]
  D --> T["Model training/evaluation"]
  T -.select hard samples.-> L

3. Characteristics by Type

Each type corresponds to a different computer vision task. Since type selection is decided not by aesthetic taste but by the kind of information the task requires, it is important to understand what information each type holds and which task it fits. Below we examine, by representative type, the information held, the corresponding task, and points to watch in practice.

Bounding Box is used for object detection, where knowing "roughly what is where" suffices. Because it expresses position with only four rectangle coordinates, the work is fast and cheap, but for an object lying diagonally or of irregular shape, the background is also included in the box, lowering precision. Nonetheless, it remains the most widely used basic type in most applications where speed and cost matter, such as real-time detection. However, detailed rules are needed—such as how to box each overlapping object and how far to include objects clipped off-screen—and without such rules the same scene is labeled differently by each worker.

Polygon and semantic segmentation are used for shape recognition and medical imaging, which need an object's exact outline and pixel boundary. A polygon traces the outline with polygon vertices, and semantic segmentation assigns a class to every pixel. For example, quantitatively measuring the area and shape of a tumor in a medical image requires a pixel-level boundary, so a box cannot substitute. Instead, the workload is heavy, requiring considerable time and skill to complete a single image.

Meanwhile, distinguishing semantic segmentation from instance segmentation is also important. Semantic segmentation only classifies "this pixel is a car" and does not distinguish multiple objects of the same class, whereas instance segmentation distinguishes individual objects, such as "car 1, car 2." In tasks such as crowd analysis, where overlapping people must be counted individually, instance-level labels are needed, so even for the same segmentation the level fitting the task must be defined precisely.

Keypoint is used for pose estimation and face recognition, where the positional relationship of joints and landmarks matters. Marking a person's joint positions with points to represent the skeleton allows recognizing motion. For example, a fitness app's exercise-posture correction or a fall-detection system judges posture from the angles and connection relationships of joint keypoints. Because a keypoint's point order and meaning (left shoulder, right elbow, etc.) must be defined consistently, strictly stipulating the point definitions in the guideline is key to quality.

Polyline is used for thin, long linear objects such as lane lines, and because a box or polygon has difficulty expressing the continuity of a line, a separate type is needed. When a lane line is dashed or partly occluded, without a rule for how far to continue drawing the line, results vary greatly by worker. 3D cuboid and point cloud are used for autonomous-driving LiDAR, which needs distance and depth, and because they must hold depth (z) information absent from a 2D image, they belong to the highest tier of work difficulty and cost.

Type Description Representative use
Bounding Box Marks an object with a rectangle Object detection
Polygon Precisely marks the outline with a polygon Shape recognition·instance segmentation
Semantic segmentation Pixel-level class classification Road/background segmentation, medical imaging
Keypoint Marks joint/landmark points Pose estimation·face recognition
Polyline Marks linear objects Lane lines·road boundaries
3D cuboid Marks with a three-dimensional cuboid Autonomous-driving LiDAR·depth recognition

For example, to avoid a collision with the car ahead in autonomous driving, one needs not only where the target is on the screen (box) but also how many meters ahead it is, so a LiDAR 3D cuboid is required. Conversely, when only judging "is a person inside the zone," as in CCTV intrusion detection, a bounding box suffices. Thus type selection is decided by the physical requirements of the task, and choosing a more precise type than necessary only wastes cost.

4. Annotation Techniques

The history of annotation-technique development is precisely the history of "how to reduce human labor." Because the limits of pure manual work became clear as data scale grew, techniques have evolved toward pushing the point of human intervention progressively later. Below we examine the spectrum from manual to active learning along with its principles.

Manual labeling has the fundamental limits of being accurate but slow and expensive, so techniques have evolved toward reducing human intervention. Early on, people labeled every image one by one, but as data scale grew to hundreds of thousands or millions of images, this approach became unsustainable in cost and time. However, since full automation carries the risk of errors, the mainstream now is the Human-in-the-loop form in which AI makes drafts and people review them.

Manual is the method in which people label directly; accuracy is high and complex judgment is possible, but it is high-cost and slow. It suits building initial datasets that are small in scale or require high-difficulty judgment. In particular, the initial seed data for training an automation model must be high-quality manual labels carefully made by people, so that subsequent semi-automation and automation work reliably. Semi-auto is the method (HITL) in which people review and correct results the AI predicted in advance; because instead of drawing from scratch people "fix the AI draft," it greatly raises work speed.

Auto-labeling is the method of automatically generating labels with a pre-trained model and then sample-validating only a portion. It can be applied quickly to large volumes of data, but there is a risk of hardening the model's mistakes as correct answers, so quality monitoring is essential. Care must be taken not to fall into a vicious cycle where a model trained on wrong auto-labels again generates wrong labels.

The reason HITL became mainstream is that combining automation's speed with human judgment is the most practical. Full manual is accurate but not scalable, and full automation is fast but not trustworthy. HITL takes the merits of both, with AI handling repetitive, mechanical draft work and people handling final judgment, exception handling, and quality validation. A key benefit here is that feeding people's review results back into model improvement steadily raises the quality of the auto-drafts, so the proportion of human intervention can be reduced over time.

Technique Description
Manual People label directly — accurate but high-cost·slow
Semi-auto AI pre-predicts, then people review·correct (HITL)
Auto-labeling Auto-generated by a pre-trained model, then sample-validated
Active Learning Preferentially label samples the model is uncertain about

In particular, Active Learning is a core strategy for efficiency. Instead of labeling all data equally, it selects the samples the model is most confused about (low prediction confidence) and preferentially assigns them to people for labeling. Labeling more of the easy samples it already gets right yields negligible performance improvement, whereas labeling the boundary cases the model finds hard, being high in information, yields greater performance improvement for the same budget. This is a strategy of concentrating a limited labeling budget on data of high information value, and its effect grows the more vast the data.

Active learning is implemented in practice as a repeating loop of "label→train→select uncertain samples→additional labeling." After making a model with a small initial amount of data, one preferentially picks the data that model is not confident about, labels it, and retrains, repeating the process. This approach is known to greatly reduce total cost especially in expert domains with high labeling unit cost (medical, legal imaging, etc.), and it well demonstrates the data-centric AI philosophy that strategically choosing "which data to label" matters more than blindly increasing data.

5. Comparison and Industry Application Cases

When setting an annotation strategy, one must judge by placing the three axes of precision, cost, and speed together. The table below compares representative types by these criteria. The values in the table should be understood as relative tendencies rather than absolute figures, and they vary by domain, tool, and skill level.

Type Precision Relative cost Work speed Representative fitting task
Bounding box Low Low Fast Real-time object detection
Polygon High Medium Moderate Instance segmentation
Semantic segmentation Very high High Slow Medical imaging·precise segmentation
3D cuboid High(spatial) Very high Slow Autonomous-driving LiDAR

The reason this difference arises is that each type differs in the degree of freedom of the information it expresses. A box's expression ends with four coordinates, but segmentation must make a judgment for every pixel of the image, so the amount of information to express is itself tens of thousands of times larger. Therefore the sharp rise in cost and time as precision goes up is not a matter of tools but an intrinsic characteristic of the expression method. In practice, considering this trade-off, a mixed strategy is sometimes used of processing most data with low-cost boxes and applying segmentation only to the small amount of data where precision is decisive.

Looking at three cases of how requirements diverge by industry makes the implications clear. First, in autonomous driving, distance cannot be judged with camera 2D boxes alone, so a 3D cuboid is attached to the LiDAR point cloud, and multiple sensors and multiple types are combined here, adding lane lines (polyline) and traffic lights (keypoint). Because safety is the top priority, it is a representative domain that accepts the cost and chooses the highest precision.

Second, in medical imaging, since a lesion's area and volume must be measured quantitatively, pixel-level segmentation is essential, and reading requires specialist-level domain knowledge. Therefore high-unit-cost expert review is unavoidable, and accuracy and reproducibility are valued more than data scale. Since the opinions of several reading physicians can differ, a procedure of synthesizing multiple experts' labels to make a consensus answer is sometimes also required.

Third, in manufacturing defect inspection, the position and shape of tiny defects matter, so polygon and segmentation are used, but because the normal/defect ratio is often extremely imbalanced (defects are rare), deliberately securing rare-class data becomes key. Thus even for the same annotation, the optimal strategy changes completely depending on the physical and economic constraints of the industry. Ultimately, the three cases together show that an annotation strategy is not a single fixed answer but a problem of designing by synthesizing the domain's risk level, cost structure, and data characteristics.

6. In Depth: Foundation Models and the Evolution of Auto-Labeling

The biggest recent change in the annotation field is the emergence of foundation models. Representatively, Meta's SAM (Segment Anything Model), with its pre-trained general-purpose segmentation capability, instantly generates the pixel boundary of an object when the user designates just one point or one box. Segmentation, which in the past required plotting a polygon point by point across dozens of points, now has drafts made in a few clicks, greatly improving work productivity.

Such prompt-based segmentation is changing the very user experience of annotation tools. Instead of drawing a detailed outline directly, the worker only needs to "point at" the target and works by confirming and correcting the boundary the model proposes. This lifts labeling from low-skill repetitive work to judgment-centered work, enabling the same headcount to process far more data.

This change is fundamentally altering the human role. People are shifting from being the subject who generates labels from scratch to being supervisors who review, correct, and quality-manage the drafts the AI makes. In other words, it is not that the total amount of labor decreases, but that the nature of labor shifts from "production" to "validation." Accordingly, annotation organizations too are trending toward being reorganized around a small number of skilled reviewers and automation pipelines rather than around large numbers of labelers.

However, it is also emphasized that automation is not a panacea. Foundation models are strong on general objects, but on special-domain data such as medical and industrial defects they still have many errors, so expert review is essential. Also, without a system to validate whether the quality of auto-generated labels can be trusted (sample audits, confidence-based filtering), there is instead a risk of mass-producing wrong labels. Therefore introducing automation should be approached not as a matter of "removing people" but of redesigning "where to concentrate human judgment."

Another trend to fundamentally reduce the labeling burden is synthetic data and weak supervision. Making images with a simulator or generative model allows automatically obtaining ground-truth labels at generation time, so data for rare or dangerous situations (accident scenes, rare defects) can be secured at low cost. However, if the distribution difference (domain gap) between synthetic and real data is not managed, performance can drop in the real environment, so mixing with and calibrating against real data is also required. These techniques show that the future of annotation is heading not toward "people plotting more" but toward "securing answers more cleverly."

7. Considerations and Implications

The core from a professional engineer's perspective is how to design the balance of quality and efficiency, and this is a problem of designing the whole pipeline, not selecting a single technique.

  • Building a quality-management system: Consistency must be ensured through clear labeling guidelines, cross-validation, and measurement of Inter-Annotator Agreement (IAA). In particular, IAA quantifies how much results agree when several workers label the same image, and it is a key indicator that reveals early the ambiguity of guidelines or a lack of worker training. When IAA comes out low, the correct order is to first check whether the guideline is ambiguous before blaming the workers.
  • Phased introduction of efficiency strategies: It is realistic to secure high-quality seed data manually at first and, on that basis, sequentially introduce semi-automation, automation, and active learning to lower cost. Since attempting full automation from the start collapses quality, an approach of gradually raising the automation ratio is safe.
  • Managing data bias: If a particular group or situation is under- or over-represented in the data, the model learns bias. For example, an autonomous-driving model trained only on daytime images becomes weak at night recognition, so the dataset's representativeness (including diverse conditions and groups) must be designed deliberately.
  • Privacy and regulatory compliance: When personal information such as faces and license plates is included, de-identification (masking, blurring) is essential, and regulations such as the Personal Information Protection Act must be complied with. When labeling is outsourced, security controls to prevent data leakage must also be considered.
  • Continuous data operations (MLOps): Annotation is not a one-off but a cyclical activity of continually labeling and augmenting new data in response to performance degradation (data drift) after model deployment. Therefore a perspective of designing annotation as a standing component of the MLOps pipeline, not as a separate task, is needed.
  • Cost/organization design and governance: Because large-scale labeling often uses outsourcing and crowdsourcing, governance including worker training, evaluation, reward systems, and data-access control governs both quality and security at once. Optimizing the cost allocation between automation-tool investment and workforce operation according to data scale and precision requirements is also a strategic task the professional engineer must judge.

References


In one line: Image annotation is the task of assigning answers matched to task precision, such as bounding box, polygon, segmentation, keypoint, and 3D; label quality governs model performance, and as it evolves manual→semi-auto→active learning·foundation models, IAA-based quality control, bias, and privacy must be controlled together.