← Back to list
AI & Data
#임베딩#벡터표현#word2vec#시맨틱검색#대조학습
Last updated · 2026-09-27

Embedding (Vector Representation)

1. Overview

A. Definition

Embedding is a representation-learning technique that maps discrete, unstructured objects—words, sentences, images, users—into low-dimensional real-valued vectors in which semantic similarity is preserved as distance, and it also refers to the resulting vectors themselves. Its core property is that 'things that are similar in meaning are also close as vectors', turning symbols that machines cannot handle into coordinates in a continuous, computable space.

The fundamental reason embeddings emerged lies in the limitations of One-Hot Encoding, the traditional way computers handled text and categorical data. If the vocabulary has 100,000 words, each word is represented as a sparse vector in which only one of the 100,000 dimensions is 1 and the rest are 0. This approach has two fatal problems. First, the dimensionality explodes to the size of the vocabulary, making storage and computation inefficient. Second, because all word vectors are mutually orthogonal, even words that are close in meaning—like 'king' and 'queen', or 'Seoul' and 'Busan'—sit at exactly the same distance from one another, so the representation captures no semantic relationships at all. In other words, one-hot can 'identify' but cannot 'understand'.

Embedding tackles this problem head-on. Each object is represented as a dense real-valued vector of hundreds of dimensions, but the coordinates are learned from data so that semantically similar objects are placed close together in space. As a result, the distance or angle between vectors becomes semantic similarity itself, and it even becomes possible to infer semantic relationships through vector arithmetic, as in the famous example king - man + woman ≈ queen. The learned coordinate system is then reused as input to downstream tasks such as search, recommendation, classification, and clustering.

By analogy, embedding is the act of assigning latitude and longitude (coordinates) of meaning to the concepts of the world. Just as you can compute the distance between two cities on a map if you know their latitude and longitude, once every object is placed on the same coordinate system, the question 'what is similar to what' reduces to a distance calculation between coordinates. This simple but powerful idea is the common principle running through search, recommendation, and generative AI as a whole.

B. Necessity and Characteristics

Today embedding has moved beyond being a tool for a few researchers to become infrastructure across entire industries. Semantic search in search engines, product recommendation in commerce, RAG in generative AI, and anomaly detection in security all run on top of embeddings. In other words, embedding is 'the first gateway that places data into a semantic space', and the quality of this gateway determines the ceiling of every intelligent service that follows.

There are three reasons embedding has established itself as a foundational technology of modern AI. First, dimensional efficiency: it compresses a 100,000-dimensional sparse vector into a 300–1,000-dimensional dense vector, dramatically reducing storage and computation burden. Second, semantic preservation: similarity is expressed as distance, so it can be applied directly to proximity-based tasks such as 'search, recommendation, and clustering'. Third, transferability: an embedding learned once on large-scale data (a pretrained model) can be reused for other tasks with little data, greatly reducing training cost. Thanks to these three properties, embedding has become the common currency of today's semantic search, RAG, recommendation engines, and anomaly detection.

Looking a little more closely at its characteristics, the decisive point is that the embedding space is continuous and differentiable. One-hot representations are discrete and do not fit well with the gradient-descent learning of neural networks, but dense vectors are coordinates in a continuous space, so they can be trained and fine-tuned together with other layers through backpropagation. That is, an embedding is not a mere encoding rule but a learning parameter that the model optimizes on its own from data. Furthermore, a single embedding space applies by the same principle not only to text but to any discrete object—users, products, nodes, images—so its generality, reducing problems from different domains to the common problem of 'vector similarity search', is another core attraction of embeddings.

2. The Principle and Learning Structure of Embedding

Learning an embedding is rooted in the 'Distributional Hypothesis', the linguistic insight that "words that appear in similar contexts have similar meanings". If you train a neural network to predict the target from the context (or vice versa), the coordinates of each object (the embedding matrix) naturally fall into alignment as a byproduct. The full pipeline is composed as a continuum of 'training → vector storage → similarity search', as shown below.

flowchart LR
  D["Source data<br/>(text, images, logs)"] --> P["Preprocessing & tokenization"]
  P --> M["Embedding model<br/>(word2vec, BERT, CLIP)"]
  M --> V["Dense vector<br/>(e.g., 768 dimensions)"]
  V --> S[("Vector store<br/>Vector DB")]
  Q["Query"] --> M
  M --> QV["Query vector"]
  QV --> ANN["Approximate nearest neighbor search<br/>(ANN: HNSW, IVF)"]
  S --> ANN
  ANN --> R["Similar results<br/>Top-k"]
  style M fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px
  style S fill:#e6f4ea,stroke:#34a853

The core of learning lies in the design of the objective. word2vec, an early static embedding, proposed two approaches. CBOW (Continuous Bag-of-Words) predicts the center word from the surrounding words, while Skip-gram does the reverse, predicting the surrounding words from the center word. Since a softmax over the entire vocabulary is too computationally expensive, this is made efficient through Negative Sampling, which trains the model to distinguish true neighbors from random non-neighbors. This is also the prototype of the contrastive learning discussed later.

That said, the learned embedding space is not always ideal. In particular, transformer-based sentence embeddings exhibit an anisotropy problem in which vectors are crowded into a narrow cone-shaped region, which can create a distortion where the cosine similarity between different sentences comes out uniformly high. To mitigate this, post-processing that evens out the space through whitening, normalization, and contrastive learning is researched and applied, and a good embedding aims for an isotropic space where 'things that differ in meaning are clearly pushed far apart'.

It is also important that the act of compressing into a low dimension itself forces generalization. If you fit 100,000 words into 300 dimensions, each dimension is shared across many words, so the model is pressured to discover latent factors (latent semantic axes) such as 'gender', 'tense', and 'positive/negative' instead of memorizing individual words. As a result, each dimension of an embedding becomes a combination of semantic factors that the data itself found, not ones a human defined in advance. This is in principle the same context as feature extraction in dimensionality reduction, which is why embedding is sometimes seen as 'learned dimensionality reduction'.

From an implementation standpoint, an embedding layer is effectively a giant lookup table. Given a vocabulary size V and dimension d, you keep a matrix of size V×d and retrieve the corresponding row using the integer index of the input token to obtain its vector. The elements of this matrix are precisely what is learned; they may be learned as a byproduct of a prediction task like word2vec, or end-to-end together with an upper transformer like BERT. The decisive difference is that whereas static embeddings stop at looking up this table as is, contextual embeddings further process the looked-up initial vector with attention to produce a final vector suited to the context.

In the embedding space obtained this way, the geometric structure reflects the semantic structure. For example, a country–capital relationship ('Korea:Seoul = Japan:Tokyo') appears as a consistent direction vector, and grammatical relations such as singular/plural, tense, and gender are also expressed as regular shifts. This shows that an embedding is not mere compression but a coordinate system that encodes relationships. Note, however, that such regularity is an approximate tendency, not a strict equality, and that it depends on the statistics of the training corpus.

Therefore, embedding quality is decisively governed by the scale, quality, and representativeness of the training data. If terms from a particular domain (medicine, law, semiconductors) are rare in a general web corpus, a general-purpose embedding cannot distinguish those terms and search accuracy plummets. For this reason, practitioners either continue pretraining a general model on a domain corpus (continued pretraining), or fine-tune it with query–answer pairs to realign the domain's semantic structure. 'A good embedding' is a matter not only of model architecture but also of what it was trained on.

3. Types of Embedding

Embeddings are classified by 'what (the target) and how (whether context is reflected)' they represent. Organizing the representative lineage of development as a flow of static → contextual → multimodal gives the following.

flowchart TB
  E["Embedding"] --> W["Word Embedding"]
  E --> C["Contextual"]
  E --> X["Multimodal"]
  W --> W1["word2vec<br/>(CBOW, Skip-gram)"]
  W --> W2["GloVe<br/>(global co-occurrence)"]
  W --> W3["FastText<br/>(subword, OOV handling)"]
  C --> C1["ELMo, BERT<br/>(in-sentence context)"]
  C --> C2["SBERT<br/>(sentence embedding)"]
  X --> X1["CLIP<br/>(text-image alignment)"]
  style E fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px

This lineage is also the history of how the problem embeddings try to solve has expanded from 'the meaning of a word' to 'meaning within context', and further to 'meaning across modalities'. The further along you go, the greater the expressive power, but so too the computational cost and operational complexity, so in practice the principle is to choose the minimum expressive power the task requires.

Word embeddings (static) assign a single fixed vector to a word. word2vec obtains coordinates from a local context window, while GloVe (Global Vectors) obtains them by factorizing a global co-occurrence statistics matrix. FastText represents a word as the sum of its character n-grams, making it strong against out-of-vocabulary (OOV) words unseen in training and against languages with heavy morphological variation (including Korean particles and inflections). The limitation of static embeddings is that they collapse a polyseme into a single vector—like 'bat' (the animal) and 'bat' (the sports equipment)—and cannot distinguish context.

In Korean processing, the choice of unit for static embeddings is especially important. Korean is an agglutinative language in which particles attach to a root so its form varies widely—as in 'hakgyo / hakgyo-eseo / hakgyo-ro' (school / at school / to school)—so training at the word-phrase level causes the vocabulary to explode and sparsity to worsen. Consequently, the key in practice is to separate roots through morphological analysis, or to use FastText's subwords or subword tokenizers (such as BPE) to robustly handle unregistered words and inflected forms. This is a representative case of a language's characteristics intervening directly in embedding design.

Contextual embeddings go beyond this limitation and generate a different vector each time for the same word depending on its position in the sentence and its surrounding words. For example, the word 'sagwa' in "sagwa-reul meogeotda" (I ate a sagwa) and in "sagwa-reul deuryeotda" (I offered a sagwa)—in Korean 'sagwa' means both apple and apology—is the same vector in a static embedding, but in a contextual embedding it diverges into different vectors, one closer to the fruit and the other to the apology. Passing through ELMo (a bidirectional LSTM), the transformer-based BERT is representative, processing the whole sentence with self-attention to reflect context. For sentence- and document-level search, SBERT (Sentence-BERT), which fine-tunes BERT as a Siamese network, is widely used; it generates sentence embeddings directly and processes similar-sentence search tens of thousands of times faster than the original BERT. This is because obtaining sentence-pair similarity with the original BERT requires re-encoding every pair, whereas SBERT vectorizes each sentence just once and compares with vector operations alone.

Multimodal embeddings align different modalities into a single shared space. The representative CLIP performs contrastive learning so that an image and its describing text come close together, enabling cross-search to 'retrieve images with text'. This forms the foundation of today's generative image and video models and of multimodal RAG.

Multimodal embeddings have recently been expanding beyond text and images toward aligning speech, video, tables, and code into a single space. If every modality is represented as a common vector, cross-search that 'finds images with speech and retrieves documents with code' becomes natural. Combined with multimodal LLMs, this forms the perception layer of agents that comprehend documents, screens, and speech in an integrated way.

Meanwhile, as a practical method for obtaining sentence and document embeddings, the pooling strategy also matters. A model like BERT emits a vector per token, so these must be combined into a single sentence vector. Rather than using the vector of the special token (CLS), mean pooling—taking the average of all token vectors—is generally superior in search quality, which is why the SBERT family widely adopts it. Beyond text, embeddings also apply to graphs: node2vec and DeepWalk learn node sequences obtained from random walks over a graph in a word2vec-like manner to build node embeddings, and are used for relational representation in social networks, knowledge graphs, and recommendation. In this way, embeddings share the principle of 'learning a representation from context', with only the target of representation changing.

Type Representative models Context reflected Strengths Limitations
Static word word2vec, GloVe, FastText No (fixed vector) Light and fast, easy to interpret Cannot distinguish polysemes
Contextual BERT, ELMo, SBERT Yes (variable vector) Captures polysemy and sentence meaning High computational cost
Multimodal CLIP, etc. Yes (cross-modal) Image-text cross-search Difficulty of training data and alignment

How to validate embedding quality is also a key practical issue. Evaluation falls into two branches. Intrinsic evaluation directly inspects the semantic structure of the embedding itself using word-similarity and analogy datasets, while extrinsic evaluation evaluates it indirectly through performance on actual downstream tasks such as search and classification. Since what ultimately matters is downstream-task performance, at adoption time it is advisable to select a model by directly measuring metrics such as search accuracy (nDCG·recall@k) on your own data.

4. Similarity Measurement and Use Cases

The reason for building embeddings is, in the end, to compute 'closeness' in a vector space and turn it into utility. Once a good coordinate system has been obtained through learning, the next step is a similarity function that quantifies how similar two vectors are, and an index that quickly finds the close ones among a great many vectors. In this section, we first look at the criteria for choosing a similarity metric, then examine how it is realized in different ways across the representative applications of search, recommendation, anomaly detection, and clustering.

The value of an embedding is realized in how you measure similarity between vectors. Notably, Cosine Similarity looks only at the direction (angle) of two vectors and excludes the influence of magnitude (length), so it is used as the standard in text search where document lengths vary widely. The Dot Product reflects both direction and magnitude, so it is used when you also want to capture popularity in recommendation, while Euclidean distance (L2) measures absolute positional difference. When vectors are normalized, cosine, dot product, and L2 become effectively equivalent, so in practice one often computes with the dot product after normalization to gain speed.

The difference among the three metrics becomes clear with simple numbers. Consider a vector A=(1, 0) and B=(2, 0), which is twice as long. Since the two vectors point in exactly the same direction, the cosine similarity is 1.0, its maximum, but the Euclidean distance is 1.0, not zero, and the dot product is 2, larger than A·A (=1). In other words, if you want to treat 'same direction as automatically similar', cosine is right, and if you want to 'also reflect magnitude (strength, frequency, popularity)', the dot product is right. Using the dot product on document embeddings without normalizing them creates a bias in which long documents unfairly rank at the top—this is the practical reason cosine became the convention in text search.

Case 1 — Semantic search and RAG. Conventional keyword search, when queried with 'car breakdown', misses documents about 'vehicle defect'. Embedding search makes the two expressions into nearby vectors and realizes meaning-based matching. For example, a typical [[rag]] pipeline converts hundreds of thousands of internal documents into 768-dimensional vectors and puts them in a vector DB, maps the user's question into the same space, extracts the top 5 by cosine similarity, and injects them into the LLM prompt. Search quality is directly governed by the performance of the embedding model. In practice, hybrid search—using both embedding-based dense retrieval and keyword-based sparse retrieval such as BM25—has established itself as the standard, because dense retrieval covers meaning while sparse retrieval covers exact matches such as proper nouns, code, and numbers, so the two make up for each other's weaknesses.

Case 2 — Recommendation systems. Netflix and online shopping malls learn users and products each as embeddings of hundreds of dimensions and recommend product vectors close to a particular user's vector. 'People who viewed this product also viewed' is implemented as nearest-neighbor search over product embeddings. The cold start problem, where a new user or product has no vector and is hard to recommend for, is mitigated with metadata embeddings.

The reason embeddings are powerful in recommendation lies in the training method. The two-tower architecture, which encodes users and products each as a vector, uses actual click and purchase logs as positive pairs to learn that 'interacting user–product pairs are close, and non-interacting pairs are far'. As a result, the latent structure of taste is captured in the vectors without explicit rules, and a new product, as long as it has features, is mapped into the same space and immediately becomes a recommendation candidate. In effect, embeddings substantially compensate for the sparsity and cold-start weaknesses of collaborative filtering.

Case 3 — Anomaly detection and deduplication. By embedding logs and transaction data, one can detect vectors far from the normal pattern as outliers, or use the proximity of document and image embeddings to filter plagiarized and duplicate content at scale. At a scale of hundreds of millions of items, approximate nearest neighbor search (ANN) indexes such as HNSW and IVF-PQ achieve millisecond-level search. Exact brute-force search costs O(N), proportional to the number of candidates N, but ANN, at the price of giving up a little accuracy (recall), finds near neighbors in sub-linear time and makes large-scale services real-time.

Case 4 — Clustering and data understanding. If you embed hundreds of thousands of customer reviews and then cluster them with k-means or HDBSCAN, topic clusters such as 'delivery complaints' and 'satisfaction with quality' emerge automatically without a human attaching labels. Combined with the t-SNE and UMAP visualizations of [[dimensionality-reduction]], this becomes a powerful tool for data exploration and quality verification.

To further raise accuracy, a retrieve-then-rerank two-stage structure is used. First, an embedding-based bi-encoder quickly narrows millions of candidates to the top few dozen, and then a cross-encoder, which takes the query and candidate together as input to assign a precise score, re-ranks them. The bi-encoder is fast because it is vectorized in advance but cannot see query–document interaction, whereas the cross-encoder is slow but precise. Combining the two as stages to obtain both speed and accuracy is the standard design for large-scale search and RAG.

The core of the comparison is that 'which similarity you use' is directly tied to the meaning of the task. Cosine suits document search where length bias must be removed, the dot product suits recommendation that must reflect popularity and strength, and Euclidean suits coordinate-type data where physical proximity carries meaning. Rather than using cosine out of inertia, you should choose according to the data and objective. Also, the balance among accuracy (recall), latency, and cost is tuned through the index type and its parameters, so the choice of similarity metric and the design of the index must be treated together as one problem.

5. Deep Dive: Recent Trends and Expected Exam Directions

Embedding technology started from static word vectors, passed through contextual and multimodal, and is now converging on the direction of 'handling multiple tasks with a single general-purpose embedding while flexibly adjusting cost'. The latest currents can be organized along three axes.

First, the rise of contrastive-learning-based general-purpose embeddings. Recent embedding models raise quality through Contrastive Learning, which without labels pulls 'positive pairs close and negative pairs far'. Sentence-embedding models such as SimCSE, E5, and BGE and CLIP-family multimodal models all share this principle, aiming at general-purpose embeddings that handle search, classification, and clustering with a single vector. As the de facto standard for comparing embedding quality, multi-task benchmarks such as MTEB (Massive Text Embedding Benchmark) are referenced.

Second, variable dimensions and lightweighting — Matryoshka embeddings. Matryoshka Representation Learning (MRL) trains a single vector so that it retains its meaning even when only the front dimensions are truncated, allowing the dimension to be adjusted by use from the same model (e.g., 1,536 dimensions for precise search, 256 dimensions for filtering large candidate sets). With OpenAI's text-embedding-3 family reflecting this as a dimensionality-reduction option, the trend of flexibly trading off storage and search cost against accuracy is pronounced. Combining this with vector quantization (Product Quantization) reduces index size by tens of times.

Third, instruction-based task-specialized embeddings. Recent embedding models are evolving to prepend a short instruction telling the purpose—such as "for query search" or "for document indexing"—to the input, so that even the same sentence yields a different vector suited to the task. This is a trend in which a single general-purpose model handles multiple tasks such as search, classification, clustering, and sentence similarity, encoding queries and documents asymmetrically to raise search quality. This shows that embeddings are moving from fixed representations to representations that adapt to context and purpose.

Fourth, past and expected exam directions. In the Professional Engineer exam, expected forms are: (1) pointing out the limitations of one-hot encoding and discussing how embedding solves them; (2) contrasting the difference between static and contextual embeddings (word2vec vs BERT) from the perspective of polysemy handling; (3) explaining the position and role of embedding in the RAG architecture that runs from embedding to vector DB to ANN to LLM; and (4) describing the selection criteria and practical implications of cosine, dot product, and Euclidean similarity. An answer developed in the order 'concept → principle (distributional hypothesis) → types → similarity and applications → trade-offs and governance' achieves high completeness.

6. Considerations and Implications

Embedding is not a technology that ends once adopted; it is infrastructure whose entire lifecycle—from model selection to bias, drift, security, and cost—must be managed. From the Professional Engineer perspective, the following should be considered comprehensively when designing, operating, and evaluating embeddings.

  1. Managing the trade-offs of model and dimension selection. The larger the dimension, the higher the expressive power, but the greater the storage and search cost and the risk of distance concentration due to the 'curse of dimensionality'. Decide the embedding model and dimension according to task characteristics (search, classification, recommendation) and data scale, and where necessary use Matryoshka and PQ to quantitatively trade off accuracy against cost. If there is a lot of domain-specific data, domain fine-tuning greatly raises search quality rather than using a general model as is.

  2. Bias and fairness. Embeddings absorb the social biases of the training corpus (gender, race, occupation stereotypes) as they are, and can encode distorted relationships such as 'programmer:man = homemaker:woman'. In sensitive areas such as hiring and credit, bias measurement, mitigation (debiasing), and impact assessment are essential, and this is directly tied to [[ai-trustworthiness]] and AI governance.

  3. Embedding drift and version management. If the data distribution changes or the model is replaced, the embedding space changes and becomes incompatible with previously indexed vectors. Since changing the model requires re-embedding the entire corpus, you must manage the embedding model version, dimension, and normalization method as metadata and reflect re-indexing cost in the operational plan. In production, drift monitoring and an A/B validation framework are required.

  4. Security and privacy. Embeddings were thought to be irreversible to the original, but recent research has shown that embedding inversion can restore a substantial part of the original text. Because the embedding of text containing personal information can itself be quasi-identifying information, you should apply access control, encryption, and de-identification measures, and must review the leakage risk when transmitting sensitive data to an external embedding API.

  5. Cost and operations (FinOps) perspective. Generating embeddings for a large corpus incurs GPU cost, and vector storage and search incur memory cost. Storing 100 million items as 1,536-dimensional float32 amounts to about 600GB, so cost must be managed through dimensionality reduction (Matryoshka), quantization (PQ, binarization), and disk-based indexes. External embedding APIs are billed in proportion to call volume, so an embedding cache that avoids recomputation and a batch-processing design govern operating expenses.

  6. Related technologies and architecture design. Embedding is not a standalone technology; its value is maximized when combined with a [[vector-database]], ANN indexes (HNSW, IVF-PQ), LLMs, and caching. An integrated perspective that designs [[rag]], recommendation, and search as a single vector pipeline and jointly tunes the chunking strategy, similarity metric, and index parameters (such as HNSW's ef and M) governs performance and cost at the same time.

References


In one line: Embedding is a representation-learning technique that maps words, sentences, images, and the like into low-dimensional dense vectors in which semantic similarity is preserved as distance; it evolved beyond the limits of one-hot into static, contextual, and multimodal forms, implements semantic search, RAG, and recommendation through similarities such as cosine and dot product, and maximizes its value when combined with vector DBs, ANN, and LLMs by managing the trade-offs of dimension, bias, drift, and privacy.