← Back to list
AI & Data
#TFIDF#텍스트마이닝#정보검색#가중치#임베딩#132회
Last updated · 2026-09-30

TF-IDF (Term Frequency – Inverse Document Frequency)

1. Overview

A. Definition

A weighting technique that quantifies how well a single word represents a particular document within a document collection (corpus). The more frequently a word appears in one document (TF) while appearing rarely across the whole set of documents (IDF), the higher the weight it is given.

TF-IDF is a classic yet still powerful weighting scheme devised to convert documents into numeric vectors in information retrieval (IR) and text mining. Because a computer cannot handle natural-language sentences as they are, each word must be assigned a real value expressing "how important this word is in this document," so that a document can be represented as a point in vector space. TF-IDF is decisively distinguished from mere frequency counting in that it determines this value by multiplying two different axes: frequency inside the document (local information) and rarity across the entire corpus (global information).

B. Background and need

The simplest way to represent a document as a vector is the BoW (Bag-of-Words) model, which merely counts word-occurrence frequencies. However, this approach causes a fatal distortion. Because stopwords—such as "the/of/and" in English or the Korean particles "eun/neun/i/ga"—appear in almost every document in large amounts, judging by frequency alone makes these common words look as if they were the key terms representing the document. In reality, it is precisely such words that carry no information for distinguishing between documents.

TF-IDF solves this problem with the clear intuition that "common words that appear all across many documents have no discriminating power." A word that appears evenly in every document is useless for distinguishing one document from another, so IDF cuts down its weight; conversely, a word that appears intensively concentrated in only a particular document is given a large weight. As a result, the keyword that characterizes each document naturally stands out. Thanks to being simple, fast, and intuitively interpretable, TF-IDF has become a long-standing standard technique for search-engine ranking, document classification, keyword extraction, and recommendation.

C. Main areas of application

Because TF-IDF is a general-purpose preprocessing step that turns a document into a keyword vector, its range of use is broad. A search engine ranks documents by the sum of the TF-IDF scores of the query terms; document classification and spam filters use this vector as the input feature for classifiers such as naive Bayes and SVM. Automatic keyword/tag extraction picks the top TF-IDF words from a document and presents them as summary tags, while content-based recommendation finds "similar articles" via cosine similarity between documents. That a single weighting formula serves as the common foundation running through search, classification, extraction, and recommendation explains the enduring vitality of TF-IDF.

Behind this technique becoming the de facto standard of information retrieval since the 1970s lies the practicality that "importance can be approximated by statistics alone, without learning." In the pre-deep-learning era, both large training data and computational resources were scarce, so TF-IDF—which derives a word's importance from frequency aggregation alone, without training a separate model—was a very attractive option. Moreover, because the result can explain "why this word matters in this document" with just two numbers, frequency and rarity, it fit well with practical settings that demand a rationale for search results. This transparency and lightness are the core reasons TF-IDF cannot be discarded even today.

2. Components and Formula

TF-IDF is defined as the product of two factors, each adjusting a word's importance from a different direction. The conceptual diagram below shows the overall structure of how the two axes combine to produce the final weight and the document vector.

flowchart LR
  subgraph local["Local information"]
    T1["Frequency of word t in document d"] --> T2["TF normalization·log scaling"]
  end
  subgraph global["Global information"]
    D1["Number of documents df containing word t"] --> D2["IDF = log(N/df)"]
  end
  T2 --> W["TF-IDF = TF × IDF"]
  D2 --> W
  W --> VEC["Per-document word weight vector"]
  VEC --> USE["Similarity·ranking·classification"]

The definition and intuitive meaning of each term are as follows. TF (Term Frequency) is a local importance measuring how frequently a word is used within a single document. IDF (Inverse Document Frequency) is a global discriminating power measuring how rare the word is across the whole corpus; letting N be the total number of documents and df(t) the number of documents in which word t appears, IDF(t) = log(N / df(t)). The final weight is the product of these two.

Item Definition Intuitive meaning
TF(t,d) Occurrence frequency of word t in document d (or normalized frequency) How frequently used in this document → local importance
IDF(t) log( N / df(t) ), N=total documents, df=documents containing t How rare across the whole corpus → discriminating power
TF-IDF TF(t,d) × IDF(t) Local frequency × global rarity

There are two reasons for taking the log in IDF. First, N/df grows excessively large when there are many documents, so the log dampens that growth gently and balances the scale with TF. Second, as a word becomes more common (df↑) the value shrinks smoothly, and if it appears in every document (df=N) then log 1 = 0, so the weight of a word with no discriminating power is naturally cancelled to 0.

The reason the two terms are multiplied rather than added also follows from the principle. To be a keyword, a word must "appear frequently in this document (high TF) while being rare in other documents (high IDF)" at the same time. If either is near 0 the word has no value as a keyword, so multiplication—which yields a large value only when both conditions are satisfied simultaneously—is logically correct. With addition, even if one side is 0 the other side's value survives intact, creating the problem that a "common word that happens to appear often in this document" is overvalued.

In practice, several variants are used. For TF, rather than using the raw frequency, log scaling (1 + log TF) or normalization against the maximum frequency is often applied. Because it is hard to regard a word appearing 3 times as exactly 3 times more important than one appearing once, this reflects the diminishing marginal utility (saturation) of frequency. For IDF, a smoothing variant that adds 1 to the denominator to prevent it from exploding when df is 0—that is, the form log(N / (1 + df)) + 1—is widely used. This smoothing/normalization variant is exactly what scikit-learn's TfidfVectorizer adopts by default, and since the detailed formula differs slightly across implementations, it is safer to attach meaning to the relative ranking rather than the absolute magnitude of the values.

3. Calculation Process (Example)

The calculation is the simple flow of computing TF and IDF separately and then multiplying them. The process diagram below shows the stages from the raw corpus to the final weight vector in detail.

flowchart TD
  P0["Raw corpus"] --> P1["Preprocessing: tokenize·remove stopwords·stemming"]
  P1 --> P2["Compute per-document TF"]
  P1 --> P3["Aggregate df per word → IDF = log(N/df)"]
  P2 --> P4["TF-IDF = TF × IDF"]
  P3 --> P4
  P4 --> P5["Document = word weight vector"]
  P5 --> P6["Cosine similarity·query matching"]

It must be emphasized first that the preprocessing stage governs result quality. Only after tokenization cuts sentences into word units, stopwords are removed, and stemming/lemmatization unifies inflected forms into one, do the TF and df aggregations become meaningful. If this stage is poor, "run/ran/running" are counted as different words, the frequency of the same meaning is scattered, and the weight of the keyword is evaluated as correspondingly lower.

Let us verify the principle with concrete numbers. Assume there are N=3 documents in total, and that the word "AI" appears in 2 of them (df=2).

  • IDF(AI) = log(3/2) = log(1.5) ≈ 0.176 (based on the common logarithm with base 10; if a log value is given in the problem, use it as is).
  • If "AI" appears 3 times in document 1 (TF=3) → TF-IDF = 3 × 0.176 ≈ 0.528
  • If a stopword like "and" appears 5 times in the same document 1 but occurs in all three documents (df=3) → IDF = log(3/3) = log 1 = 0, hence TF-IDF = 5 × 0 = 0

This contrast reveals the core of TF-IDF. The weight of the higher-frequency stopword ("and", TF=5) becomes 0, while the weight of the lower-frequency keyword ("AI", TF=3) survives at 0.528. This is exactly where the principle "a common word is not a keyword" is implemented as a formula. If a word appears in only one of the three documents (df=1), IDF = log(3/1) = log 3 ≈ 0.477, gaining the greatest discriminating power. Thus a monotonically decreasing relationship holds: the smaller df is (the rarer the word), the larger IDF becomes, and the closer df is to N (the more common the word), the more IDF converges to 0.

By representing each document as a vector of per-word TF-IDF values in this way, one can compute cosine similarity, which measures similarity by the angle two document vectors form, or a query-document matching score, computed as the inner product of the query vector and the document vector. The reason for using cosine similarity is to remove the effect of document length. A plain inner product grows larger the longer the document, but after normalizing the vectors and looking only at the angle, one can compare purely "how much they point in the same direction," that is, the similarity of word composition. The basic principle by which a search engine ranks documents for a query, news clustering groups similar articles, and a recommender system finds similar documents is all rooted in this vector-similarity computation.

4. Characteristics, Limitations, and Alternatives

The strength of TF-IDF lies in being a statistical method requiring no training, so computation is light and fast, and a person can interpret why each word got its weight. The fundamental limitation, however, is that because it treats words only as mutually independent atomic symbols, it does not understand semantics at all.

Category Content Reason
Strength Simple·fast, easy to interpret, keyword-extraction effect Statistical frequency-based, so no training needed
Limitation Semantics·context·word order not reflected Treats words only as atomic tokens
Limitation Cannot handle synonyms·polysemy Sees "car" and "automobile" as different words
Limitation High-dimensional sparse vector Dimensions as large as the vocabulary, mostly 0
Alternative Embeddings such as Word2Vec·BERT Learn context·meaning as dense vectors

For example, "bank" in "bank deposit" and "bank" in "river bank" have entirely different meanings, but TF-IDF treats them as completely the same word because the spelling matches and cannot distinguish the context (the polysemy problem). Conversely, "automobile" and "car" mean essentially the same thing, yet are treated as different words, so similarity is not captured properly (the synonym problem). Moreover, "the cat chased the mouse" and "the mouse chased the cat" have the same word composition, so their TF-IDF vectors are identical while the meanings are opposite. This reveals the fundamental limitation that TF-IDF completely ignores word order and sentence structure.

In addition, the vector dimension grows as large as the vocabulary dictionary, but only a tiny fraction of those words appear in a single document, so a high-dimensional sparse vector with most elements equal to 0 is produced, lowering memory and computational efficiency. In vectors of tens of thousands to hundreds of thousands of dimensions, it is common for only a few hundred elements to have actual values. To overcome these problems of semantics, context, and dimension, embeddings (Word2Vec, GloVe)—which learn words as low-dimensional dense vectors of a few hundred dimensions and place semantically close words near each other in vector space—and contextual embeddings (BERT, Transformers)—which represent even the same word with different vectors depending on context—emerged.

5. In Depth: BM25 and RAG Hybrid Search

TF-IDF may look like an outdated technique, but its direct descendants and derived methods are still used at the front line. The most representative is BM25 (Okapi BM25). BM25 finely compensated for TF-IDF's weaknesses by introducing a term-frequency saturation term that prevents unbounded growth of TF, and document-length normalization that keeps a long document from being favored simply for being long. Thanks to this, BM25 is adopted as the default ranking function of mainstream search engines such as Elasticsearch, OpenSearch, and Lucene, and serves as a strong baseline for pure frequency-based methods in many information-retrieval benchmarks.

BM25 has two tuning parameters: k₁ (typically 1.2–2.0), which sets the degree of frequency saturation, and b (typically 0.75), which sets the strength of document-length normalization. The larger k₁ is, the longer the additional weight persists when a word appears many times, and the closer b is to 1, the stronger the penalty on document length. Unlike TF-IDF, which was a fixed multiplication formula, the fact that BM25 can tune its ranking behavior to the characteristics of the corpus is a major reason for its practical adoption.

Recently, in the RAG (Retrieval-Augmented Generation) pipelines of generative AI, the TF-IDF/BM25 family is drawing attention again. Semantic vector (embedding) search captures synonyms and context well but is weak at exact keyword matching for proper nouns, product codes, and numbers, whereas BM25 is strong at keyword matching but does not understand meaning. For example, if a user queries "error code ORA-00942," embedding search may drift toward semantically similar documents like "database error," but BM25 pinpoints exactly the document containing that code. So hybrid search, which runs both methods together and fuses the scores with a technique such as RRF (Reciprocal Rank Fusion), has effectively become the standard. It is a strategy that raises search accuracy by mutually compensating for each side's weaknesses, and a core practical technology for making an LLM retrieve the right supporting documents.

Looking at real industrial applications, large-scale e-commerce and technical-document search systems place BM25 on fields where exact matching is decisive—product names, model numbers, statute-article numbers—and embeddings on grasping the intent of natural-language queries, then fuse the two results. In-house knowledge-base chatbots likewise commonly combine keyword search so as not to miss the proprietary terms of company rules and manuals. Thus the descendants of TF-IDF have stepped back as a standalone technique, but remain an indispensable component as one axis of a hybrid configuration.

6. Considerations and Implications

  • Preprocessing governs quality: The quality of tokenization, stopword removal, stemming, lemmatization, and normalization determines the TF-IDF result. In particular, Korean is an agglutinative language where particles and endings attach, so morphological analysis must be performed first; neglecting it scatters the same word into multiple forms and distorts the weights.
  • Still a valid foundational technology: Even as semantic embeddings have advanced, TF-IDF/BM25 require no training data or GPU, are interpretable, and are a trustworthy baseline. In cold-start situations, small corpora, or domains where explainability matters, they can be more practical than embeddings.
  • Hybrid is the mainstream: Keyword search (BM25) and semantic search (embeddings) are not substitutes but complements. In RAG and enterprise search, a hybrid structure combining the two is advantageous in both accuracy and cost, so from the professional-engineer perspective the right thing is to design "how to combine them" rather than "what to replace them with."
  • Cost·scalability trade-off: TF-IDF has low indexing/query cost and is easy to update incrementally, whereas embeddings incur large costs for model inference and vector-DB operation. The technique must be chosen by weighing corpus size, query volume, latency requirements, and explainability requirements together.
  • Room for domain specialization: Tuning the stopword dictionary, weighting variants, and per-field weighting (title vs. body) to the domain can greatly raise the performance of TF-IDF/BM25, so being a tunable, transparent algorithm is itself a practical strength.
  • Importance of evaluation·validation: Which combination of weighting variant and preprocessing is best differs by domain, so it must be quantitatively validated with search-quality metrics such as Precision, Recall, and NDCG. Since "theoretically promising" tuning often degrades performance on real query logs, running A/B tests alongside offline evaluation is a practical principle from the professional-engineer perspective.

References


In one line: TF-IDF is a technique that computes word importance as TF (frequency within a document) × IDF (log N/df) to give high weight to keywords that appear often in a particular document but rarely across the whole set; it has the advantage of automatically cancelling common words to 0 but cannot reflect semantics or context, so it evolved into embeddings and BM25, and today it is a core baseline technology used alongside semantic search in RAG hybrid retrieval.