← Back to list
AI & Data
#PLM#LLM#사전학습#SFT#RLHF#133회
Last updated · 2026-09-29

From PLM (Pre-trained Language Models) to LLMs

1. Overview

A. Definition

A PLM (Pre-trained Language Model) is a transfer-learning-based language model (BERT and early GPT families) that pre-trains general language patterns on a large corpus and then fine-tunes for individual downstream tasks. An LLM (Large Language Model) scales that PLM up by tens to hundreds of times in parameters, data, and compute, and layers instruction learning and human-preference alignment on top to become a general-purpose model that follows instructions.

The core idea of a PLM is transfer learning: "learn the general knowledge of language once at large scale, then adjust only a little for each task." Unlike the prior approach of building a model from scratch for every task, it reuses context, grammar, and common sense gleaned from vast text, so it achieves high performance even with little labeled data. This is the migration into natural-language processing of the strategy from computer vision, where CNNs pre-trained on ImageNet were reused across many tasks; since the arrival of BERT and GPT in 2018 it has become the de facto standard paradigm of NLP.

Two practical constraints drove the need for this paradigm. First, securing large volumes of labeled data per task was costly and had to be repeated wastefully for every domain. Second, static embeddings like Word2Vec and GloVe assigned a single fixed vector to each word, suffering the polysemy limitation of being unable to distinguish "bank (financial)" from "bank (river)." Pre-training solved the first problem by accumulating language knowledge from unlabeled text alone, and the second through context-based dynamic embeddings.

B. Characteristics of PLMs

The two pillars that made PLMs possible are self-supervised learning and the Transformer. Self-supervised learning creates a training signal from text itself, without humans attaching labels. BERT uses MLM (Masked LM), masking part of a sentence and predicting it, to see left and right context simultaneously, making it strong at sentence understanding (classification, named-entity recognition, etc.); GPT gains generative ability through an autoregressive approach that predicts the next token while seeing only the preceding context. This difference in objective function is precisely what created the difference in character between the BERT family (understanding-centric) and the GPT family (generation-centric), and as LLMs moved toward conversation and generation, the autoregressive approach became mainstream.

The Transformer's Self-Attention computes the relationships of all token pairs in a sentence in parallel, representing a word not by a fixed dictionary meaning but as a dynamic embedding that varies with context. Unlike RNNs and LSTMs, which processed tokens only sequentially—forgetting early information in long sentences and being hard to parallelize—Self-Attention connects relationships directly regardless of distance and enables large-scale parallel training on GPUs. This parallelism is the very hardware premise that makes the scaling law, described below, hold in practice.

Characteristic Content
Transfer learning A two-stage pre-training → fine-tuning structure reuses knowledge
Self-supervised learning Learns without labels via MLM (BERT) / next-token prediction (GPT)
Contextual embeddings Represents words dynamically by context (handles polysemy)
Transformer Self-Attention-based parallel training enables scaling up

2. The overall structure of the PLM → LLM evolution

The evolution from PLM to LLM was not a single leap but two overlapping currents: scaling up and adding an alignment stage. The conceptual diagram below shows the entire lineage from static embeddings through PLMs to aligned LLMs, and the architecture detail diagram that follows shows the flow of how tokens are processed inside the Transformer.

flowchart LR
  W["Static embeddings (Word2Vec/GloVe)"] --> B["PLM: pre-training + fine-tuning"]
  B --> G["GPT family (autoregressive generation)"]
  B --> E["BERT family (bidirectional understanding)"]
  G --> SC["Scaling (parameters/data/compute)"]
  SC --> AL["Alignment (SFT + RLHF/DPO)"]
  AL --> L["LLM (general instruction following)"]
flowchart TB
  T["Input tokens"] --> EM["Token + position embedding"]
  EM --> SA["Multi-Head Self-Attention"]
  SA --> AN["Residual connection + normalization"]
  AN --> FF["Feed-Forward Network"]
  FF --> AN2["Residual connection + normalization"]
  AN2 --> NX["Repeat for next layer (N layers)"]
  NX --> OUT["Next-token probability distribution"]

The key point of the first lineage diagram is that an LLM is a "scaled-up version" of a PLM while at the same time containing a qualitatively new stage (alignment). A giant model that has only finished pre-training merely strings together the next token smoothly and does not "follow instructions the way people want." For example, given the input "summarize the meeting notes," instead of producing a summary it keeps generating similar request sentences. Closing this gap is the job of the stages that follow.

As the second architecture diagram shows, a Transformer block repeats, across N layers, the process of gathering contextual information with Self-Attention and transforming it with a Feed-Forward Network, and each layer stabilizes deep training with residual connections and normalization. An LLM stacks dozens to hundreds of these blocks and widens each layer's hidden dimension to scale the parameters to the hundreds of billions.

Difference between PLM and LLM

PLMs and LLMs lie on a continuum, but they show a clear qualitative difference in how they are used. A PLM is closer to a "per-task specialist" that requires separate fine-tuning and labeled data for each task, while an LLM is closer to a "general-purpose assistant" where a single model handles translation, summarization, coding, and counseling merely by changing the prompt. This difference arose as scale and alignment combined and In-context Learning emerged.

Category PLM LLM
Scale Hundreds of millions to billions of parameters Tens of billions to hundreds of billions of parameters
Task adaptation Requires per-task fine-tuning Responds instantly via prompt (In-context)
Training stages Pre-training + fine-tuning Pre-training + SFT + alignment (RLHF/DPO)
Representative ability Sentence understanding/classification Generation, reasoning, instruction following, dialogue
Form of use Per-task specialist model General-purpose foundation model

The practical implication is a shift in how systems are developed and operated. In the PLM era, a separate model was built, deployed, and maintained for each of sentiment analysis, document classification, and so on; in the LLM era, many tasks can be handled by sending different prompts to a single API, reducing model-management burden and making new competencies like prompt engineering and RAG important.

3. The PLM → LLM training process

flowchart LR
  D["Large corpus"] --> P["Pre-training<br/>Pre-training"]
  P --> S["Supervised fine-tuning<br/>SFT"]
  S --> R["RLHF / DPO<br/>human preference alignment"]
  R --> L["LLM"]

An LLM is not merely a scaled-up PLM; it is completed by layering a new stage called alignment on top of pre-training. Each stage has a distinctly different purpose, and it forms an "inverted pyramid" structure in which, the later the stage, the smaller the data scale but the greater the weight of data quality and human involvement.

Stage Training characteristics Purpose
Pre-training Acquires language/world knowledge via self-supervision (next-token prediction); enormous parameters, data, and compute Forms the foundation of knowledge and language ability
Supervised fine-tuning (SFT) Learns from human-created instruction–response data Grants the ability to understand and carry out instructions
Alignment (RLHF/DPO) Human-preference reward model (RLHF) or direct optimization on preference pairs (DPO) Aligns behavior to be helpful, honest, and harmless (HHH)

A. Pre-training

The model predicts the next token over tens of billions to trillions of tokens of text, and in this process grammar, facts, and reasoning patterns are compressed into the parameters. The training data is composed by refining, deduplicating, and filtering web crawls (CommonCrawl), Wikipedia, books, code repositories, and the like, and the quality and diversity of the data greatly determine final performance. Since most of the cost (thousands of GPUs over weeks to months) arises at this stage, a Foundation Model ecosystem has formed in which only a few organizations perform pre-training and the rest download and use the results.

The pre-training objective is simple, but its byproduct is astonishing. Merely repeating the single task of "predict the next word" at trillion scale, the model acquires on its own the representations needed for translation, summarization, and commonsense reasoning. This is explained by the insight that "prediction is understanding," in that accurately predicting language requires compressed knowledge about the world. That said, the output of this stage is close to an uncontrolled "raw ore"—rich in knowledge but not steered.

Managing data quality at the pre-training stage, which determines final performance, is also very important in practice. A cleaning pipeline that filters out low-quality, duplicate, and harmful text; the exclusion of data with privacy or copyright issues; and language/domain balancing all happen at this stage. Because the principle "garbage in, garbage out" is amplified even more at trillion-token scale, recent research is moving toward selecting high-quality data rather than blindly increasing volume.

B. Supervised fine-tuning (SFT)

By learning from high-quality instruction–response examples of the form "answer questions, summarize summary requests," the pre-trained model is turned into an assistant that follows instructions. It is a process of re-training via supervised learning on thousands to tens of thousands of carefully written (instruction, ideal response) pairs; although the absolute volume of data is a tiny fraction of pre-training, it decisively changes the model's behavior. As the InstructGPT paper showed, a well-aligned 1.3-billion-parameter model can be preferred by human evaluation over an unaligned 175-billion-parameter model, which eloquently demonstrates the value of the alignment stage.

The point that the diversity and quality of SFT data determine the model's range of instruction following is also important. Only by evenly including a wide range of instruction types—summarization, translation, coding, role-play, and so on—does it become a general-purpose assistant not biased toward a specific task; recently, a trend has also appeared of lowering data-construction cost by mass-generating instruction–response pairs with a powerful model and then having humans review them, instead of writing each by hand.

C. Alignment (RLHF/DPO)

RLHF (Reinforcement Learning from Human Feedback) builds a reward model from data in which humans rank multiple answers to the same question by preference, then optimizes the policy via reinforcement learning (PPO, etc.) to maximize this reward. In contrast, DPO (Direct Preference Optimization) optimizes the policy directly from preference pairs (preferred response, dispreferred response) without a separate reward model or reinforcement-learning loop, and has recently been widely adopted for being simple to implement and stable to train. This stage suppresses harmful and false responses and determines real-use quality such as tone, format, and refusal criteria. In short, pre-training handles "what it knows," SFT "how it follows instructions," and alignment "with what values it answers."

At the alignment stage, one must also consider a trade-off called the Alignment Tax. If safety and helpfulness are aligned too strongly, the model becomes overly cautious (refusing even harmless requests) and some performance can be sacrificed, so carefully tuning the balance among "helpful, honest, harmless (HHH)" in reward-model design and preference-data composition is the crux of practice. One reason DPO is preferred over RLHF is that it reduces the instability of the reinforcement-learning loop and the risk of reward hacking, thereby achieving this balance more stably.

4. The scaling law and emergence

The keys that explain the qualitative leap to LLMs are the scaling law and emergence. Increasing parameters, data, and compute reduces the validation loss predictably in a power-law form (scaling law), and interestingly, once a certain critical scale is crossed, abilities absent in smaller models are manifested discontinuously (emergence). Multi-step reasoning (Chain-of-Thought) and In-context Learning—solving new tasks using only examples in the prompt without changing parameters—are representative.

Concept Content
Scaling Law Increasing parameters/data/compute → loss decreases predictably as a power law
Emergent abilities Reasoning, In-context Learning, etc. manifest above a critical scale
In-context Learning Performs tasks via prompt examples (Few-shot) without weight updates

The lesson the scaling law gave practice is the optimal allocation of the compute budget. Early on it was thought that merely enlarging parameters would suffice, but DeepMind's Chinchilla study (2022) showed that for a given amount of compute, model size and the number of training tokens must be grown in balance, and that large models of the time were excessively large relative to their data. For example, that Chinchilla—70 billion parameters trained on 1.4 trillion tokens—outperformed a 175-billion-parameter model on several benchmarks is a concrete case showing that the answer is not "big for its own sake" but "big in balance with the data."

Emergence too has clear measured cases. Three-digit arithmetic and multi-step reasoning abilities, negligible at GPT-2 scale, suddenly became meaningful with Few-shot at GPT-3 (175 billion parameters) scale; likewise, the phenomenon that a single line of prompting, "Let's think step by step," greatly raises reasoning accuracy is observed only above a certain scale. However, recent research points out that much of this "discontinuous emergence" can appear exaggerated because of the choice of evaluation metric (a correct/incorrect dichotomy), so an attitude of interpreting emergence prudently—not as absolute magic but as one aspect of ability improvement with scale—is needed.

The significance of In-context Learning for practice is especially large. Being able to handle a new task by putting just a few examples (Few-shot) or an instruction into the prompt without changing weights at all means one can instantly change the application without retraining the model. This ability gave rise to the new practical field of prompt engineering and became the foundation of the LLM operating model that reuses a single model for diverse tasks while omitting expensive fine-tuning.

5. In depth: domain application and recent trends

When applying a giant LLM to actual work, avoiding full retraining (whose cost and time are enormous) and combining partial adaptation techniques has become standard. The core is a dual strategy that injects ability/format via PEFT (Parameter-Efficient Fine-Tuning) and up-to-date/specialized knowledge via RAG (Retrieval-Augmented Generation).

LoRA (Low-Rank Adaptation), the representative of PEFT, freezes the original weights and additionally trains only a pair of low-rank matrices at each layer, achieving performance close to full fine-tuning while updating less than 1% of the total parameters. Combining this with 4-bit quantization, QLoRA enables even models with tens of billions of parameters to be fine-tuned on a single high-performance GPU, greatly lowering the barrier for individual companies and labs to build domain-specialized models. This led from a structure in which a few organizations monopolized pre-training to the democratization of the application stage.

RAG, instead of engraving knowledge into parameters, finds an external knowledge base (in-house documents, latest news, etc.) via vector search at query time and inserts it into the prompt. It reduces hallucination, secures recency, and can present sources, so it has established itself as a de facto essential architecture in fields where factual accuracy matters, such as in-house knowledge chatbots and customer counseling. When cost, latency, or data leakage is a concern, the strategy of placing an sLLM (small LLM) lightened via quantization and knowledge distillation on-premises is also increasing. Recently, expansion toward Agents that plan and perform tool calls, search, and code execution on their own is active.

Looking at industry application cases, the effectiveness of this combination is clear. Code auto-completion tools on software development sites connect an LLM pre-trained on vast public code to the in-house codebase via RAG to make suggestions matching the team's own conventions, and in finance and law, they are designed to answer with grounds by searching regulatory documents and precedents, lowering the risk of wrong answers due to hallucination. In customer counseling, combining RAG that searches counseling history and manuals with SFT tuned to the company tone secures consistent response quality. In this way, the three-tier composition of "general-purpose LLM + domain knowledge (RAG) + format adjustment (PEFT/SFT)" is solidifying as the standard recipe across industry.

6. Considerations and implications

  • Role-division design: Dividing responsibility as "ability via fine-tuning (PEFT), knowledge via RAG, safety via alignment/governance" to design reliability, cost, and security in balance is the core from a professional engineer's perspective. Setting the boundary of what to engrave in parameters versus what to retrieve externally governs architecture quality.
  • Risk management (hallucination, bias, misalignment): As scale grows, risks grow along with ability. Hallucination, which plausibly fabricates false answers, must be defended in layers via RAG/source citation/uncertainty expression; bias originating from training data via data cleaning and red-team evaluation; and misalignment that diverges from intent via RLHF/DPO and guardrails.
  • Trade-off of cost, latency, and sovereignty: Between the performance of ultra-large API models and the security/cost/data sovereignty of on-premises sLLMs, a design that tiers (routes) models according to task difficulty and regulatory requirements is realistic.
  • Governance and regulatory response: In line with regulations such as the EU AI Act and internal corporate policies, one must secure data provenance, copyright, personal information, and model audit trails, and also build a defense system against new security threats such as prompt injection and data leakage.
  • Evaluation and verification systems: As the emergence cases show, it is hard to conclude ability from a single benchmark metric, so one must equip a multifaceted verification system that combines per-task quantitative metrics with human evaluation, red-teaming, and real-use log analysis to catch quality degradation or safety issues early after deployment.
  • Outlook: Expansion continues toward multimodal (integrating text, image, and voice), long context, autonomous agents, and on-device sLLMs, so the ability to design a compound AI system that combines multiple models, tools, and search—rather than a single giant model—becomes increasingly important.

References


In one line: A PLM is a transfer-learning model of pre-training + fine-tuning; adding SFT and RLHF/DPO alignment and scaling it up develops it, per the scaling law, into an LLM with emergent abilities such as In-context Learning, and in practice PEFT, RAG, and alignment divide the labor of ability, knowledge, and safety.