← Back to list
AI & Data
#파운데이션모델#사전학습#LLM#Transformer#131회
Last updated · 2026-09-27

Foundation Model

1. Overview

A. Definition

A foundation model is a general-purpose, large-scale AI model that is pre-trained on vast amounts of data and can be adapted to and reused across a wide range of downstream tasks. Representative examples include GPT, BERT, CLIP, and Llama, and they are reused across many applications through fine-tuning or prompting alone. The term was coined in 2021 by researchers at Stanford HAI to mean "a model that serves as the foundation for many tasks."

What foundation models brought about is a paradigm shift in how AI is developed. The essence is the move "from building a new model for each task to adapting one general-purpose model to many tasks." In the past, a translator, a sentiment analyzer, and a chatbot each had to be trained from scratch on their own separate datasets. Each model understood only its narrow task, and whenever a new task arose, large-scale labeling and training had to be repeated all over again.

A foundation model first learns general representations of language and images (broad knowledge of the world and statistical patterns) through large-scale pre-training, and then performs diverse tasks such as translation, summarization, classification, and generation using only a small amount of data or a prompt. Many applications are stacked on top of a single "foundation." This dramatically lowers the barrier to entry for AI development and maximizes reusability, reorganizing the very structure of the AI industry from being "model-development-centric" to "model-utilization-and-adaptation-centric."

B. Background and Necessity

The rise of foundation models is intertwined with three technical developments. First, the discovery of the Scaling Law. OpenAI's research (Kaplan et al., 2020) showed that when data, compute, and model parameters are scaled together, model performance (loss) improves predictably in a power-law form. This provided the investment rationale that "making it bigger makes it better."

Second, the emergence of the Transformer architecture (2017, Attention Is All You Need). Self-Attention, which enables parallel training without recurrent structures, is a good match for large-scale parallel processing on GPUs, making ultra-large-scale training feasible. Third, Self-Supervised Learning. By learning from vast amounts of unlabeled raw text and images through tasks like "predict the next word" or "restore the masked part," it became possible to exploit large-scale data without costly human labeling. As these three factors converged, ultra-large-scale pre-trained models took off in earnest, starting with GPT-3 (175 billion parameters, 2020).

2. The Conceptual Structure of Foundation Models

The operating principle of a foundation model can be summarized as a single pipeline: "pre-training → adaptation → diverse tasks." The overall structure diagram below shows the flow from data to applications.

flowchart LR
  D["Large-scale data<br/>text·image·speech"] --> P["Pre-training<br/>Foundation Model"]
  P --> A["Adaptation<br/>fine-tuning·prompt·RAG"]
  A --> T1["Translation"]
  A --> T2["Summarization"]
  A --> T3["Classification·Generation"]
  A --> T4["Code·Reasoning"]
  style P fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px
  style A fill:#fef3e8,stroke:#ed8a2f,stroke-width:2px

The key point of the figure above is that a single model in the center (P) branches out into numerous tasks on the right. In the past, multiple independent pipelines were needed from left to right, one per task, but a foundation model performs the heavy process of pre-training only once and thereafter repeats only lightweight adaptation. This "train once, use many times" structure is the source of its cost efficiency and speed of diffusion.

The characteristics of foundation models are causally linked to one another. Large-scale parameters and data give rise to generality (using one model for many tasks), and once scale crosses a certain threshold, emergent abilities—capabilities that were not explicitly trained, such as arithmetic reasoning and multi-step thinking—suddenly appear. The general representations thus secured are adapted to specific domains through fine-tuning, prompting, and RAG, and recently they are expanding into multimodal systems that handle text, images, and speech together.

Characteristic Description Practical Implication
Generality One model used for diverse tasks Reduced per-task development cost
Large scale Pre-trained on vast parameters and data Enormous compute and power required
Emergence Unexpected abilities appear as scale grows Nonlinear leaps in performance
Adaptability Domain specialization via fine-tuning·prompt·RAG Specialization possible with little data
Multimodal Integrated processing of text·image·speech Expanded application scope

The difference between traditional AI models and foundation models is especially striking from the perspective of development economics. In the traditional approach, the entire process of data collection, labeling, model design, training, and deployment was repeated for each task, so if there were N tasks the cost roughly grew N-fold as well. In the foundation model approach, once the fixed cost of pre-training is paid a single time, the marginal cost incurred per task plummets to the level of writing a prompt or doing a small-scale fine-tune. This shift in cost structure is the economic driver behind the explosive spread of generative AI, and at the same time it is the backdrop for a monopolization in which capability concentrates in the few firms that can afford pre-training.

3. Underlying Technologies and Training Architecture

A foundation model is an aggregate of several technologies. The detailed diagram below shows the layers of the technology stack, from pre-training to alignment and utilization.

flowchart TB
  subgraph PRE["Pre-training stage"]
    T["Transformer·Self-Attention"]
    SSL["Self-supervised learning"]
    HW["Distributed training·multi-GPU (HBM·InfiniBand)"]
  end
  subgraph ADAPT["Adaptation·alignment stage"]
    FT["Fine-tuning·PEFT (LoRA)"]
    AL["Alignment (RLHF·DPO)"]
    RAG["RAG (retrieval-augmented generation)"]
  end
  PRE --> ADAPT --> APP["Real application services"]
  style PRE fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px
  style ADAPT fill:#fef3e8,stroke:#ed8a2f,stroke-width:2px

First, the Self-Attention of the Transformer is the core structure by which every token in a sentence references every other, learning long-range dependencies of context. Unlike RNNs, which process sequentially, it allows parallel computation and thus resolves the speed bottleneck of large-scale training. Second, self-supervised learning enables pre-training by masking part of the raw data and having the model predict it, without labels. The GPT family uses next-token prediction (autoregressive), while the BERT family uses masked-token restoration (Masked LM).

Third, distributed training and multi-GPU infrastructure (HBM high-bandwidth memory, InfiniBand ultra-high-speed networking, and data/tensor/pipeline parallelism) binds thousands of GPUs into a single training cluster, physically underpinning the training of ultra-large-scale models. Once parameters reach hundreds of billions, the model no longer fits in a single GPU's memory, so a three-dimensional parallelism that combines tensor/pipeline parallelism—splitting the model across multiple devices—with data parallelism—dividing the data for processing—becomes essential.

Fourth, alignment fills the gaps that pre-training alone leaves. A pre-trained model possesses vast knowledge but has not been tuned to respond in the way humans want, so it is first made to follow commands through instruction tuning, and then usefulness, harmlessness, and honesty are reinforced through reinforcement learning from human feedback (RLHF) or the simpler direct preference optimization (DPO). Only with this alignment stage does the model reach the quality usable in actual conversational services.

Fifth, RAG (retrieval-augmented generation) and PEFT (lightweight fine-tuning such as LoRA) efficiently inject up-to-date and domain knowledge. RAG searches an external knowledge base at query time and places the supporting documents into the prompt, thereby reflecting the latest information and reducing hallucination without retraining the model. LoRA keeps the huge original parameters frozen and trains only a small number of low-rank matrices, cutting specialization cost and storage space by tens of times.

Technology Role Representative Example
Transformer·Self-Attention Parallel learning of long-range context GPT, BERT
Self-supervised learning Large-scale pre-training without labels Next-token prediction, MLM
Distributed training·multi-GPU Training ultra-large-scale models Data·tensor parallelism
Alignment (RLHF·DPO) Fine-tuning to reflect human preference ChatGPT, Claude
RAG·PEFT (LoRA) Efficient injection of latest·domain knowledge Enterprise knowledge chatbots

4. Use Cases and Implementation Considerations

Foundation models are spreading across all industries. Concretely, Microsoft's GitHub Copilot leverages a code foundation model, and research has reported that it meaningfully increased developer productivity; in finance and healthcare, domain-specific fine-tuning (e.g., BloombergGPT) is used to analyze specialized documents. The CLIP and DALL·E families, which jointly learn images and text, are used for search, generation, and content creation. As shown here, the current trend is for industry-specific applications to be stacked layer upon layer on top of a single foundation.

The ways enterprises adopt foundation models generally fall into three branches. The first is calling a commercial API, which has almost no upfront cost and can be used immediately, but entails a dependency in which data leaves the organization and billing scales with usage. The second is putting an open-weight model (e.g., the Llama family) on one's own infrastructure and fine-tuning it, which is advantageous for data sovereignty and regulatory compliance but demands GPU infrastructure and MLOps capability. The third is placing a small specialized model (sLLM) on-premises so that sensitive data is processed only internally. A professional engineer must synthesize data sensitivity, regulation, cost, and performance requirements to judge the trade-offs among these three approaches, and, when necessary, be able to design a hybrid configuration that combines internal knowledge via RAG.

Foundation models can be broadly divided into three lineages, and the choice differs by application purpose. The encoder family, such as BERT, is strong at understanding an entire sentence bidirectionally, making it well suited to comprehension tasks like classification, search, and named-entity recognition. The decoder family, such as GPT, is specialized for sequentially generating the next token and is strong at generation, conversation, and summarization. The encoder-decoder family (e.g., T5) is well balanced for translation- and summarization-style tasks that understand an input and generate a new output. In practice, as conversational services have proliferated, the decoder family has become mainstream, but the encoder family is still widely used for search and embedding purposes.

However, with such power come responsibilities on three levels. Legally, the copyright and licensing of training data (real lawsuits are underway, such as The New York Times' suit against OpenAI), personal information, accountability for outputs, and compliance with regulations such as the EU AI Act and Korea's AI Framework Act become issues. Environmentally, training ultra-large-scale models consumes enormous power and carbon, so Green AI and model compression are demanded. Socially, discrimination from amplified bias in training data, misinformation and deepfakes, changes to jobs, misuse, and the model monopoly of a few big-tech firms and fair-access issues are raised.

Aspect Consideration Direction of Response
Legal Training-data copyright·licensing, personal information, accountability, regulation Data provenance management, regulatory compliance framework
Environmental Power and carbon emissions of training Green AI·compression·efficient training
Social Bias·discrimination, misinformation·deepfakes, jobs·monopoly Alignment·guardrails, open models

5. In-Depth: Latest Trends and Technological Evolution

The foundation model ecosystem is evolving rapidly. First, the trend toward efficiency and lightweighting (sLLM). Beyond giant models with hundreds of billions of parameters, small models with a few billion parameters (sLLM) are increasingly optimized via quantization, distillation, and MoE (Mixture of Experts) for on-device and cost-reducing applications. MoE activates only a subset of expert sub-networks for each input, reducing computation while maintaining large-scale capacity, and has established itself as a core technology in the latest models.

Second, the expansion into agents. As foundation models move beyond simple responses to serve as the brain of an agent that uses tools (search, code execution, API calls) on its own and plans and carries out multi-step tasks, the autonomy of applications is greatly increasing. Third, multimodal integration is accelerating, moving toward processing text, images, speech, and video in a single model.

Fourth, the refinement of evaluation and governance. As models grow larger and their use broadens, evaluation is moving beyond simple accuracy metrics to standardize benchmarks (e.g., MMLU) and assessments of safety, bias, and harmfulness, and procedures for pre-checking potential for abuse through red-teaming are being incorporated into the development pipeline. Transparency practices that document training data and limitations via model cards and datasheets are also spreading.

From the perspective of past questions and exam trends, the Professional Engineer Information Management exam has posed foundation models in a form that links them with LLMs, generative AI, Transformers, and RAG, asking about concepts, underlying technologies, and risks. An answer can secure both depth and currency if it is organized along the flow "meaning of the paradigm shift → scaling law and underlying technologies → adaptation and alignment mechanisms → legal, environmental, and social risks → latest trends such as lightweighting and agents."

6. Considerations and Implications

  1. The spread of the application ecosystem is the core trend. Countless applications are built on top of foundation models through fine-tuning, prompting, and agents, and corporate competitiveness is shifting from "the ability to build models" to "the ability to adapt and combine models well." A professional engineer must be able to advise on the strategic choice among in-house development, API use, and fine-tuning of open-source models.

  2. Reliability, safety, and cost are the keys to practical use. Reliability that reduces hallucination and bias, safety through alignment and guardrails, and efficiency that lowers inference cost via sLLM, quantization, and MoE determine the success or failure of actual service adoption. In particular, RAG has established itself as a practical standard for reducing hallucination through currency and evidence presentation.

  3. Embedding legal, environmental, and social risks at the design stage (Responsible AI by Design) is necessary. The more powerful a general-purpose technology, the greater its side effects, so data governance, copyright management, bias mitigation, and carbon measurement must be integrated from the earliest stages of development. After-the-fact responses incur high costs and loss of trust.

  4. Sovereignty and monopoly risks and the open ecosystem. The model monopoly of a few big-tech firms produces dependency and unequal access, so securing open models (open weights) and national-level AI infrastructure (sovereign AI) is emerging as a strategic task. It is also important from the standpoint of data sovereignty and security.

  5. Securing regulatory alignment. A regulatory environment such as the EU AI Act's risk-based regulation and Korea's AI Framework Act is forming rapidly, so the higher the risk of a field of use, the more transparency, explainability, and record-keeping obligations must be reflected in the design in advance.

References


In one line: A foundation model is a general-purpose AI pre-trained at large scale on the basis of the scaling law, the Transformer, and self-supervised learning; it is adapted to diverse tasks through fine-tuning, prompting, alignment (RLHF), and RAG, and has recently evolved toward sLLM, MoE, agents, and multimodality—yet legal, environmental, and social risks such as training-data copyright, power, and bias must be managed in an integrated way from the design stage.