← Back to list
Hardware & Semiconductor
#FPGA#재구성형컴퓨팅#ASIC#하드웨어가속#HLS
Last updated · 2026-10-04

FPGA (Field Programmable Gate Array)

1. Overview

A. Definition

An FPGA (Field Programmable Gate Array) is a semiconductor whose internal logic-circuit connections can be freely reconfigured by the user in the field even after manufacturing; it arranges logic blocks (LUTs and flip-flops) and the programmable interconnect that links them in an array, and implements an arbitrary digital circuit in hardware by injecting the connection information as a bitstream — a reconfigurable-computing device.

The fundamental reason FPGAs emerged lies in the gap between the rigidity of an ASIC (Application Specific Integrated Circuit) and the inefficiency of a CPU. An ASIC freezes a specific function into a dedicated circuit, giving the best performance and power efficiency, but once fabricated the circuit cannot be changed, and mask production incurs non-recurring engineering (NRE) costs of billions to tens of billions of won and a design-and-verification cycle of several months to more than a year. Conversely, general-purpose processors such as CPUs/GPUs can execute anything in software, but because of the von Neumann overhead of instruction fetch, decode, and execute, their performance per watt drops when a specific operation is repeated. The FPGA occupies the unique position between the two of "having hardware's parallelism and deterministic low latency while still being able to change its design like software."

This value grows as the industrial environment changes. In domains where specifications change frequently, such as communication standards (5G·Wi-Fi), video codecs, and AI models, an ASIC that freezes hardware before the standard is finalized is risky. An FPGA only needs a new bitstream when a standard is revised or an algorithm changes, dramatically shortening Time To Market (TTM). Moreover, in small-volume, high-mix production or in the prototyping stage before ASIC mass production, the FPGA is economically overwhelming. Recently, in domains where latency is value itself — data-center AI-inference acceleration, network acceleration (SmartNIC), and high-frequency trading (HFT) — the FPGA's sub-microsecond deterministic response is drawing renewed attention.

B. Background and Necessity

Since Xilinx released the first commercial FPGA (XC2064) in 1985, FPGAs have expanded from a simple glue-logic replacement to today's data-center accelerators. Three pressures lie behind this expansion. First, with the slowing of Moore's Law and the power wall, clock and core scaling of general-purpose CPUs hit a ceiling, and domain-specific architectures (DSA) that accelerate specific workloads with dedicated hardware rose to prominence. Second, the rapid evolution of AI and communication standards increased demand to accelerate without freezing hardware. Third, in domains such as edge, defense, and aerospace that require small-volume production, field upgrades, and radiation-hard (rad-hard) characteristics, reconfigurability became essential.

C. Characteristics

The nature of FPGAs can be summarized by the following characteristics.

  • Reconfigurability: The circuit can be changed simply by swapping the bitstream, allowing flexible response to bug fixes, feature additions, and standard changes.
  • Spatial Parallelism: Thousands to millions of logic cells operate simultaneously, unrolling iterative and pipelined computation across hardware.
  • Deterministic Low Latency: With no OS, cache, or branch prediction, response time is predictable at the cycle level.
  • A middle point in power efficiency: For the same operation, performance per watt is higher than a CPU/GPU, but because of reconfiguration overhead it falls short of an ASIC.
  • High design difficulty and resource constraints: Knowledge of RTL and timing closure is required, and the design must fit within the available logic and memory resources.
  • Field-upgrade capability: Even after deployment the bitstream can be updated remotely, enabling performance improvements, security patches, and support for new standards.

2. Internal Structure and Components of an FPGA

An FPGA is clearest when understood as three layers: 'programmable logic + programmable interconnect + embedded hard blocks'. The key point is that it builds arbitrary combinational logic with small memories that hold truth tables (LUTs), stores state with flip-flops, and freely connects these with array-shaped wiring. The diagram below shows the overall composition of the FPGA fabric.

flowchart TB
  subgraph FPGA["FPGA fabric"]
    direction TB
    IOB["I/O block(IOB)<br/>LVDS·high-speed transceiver"]
    subgraph FABRIC["programmable logic array"]
      CLB1["CLB(LUT+FF)"]
      CLB2["CLB(LUT+FF)"]
      SW["switch matrix<br/>(programmable wiring)"]
      BRAM["block RAM(BRAM)"]
      DSP["DSP slice<br/>(multiply-accumulate MAC)"]
    end
    CLK["clock tree·PLL/MMCM"]
    HARD["hard IP<br/>(PCIe·DDR·Ethernet MAC)"]
  end
  IOB --> FABRIC
  CLB1 <--> SW
  CLB2 <--> SW
  SW <--> BRAM
  SW <--> DSP
  CLK --> FABRIC
  HARD <--> FABRIC
  style FABRIC fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px
  style DSP fill:#eafbea,stroke:#2f9e44,stroke-width:1px
  style BRAM fill:#eafbea,stroke:#2f9e44,stroke-width:1px

A. Configurable Logic Block (CLB). This is the heart of the FPGA. At its core is the Look-Up Table (LUT): a typical 6-input LUT stores a truth table in 64-bit SRAM and computes a 6-variable combinational function in one cycle. That is, rather than building logic gates directly, it implements an arbitrary Boolean function by "writing output values for input combinations into memory in advance and reading them." Alongside the LUT, a CLB contains flip-flops (FF) to create sequential logic (registers) and uses a dedicated carry chain to perform arithmetic such as addition quickly. Hundreds of thousands to millions of CLBs come together to form a huge parallel circuit.

B. Programmable Interconnect and Switch Matrix. No matter how many CLBs there are, they are meaningless unless freely connected. A substantial portion of the FPGA area is in fact occupied by wiring and the switch matrix, and the bitstream determines the on/off of these switches to form paths between logic blocks. This is why design performance (maximum operating frequency) depends heavily on place-and-route quality — if a signal has to detour to a distant CLB, wiring delay grows and timing cannot be met.

C. Embedded Hard Blocks (BRAM·DSP·hard IP). Building everything out of LUTs is inefficient, so frequently used functions are embedded as dedicated silicon (hard blocks). Block RAM (BRAM) is on-chip memory in units of tens of Kb used for buffers, FIFOs, and caches, while a DSP slice is a hardwired multiplier-accumulator (MAC) that is central to signal processing and AI computation. Advanced FPGAs embed PCIe, DDR memory controllers, 100G Ethernet MACs, and high-speed SerDes transceivers as hard IP, avoiding the resource and performance loss of implementing these in LUTs.

D. Clock·I/O·SoC Integration. A dedicated clock tree and PLL/MMCM supply low-jitter clocks across the entire chip and manage multiple clock domains. The I/O blocks support various electrical standards (LVDS·LVCMOS) and transceivers of tens of Gbps. Furthermore, an SoC FPGA (e.g., AMD Zynq, Intel Agilex SoC) integrates a hard processor such as an ARM core (PS, Processing System) and the FPGA fabric (PL, Programmable Logic) on a single chip, combining software flexibility with hardware acceleration.

Component Role Analogy
LUT Implements arbitrary combinational logic (truth-table memory) The basic Lego block of logic
flip-flop (FF) Stores state (sequential logic) A 1-bit memory element
interconnect/switch Connects wiring between blocks The road network of the circuit
BRAM On-chip memory buffer A working scratchpad
DSP slice Accelerates multiply-accumulate (MAC) A dedicated calculator
hard IP PCIe·DDR·Ethernet, etc. Off-the-shelf parts

3. FPGA Design Flow

FPGA development resembles software compilation, but it differs decisively in that the final output is not an 'executable' but a 'circuit connection table (bitstream)'. The designer describes, in a hardware description language, not 'what to compute' but 'what circuit to build', and the tools convert this into actual logic-cell placement and wiring. Below is a typical design flow.

flowchart LR
  A["design entry<br/>(Verilog/VHDL RTL · HLS C/C++)"] --> B["logic synthesis<br/>(Synthesis)"]
  B --> C["place and route<br/>(Place and Route)"]
  C --> D["static timing analysis<br/>(STA)"]
  D -->|timing violation| B
  D -->|pass| E["bitstream generation<br/>(Bitstream)"]
  E --> F["device configuration<br/>(Configuration/Download)"]
  A -.verify.-> V["functional simulation<br/>(Testbench)"]
  C -.verify.-> V
  style B fill:#e8f0fe,stroke:#2f6fed,stroke-width:1px
  style C fill:#e8f0fe,stroke:#2f6fed,stroke-width:1px
  style E fill:#eafbea,stroke:#2f9e44,stroke-width:2px

A. Design Entry and Verification. Traditionally, the circuit's behavior is described at the register and combinational-logic level in an RTL (Register Transfer Level) language such as Verilog·VHDL. At each stage, testbench-based functional simulation verifies that the design behaves as intended. Recently, HLS (High-Level Synthesis) that automatically converts C/C++/OpenCL into RTL to raise productivity has been spreading.

B. Logic Synthesis. This stage maps RTL code to the FPGA's actual resources (LUT·FF·DSP·BRAM) and corresponds to software compilation. The synthesis tool performs Boolean-algebra optimization and technology mapping to produce a netlist. Because resource usage and estimated delay are determined here, coding style directly affects the Quality of Results (QoR).

C. Place and Route, and Timing Closure. Each logic element of the netlist is placed at a physical location on the chip and connected (routed) through the switch matrix. Then static timing analysis (STA) checks whether every signal path arrives within the target clock period. If there is a violation (negative slack), the timing-closure iteration — adjusting constraints or revising the design and re-synthesizing — is the most demanding process in FPGA development. For example, if a design targeting 300 MHz yields only 280 MHz due to wiring delay, pipeline stages must be added or logic relocated.

D. Bitstream Generation and Configuration. The final place-and-route result is encoded into a bitstream, and downloading it to the FPGA sets the switches and LUT contents so the circuit 'comes alive'. Because an SRAM-based FPGA loses its configuration when power is off, it loads the bitstream from external flash at every boot.

4. FPGA Type Comparison and Application Cases

A. Types by Configuration Memory. FPGAs are divided by how they store the circuit configuration. SRAM-based types reconfigure quickly and are favorable for process scaling, so they are mainstream (AMD·Intel), but being volatile they require external boot memory and risk configuration exposure at power-on. Flash-based types (Microchip/Microsemi) are non-volatile, so they boot instantly, consume little power, and are favorable for security. Antifuse types can be programmed only once but have excellent radiation-hardness and are used in satellites and defense. Thus, even for the 'same FPGA', the choice differs by application field.

B. Comparison with CPU·GPU·ASIC. The difference among the four devices is explained as a 'trade-off between flexibility and efficiency'. The CPU is the most flexible but has low throughput and power efficiency; the GPU is strong at large-scale data parallelism (SIMT) but has high latency and large power consumption. The ASIC has the best efficiency but zero flexibility and enormous NRE. The FPGA, in between, offers low latency and high power efficiency together with flexibility.

Category CPU GPU FPGA ASIC
flexibility best (SW) high (SW) high (reconfigurable) none (fixed)
latency medium high very low (deterministic) very low
power efficiency low medium high best
NRE none none low very high
TTM immediate immediate weeks–months months–1 year+
suitable domain general control AI training·mass parallelism low-latency acceleration·small-volume high-mix high-volume mass production

This trade-off is also explained by break-even volume. An ASIC has large NRE but a very low unit cost, so once production volume exceeds a certain scale (typically tens of thousands to hundreds of thousands of units) its total cost becomes cheaper than an FPGA. Conversely, when volume is small or the likelihood of design changes is high, the FPGA's total cost of ownership (TCO) — with essentially no NRE and low redesign cost — is favorable. Therefore, dividing FPGA and ASIC along the axes of 'expected volume × lifetime × change frequency' is the starting point of practical decision-making, and many companies adopt a staged strategy of launching and validating first with an FPGA, then switching to an ASIC once volume is confirmed.

C. Industrial Application Cases (specific, with figures). First, in data-center AI inference, Microsoft deployed FPGAs at scale into Bing search and Azure through its 'Catapult/Brainwave' project, providing real-time AI inference with low latency even at batch size 1. Second, in financial high-frequency trading (HFT), network packet processing and order logic are implemented on FPGAs to achieve 'tick-to-trade' latency at the level of hundreds of nanoseconds to 1 microsecond, securing a trading edge with responses tens of times faster than a CPU. Third, in 5G base stations and SmartNICs, baseband signal processing and network offload (encryption, virtual switching) — whose standards keep evolving — are accelerated with FPGAs, so a standard revision is handled by a firmware update. Fourth, cloud FPGAs (AWS EC2 F1/F2 instances) let users rent FPGA accelerators by the hour for genomic analysis, video transcoding, and financial risk computation.

5. Deep Dive: Latest Trends and Evolution into Adaptive Computing

The FPGA ecosystem is evolving rapidly along two axes: 'the democratization of hardware design' and 'heterogeneous integration'.

A. High-Level Synthesis (HLS) and Development Productivity. The biggest barrier to FPGA adoption was the difficulty of RTL design. To lower it, HLS converts C/C++/OpenCL into RTL, and Python-based overlays (e.g., PYNQ) that let software developers use accelerators have spread. This is a change that broadens the FPGA from 'the domain of HDL experts' into 'an acceleration tool for general developers'.

B. Partial Reconfiguration. This is a technique that replaces only a part of the chip while it operates without stopping the whole chip, so different accelerators can be swapped over time in a single FPGA. It is used to dynamically switch per-tenant acceleration functions in data centers or to switch communication modes in wireless base stations. This gives the effect of reusing limited logic resources by 'time-division', letting even small devices accommodate more functions. However, it comes at the cost of increased design complexity, since interface design between replaceable regions and ensuring safety during configuration are difficult.

C. Adaptive Computing (ACAP) and AI Engines. Next-generation products such as AMD's Versal and Intel's Agilex have evolved into 'Adaptive Compute Acceleration Platforms (ACAP)' that integrate, into the traditional FPGA fabric, vector-type AI engines (arrays of VLIW/SIMD cores), a NoC (Network-on-Chip), and a hard processor. The matrix operations of AI inference are handled by the dedicated AI engines, user-defined pre/post-processing by the FPGA fabric, and control by the ARM core, processing heterogeneous workloads on a single chip.

D. Open-Source Toolchains and Ecosystem Opening. For a long time, FPGA development was tied to vendor-locked closed tools, but recently open-source synthesis and place-and-route tools such as Yosys and nextpnr, together with research on open bitstream formats (e.g., Project IceStorm), have advanced, forming an open development flow centered on small FPGAs. This lowers the entry barrier for academia and startups and promotes reproducible hardware design, serving as a long-term driver that broadens the base of FPGA use.

E. Chiplet·eFPGA·CXL Linkage. With the chiplet approach, FPGA dies and HBM·I/O dies are integrated in 2.5D/3D to raise bandwidth, and eFPGA (embedded FPGA), which embeds a small FPGA IP inside an SoC, grants partial reconfigurability to an ASIC. In addition, FPGA accelerators that connect cache-coherently with the CPU and memory over the CXL interconnect are being researched and productized, emerging as a core of memory-sharing acceleration. (See [[cxl-compute-express-link]], [[chiplet]])

6. Considerations and Implications

FPGA adoption must be judged comprehensively — 'when, what, and how' — from a professional-engineer perspective.

  • Adoption strategy (when is it an FPGA?): Judge by volume, lifetime, and change frequency. If high-volume mass production is confirmed and specifications are stable, an ASIC is favorable; if production is small-volume high-mix or standards/algorithms are fluid and prototyping is needed before ASIC mass production, an FPGA is suitable. In data centers, deploy it selectively to workloads where 'latency is value'.
  • Managing trade-offs: An FPGA has better power efficiency and latency than a CPU/GPU, but its design difficulty is high, its unit cost is higher than an ASIC, and due to reconfiguration overhead it falls short of an ASIC's peak performance. An organization must weigh development capability (RTL and timing-closure staff), TTM, and total cost of ownership (TCO) together.
  • Security (bitstream·supply chain): Because an SRAM-based device loads the bitstream from external memory at boot, bitstream encryption and authentication (AES·HMAC) and anti-tamper are essential. Reconfigurability also creates new supply-chain threats such as 'hardware Trojans' and malicious bitstream injection, so integrity verification of the design IP and the configuration process is important.
  • Organizational capability and TCO: The success of an FPGA ultimately depends on people. Securing staff with RTL design, timing-closure, and verification-automation capability, and estimating the total cost of ownership — including vendor-tool licenses and development time — must be done from the outset. Although HLS and high-level frameworks are lowering the entry barrier, the strategy must reflect that extracting high performance still requires hardware thinking.
  • Outlook and related technologies: The FPGA is being reorganized beyond a standalone device into the center of heterogeneous integration, adaptive computing, and memory-sharing acceleration. Amid a division of roles with GPU ([[multi-gpu]]), NPU ([[npu]]), ASIC, chiplet ([[chiplet]]), and CXL ([[cxl-compute-express-link]]), the FPGA will retain its unique position of 'acceleration that needs flexibility' and is expected to remain a core axis of AI, communications, and the edge.

References


In one line: An FPGA is a reconfigurable device that implements an arbitrary digital circuit in hardware by reconfiguring LUTs, flip-flops, and programmable wiring via a bitstream; positioned between the efficiency of an ASIC and the flexibility of a CPU, it is evolving into a core axis of AI, communications, and edge acceleration with low-latency parallel acceleration and fast TTM as its weapons.