NAND Flash-Based SSD (FTL, Wear Leveling, Garbage Collection)
1. Overview
A. Definition
An SSD (Solid State Drive) is a semiconductor storage device that uses NAND flash memory instead of a rotating magnetic disk and behaves like a block device through the FTL (Flash Translation Layer), a software layer that mediates between the host's logical addresses and the flash's physical addresses. The FTL hides the physical constraints of flash cells (no overwrite, limited lifetime) so that the operating system sees the same interface as a conventional disk.
The starting point for understanding an SSD is the physical fact that "NAND flash cannot be overwritten in place like a disk." A hard disk can overwrite a new value directly onto a given sector, but NAND flash performs writes (program) at page granularity while erases are only possible at the much larger block granularity, and a page that already holds a value cannot be rewritten until the entire block is erased. This asymmetry (page write vs. block erase) is the root of almost every design challenge of the SSD — garbage collection, wear leveling, write amplification. The performance and lifetime of an SSD ultimately depend on how cleverly the FTL hides this constraint.
B. Background and Necessity
Since the late 2000s, CPU and memory performance kept rising, but storage remained bound by the magnetic disk's mechanical seek latency (seek time, several ms), deepening the I/O bottleneck. Datacenter database, virtualization, and analytics workloads demand large volumes of random I/O, yet a rotating disk is especially weak at random access because the head must physically move. NAND flash has no mechanical parts, providing microsecond-level access latency and high IOPS through parallel channels, and thus became the means to resolve this bottleneck. However, because flash cells have a finite lifetime as their insulating film degrades with repeated writes and erases, flash became a practical storage medium only once the controller and firmware technologies that manage lifetime and hide the constraints matured together. Today the SSD has become the de facto standard medium from laptops to all-flash arrays and AI training storage. In addition, the SSD is free from the structural limits of mechanical disks — noise, vibration, shock vulnerability, high idle power — so its range of application is wide, from mobile and edge devices to dense server racks. Yet the premise of all these advantages is "firmware that hides the physical constraints of flash," which makes the SSD not a mere component but a product of hardware-software co-design; thus, even with the same NAND, performance, lifetime, and reliability diverge greatly by controller and FTL quality.
2. SSD Structure and Core Components
An SSD is broadly composed of four parts: host interface — controller (FTL, ECC, channel management) — DRAM cache — NAND flash array. Each component buffers the physical constraints of flash while raising performance and reliability, sharing those roles.
flowchart TB
Host["Host (OS/filesystem)"] -->|"NVMe/SATA command"| IF["Host interface"]
IF --> CTRL["SSD controller"]
subgraph CTRL["SSD controller"]
FTL["FTL (address mapping)"]
GC["Garbage collection engine"]
WL["Wear leveling"]
ECC["ECC/LDPC engine"]
end
CTRL <--> DRAM["DRAM cache (mapping table)"]
CTRL -->|"multi-channel parallel"| NAND["NAND flash array"]
subgraph NAND["NAND flash array"]
CH0["Channel 0: die·block·page"]
CH1["Channel 1: die·block·page"]
CHn["Channel N: die·block·page"]
end
style CTRL fill:#e8f0fe,stroke:#2f6fed,stroke-width:2px
style NAND fill:#fef7e8,stroke:#e0a500,stroke-width:1px
The host interface is the channel linking the SSD to the system. Early SSDs reused the hard-disk SATA (AHCI) as-is, but AHCI, designed with a single queue and low queue depth, could not exploit flash parallelism. To resolve this, NVMe (NVM Express) appeared over PCIe, supporting many queues (up to 64K queues × 64K commands) so that even multicore CPUs pouring in commands simultaneously hit no bottleneck. In fact, a SATA SSD saturates at about 550MB/s, whereas a PCIe 4.0 NVMe SSD delivers over 7,000MB/s sequential read. This is a representative example of how interface choice governs perceived performance.
The controller is the brain of the SSD, housing the FTL, garbage collection, wear leveling, and ECC engines inside. The controller takes the host's logical block address (LBA), looks up the FTL mapping table to translate it into a physical address, and distributes writes across multiple channels and dies (interleaving) to draw out parallelism. The more channels, the more flash banks accessible at once, increasing bandwidth.
The DRAM cache is mainly the space that holds the FTL mapping table resident. The table must be in DRAM for fast address translation, and its size is roughly 1/1000 of capacity (about 1GB per 1TB). Low-cost DRAM-less SSDs replace part of this table with the HMB (Host Memory Buffer) that places it in host memory, but with the trade-off of lower random performance.
The NAND flash array is structured as the hierarchy die → plane → block → page. Writes and reads are at page (e.g., 16KB) granularity, while erases are at block (e.g., hundreds to thousands of pages) granularity — and this asymmetry originates right here. The table below organizes this hierarchy and the operations possible at each unit; why the erase unit is larger than the write unit is the starting point of all design.
| Hierarchy unit | Approx. size | Possible operation | Implication |
|---|---|---|---|
| Cell | 1~5 bit | charge storage | more levels → more errors·wear↑ |
| Page | 4~16KB | read·write (program) | minimum write unit |
| Block | hundreds~thousands of pages | erase | minimum erase unit → root of no-overwrite |
| Plane/die | many blocks | parallel access | interleaving → bandwidth↑ |
The performance by which SSDs overwhelm rotating disks comes from the parallelism of this hierarchy. The controller splits one large write finely across multiple channels, dies, and planes and performs interleaving that records them simultaneously, so despite the slow write of a single flash die (hundreds of μs), total bandwidth reaches several GB/s. In other words, SSD performance depends not on individual cell speed but on how many banks are struck at once, and that is why performance differs greatly between a fresh state with ample free blocks allowing parallel writes and an aged, highly occupied state where GC is busy.
3. FTL Operation and Garbage Collection·Wear Leveling
The FTL is the heart of the SSD, the core firmware that hides the flash's constraints from the host. The diagram below shows how the FTL circumvents the "no in-place overwrite" constraint through log-structured recording and garbage collection.
flowchart LR
W["Rewrite request for LBA 10"] --> M["FTL: write to a new physical page"]
M --> O["mark old physical page as stale"]
O --> MAP["update mapping table (LBA10 to new PPA)"]
MAP --> CHK{"free blocks scarce?"}
CHK -->|"no"| DONE["done"]
CHK -->|"yes"| GC["start garbage collection"]
GC --> V["copy only valid pages to a new block"]
V --> E["erase the original block"]
E --> FREE["secure a free block"]
FREE --> DONE
style GC fill:#fdecec,stroke:#d33,stroke-width:2px
A. Address Mapping
The most basic function of the FTL is the mapping that links the host's logical address (LBA) to the flash's physical address (PPA). By mapping granularity it divides into page mapping (fine-grained but a large table), block mapping (small table but low flexibility), and hybrid mapping that compromises between the two. Most high-performance SSDs operate close to page mapping for performance, and that is why they need DRAM amounting to the "1/1000 of capacity" mentioned earlier.
The core design principle is that "it does not do update-in-place." When a request comes to change the value of LBA 10, the FTL, instead of overwriting the existing page, writes anew to a free page and updates only the mapping table to the new physical address. This way, writes always occur sequentially (log-structured), fitting the flash characteristics, but with the side effect that the old pages accumulate as invalid data (garbage). Reclaiming these invalid pages is garbage collection.
Also, because the mapping table must not be lost even on a sudden power cut, the FTL periodically stores the mapping-change history to flash and reconstructs it on power recovery. Enterprise SSDs mount a capacitor (PLP, Power-Loss Protection) so that even if power is cut, unflushed data in DRAM is pushed into flash — this too is to preserve the consistency of this mapping and data. In other words, the FTL is less a simple address translator than a mini storage engine responsible for atomicity and durability.
B. Garbage Collection (GC)
A block mixed with invalid pages cannot be reused as-is. Because erase is per block, only after the valid pages remaining in that block are first copied to another free block and the original block is erased in its entirety does it become a free block. This process is GC. GC induces internal writes (copies) not requested by the host, so the amount of data actually written to flash exceeds what the host requested. This ratio is the WAF (Write Amplification Factor); if WAF is 2, even when the host writes 1GB, 2GB is written to flash, cutting both lifetime and performance.
When and which blocks GC targets is the key to performance. Background GC that runs ahead during idle time reduces user-perceived latency, but when free blocks are exhausted during a write surge, foreground GC blocks user I/O and induces latency spikes. For GC target-block selection, a policy of reclaiming blocks with a high ratio of invalid pages first (greedy), or a cost-benefit-based policy that also considers wear balance, is used. The more invalid pages, the fewer valid pages to copy, so reclamation efficiency is better.
To mitigate this, over-provisioning (OP) — spare space invisible to the user — is reserved. Securing an OP of typically 7~28% lowers the ratio of valid pages GC must copy, reducing WAF. As a simple intuition, when the device is nearly full (spare space scarce), each block GC erases is packed with valid pages, so copy volume explodes and WAF soars, whereas with ample OP it can on average pick 'emptier' blocks to erase, so copy volume drops. This is why enterprise SSDs set a larger OP than consumer ones (e.g., consumer 7% vs. enterprise 2028%). In practice, users sometimes deliberately under-allocate partitions (akin to short-stroking) to increase OP and improve write endurance and latency.
C. Wear Leveling
NAND cells have a limit on the number of P/E (Program/Erase) cycles (see table below). If writes concentrate on only certain blocks, those blocks wear out first and shorten overall lifetime. Wear leveling distributes write and erase counts evenly across all blocks to prevent the early death of particular blocks. It places frequently changing data (hot) and rarely changing data (cold) separately, and deliberately moves the data of long-unchanged blocks (static wear leveling) so those blocks also get a chance to be erased. Thanks to wear leveling, even cells rated for hundreds to thousands of cycles can be used stably for years.
D. ECC·Reliability Correction
As miniaturization and multi-level storage progress, the charge stored in flash cells shakes due to leakage and interference, raising the read error rate. In particular, when cells are left for long the charge gradually leaks out — retention degradation; read disturb that arises from repeatedly reading neighboring cells; and insulating-film degradation from accumulated P/E — these overlap. What corrects them is ECC (Error Correction Code); beyond the old BCH, today's high-density NAND uses LDPC (Low-Density Parity-Check). LDPC can recover more errors with soft-decision decoding, becoming the precondition for commercializing error-prone cells like TLC and QLC. In other words, it is a structure that compensates for the physical reliability decline of cells with signal processing, and the controller even performs data refresh (scrub), preemptively moving data to another block when the error rate exceeds a threshold, to maintain integrity.
4. Cell Type·Interface Comparison and Application Cases
Depending on how many bits a single cell holds, capacity, unit cost, lifetime, and performance change systematically. The more bits put into one cell (the more finely the voltage range is divided), the more consistently the trade-off works that the per-capacity unit cost drops but the voltage margin narrows, making it vulnerable to errors and wear.
| Cell type | Bits/cell | Approx. P/E lifetime | Characteristics·use |
|---|---|---|---|
| SLC | 1 | tens of thousands~100K | top endurance·speed, expensive → cache·industrial |
| MLC | 2 | thousands~10K | balanced → past enterprise |
| TLC | 3 | hundreds~3K | mainstream consumer·datacenter |
| QLC | 4 | hundreds~1K | high capacity·read-centric (archive) |
| PLC | 5 | experimental·early | aiming ultra-low cost, reliability challenge |
The figures above must be seen as approximate trends since they vary greatly by process and generation. Real products operate even QLC drives with some regions used like SLC — a pSLC cache — to boost write-burst performance. For example, a consumer QLC SSD quickly absorbs writes via the SLC cache when there is plenty of free space, but once the cache fills it writes directly to the QLC main area and speed drops sharply. This is a representative cause of benchmark figures differing from real-use feel.
Generational differences in the interface itself are also reflected directly in perceived performance. The table below compares the approximate characteristics of the major interfaces, showing that to draw out the flash's internal parallelism, the queue model and physical bandwidth must support it together.
| Interface | Protocol | Approx. sequential bandwidth | Queue model | Note |
|---|---|---|---|---|
| SATA 3 | AHCI | ~550 MB/s | single queue·32 | legacy compatible |
| PCIe 3.0 ×4 | NVMe | ~3,500 MB/s | multi-queue | entry NVMe |
| PCIe 4.0 ×4 | NVMe | ~7,000 MB/s | multi-queue | mainstream high-performance |
| PCIe 5.0 ×4 | NVMe | ~14,000 MB/s | multi-queue | latest·high heat |
As an application case on the interface side, a database OLTP server values random writes and low latency, so it uses an enterprise TLC NVMe with a large OP, while a cold-data archive is read-centric, so high-capacity QLC is economical. On the other hand, uses with extreme writes such as logs and journals still suit high-endurance SLC/pSLC. For example, a large CDN or video service's cache node has a clear "write once, read many" pattern, so it lowers cost with high-capacity QLC, whereas a financial transaction log store requires extreme write endurance and so selects a product with high DWPD. Even the same "SSD" — workload-matched choices of cell, OP, and interface govern the total cost of ownership (TCO).
5. Deep Dive — Latest Trends and Host-Cooperative Architecture
The recent trend in the NAND industry is to break through the limit of horizontal miniaturization with vertical stacking. As planar (2D) miniaturization reached its limit due to cell-to-cell interference, 3D NAND that stacks cells vertically became mainstream, and the stacking count is rapidly growing from dozens of layers to over 200 layers. Stacking increases capacity per area but creates new process challenges of etching deep holes (channel holes) and securing uniformity, making a technology gap by manufacturer. Recently, competition is fierce to raise the layer count further with bonding·string-stacking techniques that fabricate two wafers separately and join them, and to raise integration and performance together with structures like CBA (Cell-over-Bonded-Array) that increase cell density.
These trends are summarized as follows.
- Density: 2D→3D stacking, 200-layer+ high stacking, wafer bonding — continued decline in per-capacity cost
- Interface: SATA→NVMe, PCIe generation upgrade (4.0→5.0→6.0) — bandwidth expansion
- Role sharing: ZNS·FDP share FTL functions with the host, minimizing WAF
- Compute proximity: computational storage·CXL·direct GPU link minimize data movement
On the software·architecture side, the trend to share the FTL's role with the host is clear. The ZNS (Zoned Namespace) SSD divides storage space into zones that allow only sequential writes, letting the host (filesystem·DB) directly control data placement. This reduces the write amplification that arose from the device-internal GC and the host's data-lifetime management being out of sync, lowering WAF close to 1 and reducing OP. Similarly, FDP (Flexible Data Placement) lets the host give data-lifetime hints so data of the same lifetime is written together. Both approaches show the common direction of moving away from the traditional FTL that "the device hid all alone" toward host-device cooperation.
On the memory-hierarchy side, CXL (Compute Express Link)-based memory expansion and computational storage that performs computation inside the storage device are rising. Processing filtering, compression, and search inside the SSD without pulling large data to the CPU can reduce data movement and power, and research and commercialization are underway as a mitigation for the I/O bottleneck of AI·analytics pipelines. This, together with PIM, is the storage-device version of the larger trend of "moving computation closer to the data."
AI training·inference demand is also changing SSD design. As storage bandwidth emerged as a new bottleneck in the process where GPUs repeatedly read datasets and checkpoints of hundreds of GB to several TB, a path where the GPU reads data directly from the SSD without going through the CPU (e.g., the GPUDirect Storage family) and KV cache·vector offloading designs that place sparse embeddings and vectors in flash and read only the needed parts are drawing attention. In short, the SSD is being redefined not as a passive storage medium but as an active component of the data pipeline, and this directly touches vector databases and large-scale embedding serving.
6. Considerations and Implications
SSD adoption·operation is not a mere medium swap but a problem of multidimensional optimization of performance, lifetime, and cost. From a professional engineer's perspective, the following should be considered.
- Workload-based cell·OP selection (application strategy): Place high-endurance TLC/SLC with large OP for write-intensive use (DB journals·logs), and QLC for read-centric archives. Design lifetime margin quantitatively by contrasting DWPD (Drive Writes Per Day)·TBW with the workload's write volume, and make it the warranty basis at procurement.
- Performance stability vs. benchmark (trade-off): Accounting for latency spikes (tail latency) from SLC-cache exhaustion and GC foreground operation, verify not with peak figures but with sustained performance and p99 latency. Reconcile the conflict that increasing OP improves WAF and latency but worsens usable capacity and unit cost.
- Data security·disposal: Because flash has no in-place overwrite, there is a risk of old-data remnant, so adopt not simple deletion but encryption-based instant disposal (Instant Secure Erase)·Sanitize. Self-encrypting (SED)·full-lifetime encryption is essential not only for disposal but also for preventing information leakage on loss or return.
- Related technology·outlook: A design that vertically integrates NVMe namespaces·OP·ZNS/FDP with the filesystem·DB to lower WAF governs TCO. In the mid-to-long term, with ties to CXL memory pooling·computational storage, the boundary between storage and memory is expected to blur, so a perspective that designs the storage layer not as a fixed medium but as the target of a tiered data-placement (hot/warm/cold) policy is required.
References
- NVM Express Base Specification, https://nvmexpress.org/specifications/
- Western Digital, "Zoned Storage (ZNS)", https://zonedstorage.io/
- JEDEC SSD Endurance Workloads (JESD219), https://www.jedec.org/
In one line: An SSD is a semiconductor storage device that hides NAND flash's constraints of "no in-place overwrite, block-granular erase, finite lifetime" with the FTL's log-structured mapping, garbage collection, wear leveling, and ECC to present a block device, and workload-matched choices of cell·OP·interface together with host-cooperative architectures like ZNS·computational storage govern performance, lifetime, and TCO.