NVMe (Non-Volatile Memory Express) and NVMe-oF
1. Overview
A. Definition
NVMe (Non-Volatile Memory Express) is a host–controller interface specification and command set designed to fully extract the parallelism and low-latency characteristics of non-volatile storage devices (mainly NAND flash and next-generation memory SSDs) connected directly to the PCIe (PCI Express) bus. NVMe-oF (NVMe over Fabrics) is a transport specification that extends this NVMe command and queue model over network fabrics (RDMA, Fibre Channel, TCP), allowing remote storage to be accessed with low latency as if it were a local NVMe device.
The essence of NVMe lies in resolving the mismatch that "the storage medium changed but the access protocol stayed the same." Older AHCI/SATA and SAS were interfaces built on the premise of rotating disks (HDDs), assuming a single command queue, shallow queue depth, and a long command-processing path. Flash SSDs, however, have no mechanical seek latency and can internally operate dozens of channels and dies in parallel; using an HDD-era protocol unchanged therefore lets the software layer become a bottleneck that blocks the medium's parallelism. NVMe is a specification that makes multi-queue, low command overhead, and per-core parallel submission first-class design goals precisely to remove this bottleneck.
B. Background and Need
As flash SSDs became widespread in the late 2000s, the access latency of the disk itself dropped from milliseconds (ms) to tens of microseconds (µs), a reduction of two orders of magnitude or more. Yet connecting these SSDs via conventional SATA/AHCI ran into a structural limit: AHCI permits only 1 queue with a queue depth of 32, so no matter how fast the medium is, commands cannot be pushed in fast enough. Moreover, processing a single AHCI command requires multiple register accesses (MMIO) and interrupt handling, so the per-command software overhead became relatively large compared with the medium's latency. In short, a structural mismatch of "a slow protocol on a fast medium" arose.
NVMe fundamentally redesigned this mismatch. The host places Submission Queues (SQ) and Completion Queues (CQ) in pairs in memory, writes a command into the queue, and then notifies the controller with a single doorbell register write. In theory up to 64K queues (65,535) can exist, each holding up to 64K commands, so each core of a multi-core CPU submits commands through its own queue in parallel, without lock contention. The command set is also simplified to a small number of required commands, shortening the per-command path. As a result, the same SSD attached via NVMe can deliver hundreds of thousands to millions of IOPS and tens-of-µs latency.
On top of this, as flash got faster, the physical constraint of "a local SSD plugged into a single server" became a new bottleneck. In cloud and virtualized environments, compute and storage must be disaggregated so resources can be scaled and shared independently, but if the low latency of local NVMe cannot be preserved across the network, the benefit of disaggregation is canceled out. NVMe-oF began from exactly this need, encapsulating NVMe's queue and command model almost verbatim over a network transport to minimize added latency on remote access (targeting a single-digit-µs addition with RDMA transport). In short, NVMe resolves the "medium–protocol mismatch" and NVMe-oF the "local dependency" — two complementary axes.
C. Key Characteristics
The nature of the NVMe family can be summarized in four points. First, deep and wide parallelism — numerous queues and deep queue depth ensure software does not mask the medium's internal parallelism. Second, low protocol overhead — doorbell-based submission, MSI-X interrupts, and a simplified command set minimize per-command CPU cost. Third, a scalable layered structure — namespaces logically partition a single controller, and the transport layer is abstracted so that PCIe and fabrics are handled with the same command model. Fourth, transport independence (transport abstraction) — in NVMe-oF the upper command semantics are identical regardless of whether RDMA, FC, or TCP is used. These characteristics form the basis for the later judgment of "local or disaggregated, and which transport to choose."
2. NVMe Architecture and Operating Principles
A. Overall Structure and Core Components
The overall structure of NVMe can be divided into four layers: host (driver), transport (PCIe), controller, and namespace. The host-side NVMe driver creates and manages queues between the OS block layer and the device, while the controller pulls commands piled in the queues and reflects them onto the medium through the internal Flash Translation Layer (FTL). A namespace is a unit that logically partitions a single physical SSD; it corresponds to a SCSI LUN and provides storage space isolated per virtual machine or tenant.
flowchart TB
subgraph HOST["Host"]
APP["Application / Filesystem"]
DRV["NVMe driver<br/>(queue creation·doorbell mgmt)"]
AQ["Admin Queue (Admin SQ/CQ)"]
IQ["I/O queue pairs (many per-core SQ/CQ)"]
end
subgraph BUS["Transport (PCIe)"]
DB["Doorbell register (MMIO)"]
end
subgraph CTRL["NVMe controller"]
ARB["Command arbitration"]
FTL["FTL / mapping·GC·wear-leveling"]
NS1["Namespace NS1"]
NS2["Namespace NS2"]
end
APP --> DRV --> AQ
DRV --> IQ
IQ --> DB --> ARB
AQ --> DB
ARB --> FTL --> NS1
FTL --> NS2
ARB -. "MSI-X interrupt (completion notice)" .-> DRV
The main components are summarized as follows.
- Admin Queue: a single queue pair that handles control commands such as controller identification, I/O queue creation/deletion, and namespace management.
- I/O queue pairs (SQ/CQ): queues over which actual read/write commands flow, with one or more per core to enable parallel submission.
- Doorbell register: the MMIO window through which the host, after placing a command in a queue, notifies the controller of the tail position, and updates the head position after completion processing.
- Namespace: the set of logical blocks the controller exposes, serving as the unit of formatting, encryption, and isolation.
- FTL (Flash Translation Layer): the controller's internal firmware layer that performs logical-to-physical address mapping, garbage collection, and wear leveling.
B. Command Processing Flow and Queue Model
NVMe command processing follows a simple producer–consumer model of "put it in a queue → notify via doorbell → receive completion via interrupt." The sequence diagram below shows the path from submission to completion of one read command. The key point is that command submission ends with a single register write, and completion notification can be batched per queue (interrupt coalescing), making the per-command CPU cost very low.
sequenceDiagram
participant H as Host driver
participant SQ as Submission Queue (SQ)
participant DB as Doorbell
participant C as Controller
participant CQ as Completion Queue (CQ)
H->>SQ: Write command descriptor (tail++)
H->>DB: SQ Tail doorbell write (once)
DB->>C: "Command available" notice
C->>C: Fetch command·transfer data via DMA
C->>CQ: Write completion entry (status·result)
C-->>H: MSI-X interrupt (completion)
H->>DB: CQ Head doorbell write (reflect processing done)
In this model, the factors governing performance are the queue's count, depth, and arbitration policy. If there are as many queues as cores, inter-core lock contention disappears; if queue depth is deep, the medium's internal parallel channels can be filled. When pulling commands from multiple SQs, the controller applies round-robin or weighted-round-robin (WRR)/priority-based arbitration, so it can give QoS such as processing a latency-sensitive transaction queue ahead of a batch queue. In practice, when a latency-sensitive workload like a database and a throughput-oriented workload like backup share one device, queue and arbitration design becomes the key variable that determines service quality.
C. Structural Differences versus AHCI/SATA
The difference between NVMe and AHCI stems not from mere performance numbers but from a difference in design philosophy. AHCI, premised on a single queue and shallow depth, is tuned for HDD serial processing, whereas NVMe, with multiple queues and deep depth, presumes flash parallelism. The table below summarizes the core differences, while "why" each item matters is explained in the paragraph after the table.
| Category | AHCI / SATA | NVMe |
|---|---|---|
| Queue count | 1 | up to ~64K |
| Queue depth | 32 | up to ~64K per queue |
| Command submission | multiple register accesses | single doorbell write |
| Interrupts | single | MSI-X (multiple per queue) |
| Transport | SATA 6Gbps, etc. | PCIe lanes (several GB/s↑ per generation) |
| Typical latency | tens to ~hundred µs | a few to tens of µs |
What matters more than the numbers themselves are their implications. With only 1 queue, every core on a multi-core server contends on that one queue, so adding cores does not scale IOPS linearly. NVMe removes this contention with per-core queues, securing scalability proportional to core count. Also, AHCI needs several register accesses per command, adding a few µs of software overhead per command — a share that cannot be ignored on flash whose medium latency has shrunk to tens of µs. NVMe's single doorbell write and interrupt coalescing drive down exactly this share. For example, even with the same TLC-based SSD, SATA models typically stay in the range of tens of thousands to ~100K IOPS, whereas PCIe 4.0 NVMe models commonly deliver hundreds of thousands to over 1 million IOPS, and much of this gap comes not from the medium but from interface design.
3. NVMe over Fabrics (NVMe-oF) Architecture
A. Design Goals and Transport-Layer Abstraction
With the goal of "preserving the low latency of local NVMe across the network," NVMe-oF carries NVMe's queue and command capsules over a network transport. The core design is a separation of the command layer and the transport layer. The upper NVMe command semantics are kept identical whether local or remote, while which medium carries that command is handled by the transport binding (RDMA, FC, TCP). Thanks to this abstraction, applications and upper driver layers are almost unaware of the transport type.
flowchart LR
subgraph INIT["Initiator (Host)"]
FE["NVMe core (queues·command capsules)"]
T1["RDMA transport"]
T2["TCP transport"]
T3["FC transport"]
end
subgraph NET["Network fabric"]
ROCE["RoCE/iWARP/InfiniBand"]
ETH["Ethernet (TCP/IP)"]
FCN["Fibre Channel"]
end
subgraph TGT["Target (Storage)"]
TT["Transport endpoint"]
TC["NVMe controller"]
NS["Namespace (remote block)"]
end
FE --> T1 --> ROCE --> TT
FE --> T2 --> ETH --> TT
FE --> T3 --> FCN --> TT
TT --> TC --> NS
B. Characteristics by Transport and Selection Criteria
The transports NVMe-oF supports are distinctly different in nature, and the choice depends on latency, cost, and existing infrastructure. The table below compares the representative transports, while the basis for the actual choice is explained in the following paragraph.
| Transport | Added latency | Required infrastructure | Characteristics |
|---|---|---|---|
| RDMA (RoCE v2/iWARP/IB) | lowest (single-digit µs target) | RDMA NIC·lossless Ethernet (PFC/ECN) | top performance, high setup·operation difficulty |
| Fibre Channel (FC-NVMe) | low | FC HBA·SAN switch | reuse of existing FC-SAN assets, stable |
| TCP (NVMe/TCP) | relatively higher | standard Ethernet·NIC | general-purpose·low cost, easy large-scale expansion |
RDMA-based transport bypasses the kernel and CPU to directly put and get data in remote memory (zero-copy), giving the lowest added latency, but using RoCE v2 requires configuring lossless Ethernet (PFC·ECN·congestion control), raising network operation difficulty. FC-NVMe is advantageous when an organization already running an FC-SAN, such as finance or mission-critical systems, reuses its existing HBA and switch assets while switching only the block protocol from SCSI to NVMe. NVMe/TCP runs over standard Ethernet without special NICs, so its general applicability and cost efficiency are the best, and it is establishing itself as the de facto default choice when hyperscale and cloud providers disaggregate storage at large scale. In other words, "RDMA for lowest latency, FC-NVMe for existing FC assets, TCP for general-purpose large scale" is the primary criterion.
4. Comparisons and Application Cases
A. NVMe/TCP versus iSCSI
iSCSI, the traditional standard for remote block storage, carries SCSI commands over TCP/IP, but the underlying SCSI model is itself single-queue-oriented, limiting multi-core parallelism and low latency. NVMe/TCP, while using the same standard Ethernet, transports NVMe's multi-queue model as-is, so it often achieves higher IOPS and lower latency on the same network. That said, NVMe/TCP also traverses the TCP stack, so its latency is higher than RDMA and it is affected by congestion control and retransmission, meaning network quality directly determines performance. In practice, a hybrid configuration is common where an existing iSCSI SAN is gradually migrated to NVMe/TCP, while only a few especially latency-sensitive workloads are separately placed on an RDMA fabric.
B. Industry Application Cases
First, in cloud block storage, hyperscalers disaggregate compute nodes and storage nodes and connect them with NVMe-oF (mostly TCP), letting a VM use a remote volume that appears as a local disk with low latency. This allows storage capacity and compute to be scaled independently and the volume of a failed node to be quickly re-attached to another node. Second, in AI training infrastructure, so that loading datasets at hundreds of GB/s does not starve the GPU, NVMe-oF (RDMA) combined with GPUDirect Storage pushes data directly from storage into GPU memory. Third, in financial mission-critical systems, introducing FC-NVMe over an existing FC-SAN has been reported to lower transaction-system storage latency by tens of percent while keeping the proven stability of the FC infrastructure. All three cases show the common lesson that "even with the same medium, interface and transport choices govern overall performance and the operating model."
5. In-Depth — NVMe 2.0 and Next-Generation Extensions
The recent trends in the NVMe ecosystem can be summarized as modularization of the specification and data placement tailored to medium characteristics. NVMe 2.0, released in 2021, split the single giant specification into a base specification + command-set specifications (NVM·ZNS·Key-Value) + transport specifications (PCIe·RDMA·TCP·FC) + management interface (NVMe-MI). This modularization lets a new command set or transport be added independently without shaking the whole specification, accelerating the ecosystem's evolution.
The most noted extension is ZNS (Zoned Namespaces). ZNS divides a namespace into many zones and constrains each zone to allow only sequential writes. Doing so greatly reduces garbage collection inside the SSD and the resulting write amplification, and lowers over-provisioning and DRAM requirements. Especially on QLC flash, which has many bits per cell and thus weak endurance, ZNS is regarded as a means to improve lifespan and cost efficiency at once. The principle is that when the host places zones according to data lifetime, data that will be erased together gathers in the same physical region, raising erase efficiency. A more flexible approach is the relatively recently standardized FDP (Flexible Data Placement), which, without enforcing sequential-write constraints like ZNS, lets the host give hints about data placement to reduce write amplification.
The relationship with CXL (Compute Express Link) is also an in-depth point. If NVMe represents block-unit, queue-based storage access, CXL aims at byte-unit, cache-coherent memory access. The two are not competitors but complement each other by dividing layers: a memory–storage tiering that places hot data in the CXL memory tier and persistent, large-volume data in the NVMe tier is discussed as the direction of next-generation server architecture. In addition, features such as CMB (Controller Memory Buffer) and PMR (Persistent Memory Region) provide peer-to-peer DMA that bypasses host memory and a persistent buffer, further shortening the data-movement path between storage and accelerators or networks.
6. Considerations and Implications
From a professional engineer's perspective, adopting NVMe/NVMe-oF should be approached not as a simple device swap but as a comprehensive design of architecture, operations, and economics.
- Trade-offs in transport choice: RDMA gives the lowest latency but demands lossless-Ethernet configuration and operational expertise, while TCP is general-purpose and low-cost but relatively higher latency. Evaluating the organization's network maturity, workload latency sensitivity, and existing assets (whether an FC-SAN exists) together, it is more realistic to design a mix of transports by workload rather than a single uniform choice.
- Outlook on bottleneck migration: once NVMe removes the medium bottleneck, the bottleneck shifts to the CPU, network, and software stack. Therefore, upper-layer optimizations such as polling-based I/O (user-space drivers like SPDK), interrupt coalescing, and offloading the storage datapath to a DPU must proceed in parallel for NVMe's performance to be actually felt.
- Security and multi-tenancy: namespace isolation alone is insufficient; for remote access, transport encryption (the secure channels of FC-NVMe·NVMe/TCP), authentication·access control, and medium encryption based on self-encrypting drives (SED)·TCG Opal must be considered together. In particular, when multiple tenants share one physical SSD in the cloud, QoS isolation and secure-erase guarantees are directly tied to regulatory and audit requirements.
- Adoption strategy and related technologies: the value of NVMe-oF is maximized when combined with compute-storage disaggregation, software-defined storage (SDS), and dynamic volume provisioning through Kubernetes' CSI driver. Furthermore, through a QLC-based large-capacity low-cost tier using ZNS·FDP, tiering with a CXL memory tier, and combinations with data-protection techniques such as erasure coding and snapshots, the balance point among performance, durability, and cost must be designed at the architectural level.
References
- NVM Express, "NVM Express Base Specification 2.0" — https://nvmexpress.org/specifications/
- NVM Express, "NVMe over Fabrics" overview — https://nvmexpress.org/
- SNIA, "NVMe and NVMe-oF educational materials" — https://www.snia.org/
In one line: NVMe is a PCIe storage interface that extracts flash parallelism through a multi-queue, low-latency command model, and NVMe-oF is a remote-transport specification that extends this model over RDMA·FC·TCP fabrics to enable compute-storage disaggregation.