DPU (Data Processing Unit) and SmartNIC
1. Overview
A. Definition
A DPU (Data Processing Unit) is a data-movement-centric programmable processor that takes over infrastructure processing (networking, storage, security, virtualization, etc.) previously handled by the server's CPU and executes it on dedicated hardware. It is an evolution of the SmartNIC, which added programmable compute capability to the conventional NIC (Network Interface Card), and is often called "the third pillar of datacenter computing after the CPU and GPU."
To define the DPU in one phrase, it is "a chip responsible not for the result of computation but for the movement of data and the auxiliary processing attached to that movement." Whereas the CPU handles general-purpose control and serial logic and the GPU/NPU handle massively parallel computation, the DPU directly executes the entire data path—encryption/decryption, firewall rule application, virtual switching, storage protocol conversion—from the moment a packet enters the NIC until it reaches application memory, without going through the CPU. In short, the DPU is a domain-specific processor specialized not in "what to compute" but in "how to move data safely and quickly."
In terms of terminology, depending on the vendor, DPU (NVIDIA), IPU (Infrastructure Processing Unit, Intel), and more broadly SmartNIC are used interchangeably, but they all point to the same essence. That is, the concept of "moving the infrastructure burden the host CPU used to bear onto a dedicated processor sitting at the chokepoint the data passes through," and this article treats them collectively as the DPU. Differences in name are merely differences in marketing and architectural emphasis; the goals of offload, acceleration, and isolation are common to all.
The SmartNIC and the DPU lie on a continuous lineage but differ in maturity. Whereas early SmartNICs merely accelerated specific offloads (e.g., TCP checksum, VXLAN encapsulation) with fixed circuits, the DPU has general-purpose CPU cores (mainly Arm), dedicated accelerators, and its own memory and operating system, making it closer to a small server that runs independently of the host. This independence turns the DPU from a mere acceleration component into a new control point for security and operations.
B. Background and Necessity
The fundamental reason for the DPU's rise is the surge in infrastructure overhead known as the "Datacenter Tax." In cloud and microservice environments, a substantial portion of CPU cores is consumed by infrastructure processing such as virtual switching, overlay networking (VXLAN/Geneve), encryption, and storage virtualization that support the application traffic rather than the traffic itself. Various studies and vendor materials report that more than 20–30% of datacenter CPU cycles are spent on such infrastructure processing, which means the compute resource sellable to customers shrinks by that much. From a cloud provider's viewpoint, collecting this "tax" with a dedicated chip allows selling more vCPUs on the same hardware, leading directly to improved profitability.
The second reason is the explosion of network bandwidth. As server interfaces climbed from 25GbE to 100/200/400GbE, processing tens of millions of packets per second in software on a general-purpose CPU hit a physical limit. To process 64-byte small packets at line rate on a 400GbE link, one must handle roughly 600 million packets per second, which is hard to sustain with a per-packet CPU processing budget of tens to hundreds of nanoseconds. Without circuits dedicated to data movement, software can no longer keep up with the bandwidth increase.
The third is the demand for Zero Trust and security isolation. In multi-tenant environments densely packed with virtual machines and containers, controlling East-West traffic finely and physically separating the trust boundary between tenants requires a control point independent of the host CPU. The DPU can enforce firewalls, encryption, and policy from outside the host even if the host OS is compromised, becoming the foundation of a security architecture that separates the infrastructure control plane from the application execution environment. As these three pressures overlapped, the trend of "detaching infrastructure from the CPU onto dedicated chips" became a standard datacenter design.
C. Characteristics
The nature of the DPU is summarized by the so-called "3A (or O-A-I)" principle: ① infrastructure processing Offload, ② data-path Accelerate, and ③ infrastructure control Isolate. Freeing the CPU through offload, processing the data path at line rate with dedicated accelerators, and enforcing security and policy in an execution domain separated from the host—these three axes interlock to complete the DPU's value.
These three characteristics are not independent. For offload (①) to hold, data-path acceleration (②) must guarantee line rate, and because that processing is performed in an independent domain outside the host, isolation (③) naturally follows. Conversely, if a workload has a low proportion of infrastructure processing and is purely compute-centric, these three benefits weaken simultaneously. Therefore, the justification for adopting a DPU is directly proportional to "how large the infrastructure tax is in that environment."
Another essential characteristic is that the DPU is "Programmable." Unlike past offload NICs that provided only fixed functions, the DPU can redefine functions like software, so new overlay protocols, security policies, and telemetry-collection logic can be deployed without replacing hardware. Because of this programmability, cloud providers absorb rapidly changing requirements without recalling cards, and the DPU becomes not a component but the execution point of continuously evolving software-defined infrastructure.
2. Architecture and Operating Principles
The DPU's performance and isolation come from a structure that combines general-purpose cores, dedicated accelerators, and high-speed interconnects on a single card. The structural diagram below shows a typical system configuration in which the host server and the DPU cooperate.
flowchart LR
subgraph HOST["Host Server"]
HCPU["Host CPU (applications/VMs)"]
HMEM["Host memory"]
HCPU --- HMEM
end
subgraph DPU["DPU / SmartNIC"]
ARM["General-purpose cores (Arm multicore)"]
ACC["Dedicated accelerators (crypto/compression/regex)"]
ESW["Embedded switch (eSwitch)"]
MEM["On-board memory (DDR)"]
ARM --- MEM
ACC --- MEM
ESW --- ARM
end
HCPU -->|"PCIe high-speed bus"| ESW
ESW -->|"Ethernet 400G"| NET["Network fabric"]
The core of the structure is a three-tier division of labor. First, the general-purpose cores (mainly Arm Neoverse-class multicore) handle the Control Plane, processing complex but low-frequency logic such as routing-table management, policy updates, and running management agents. Second, the dedicated accelerators handle the Data Plane, processing high-frequency, regular operations such as encryption/decryption (IPsec/TLS), compression, regex matching, and hashing at line rate with fixed or semi-fixed circuits. Third, the embedded switch (eSwitch) classifies and forwards packets between physical ports and the host's virtual functions (VFs), replacing the software switch by offloading virtual-switching rules to hardware.
Viewing the operating principle from the data-path perspective, an arriving packet flows inside the DPU in the order classify → apply policy → accelerate → DMA transfer without going through the host CPU, and is placed directly into application memory. Here, connecting virtual functions directly to each VM via SR-IOV (Single Root I/O Virtualization) and combining RDMA (Remote Direct Memory Access) with kernel bypass (DPDK, etc.) minimizes CPU intervention and memory copies, lowering latency to the microsecond level. In short, the DPU implements in hardware "a shortcut for packets to reach destination memory while touching the CPU as little as possible."
Digging a little deeper into why this shortcut contributes to performance, one must understand that the main culprits eating away at performance in the traditional network stack are interrupts, context switching, and memory copying. In the software processing path, a kernel interrupt occurs for each packet, data is copied into user space, and in that process the CPU cache is polluted. The DPU maps hardware queues directly into application memory (zero-copy) and eliminates interrupts through polling-based processing, fundamentally removing these three overheads. As bandwidth rises to 400GbE, the allowable per-packet processing time shrinks to the nanosecond level, and under this physical constraint software optimization alone has clear limits, so a hardware data path becomes the only solution.
A particularly notable point is that the DPU boots its own operating system (mainly a lightweight Linux) and runs independently of the host. Thanks to this, even if the host hypervisor is compromised, the DPU maintains security policy in a separate trust domain, realizing the "off-host-ing of the infrastructure control plane" that physically separates the management plane. AWS Nitro, which detached virtualization, networking, storage, and security onto dedicated cards and chips to turn the host into a pure compute resource, is a representative commercial implementation of this design philosophy.
Understanding the two deployment modes clarifies DPU operational design. The first is Separated mode, in which the host and the DPU maintain their own management domains and cooperate only on offload. It can be adopted incrementally without greatly changing the existing server operations, making it suitable for early adoption. The second is DPU-managed mode, in which the DPU fully holds control over infrastructure provisioning and security policy and treats the host as an untrusted, pure compute resource. It is a form that pushes Zero Trust to the extreme and fits the requirement of cloud providers to guarantee isolation while not trusting the tenant host. Since the responsibility boundaries of the operations organization and the design of the automation pipeline differ depending on which mode is chosen, this is not a mere technical choice but an operational governance decision.
3. Major Offload Types and Processing Flow
The infrastructure processing the DPU collects is broadly divided into three domains: networking, storage, and security. The detailed process diagram below shows the typical flow in which virtual-switching offload occurs.
flowchart TD
A["Packet arrival (physical port)"] --> B{"Flow cache lookup"}
B -->|"Cache hit Fast Path"| C["Apply hardware rule (forwarding/NAT)"]
B -->|"Cache miss Slow Path"| D["Policy decision on DPU core (OVS control)"]
D --> E["Install rule into hardware table"]
E --> C
C --> F["Accelerator processing (encryption/encapsulation)"]
F --> G["Deliver to destination VM memory via DMA"]
The core of networking offload is lowering the data path of virtual switches such as OVS (Open vSwitch) into hardware. As in the diagram above, for the first packet (Slow Path) the DPU core decides the policy and installs the rule, and thereafter subsequent packets of the same flow (Fast Path) are processed directly in the hardware flow table. This makes the software switching that ran on the host CPU disappear, allowing several cores to be reclaimed. VXLAN/Geneve overlay encapsulation, NAT, and load-balancing (L4) rules are also offloaded to hardware in the same way.
In storage offload, remote storage is provided to the host as if it were a local disk via NVMe-oF (NVMe over Fabrics). Since the DPU handles storage protocol conversion, encryption, and compression on its behalf, the host sees only a standard NVMe device, and the complexity of the distributed storage behind it is completely hidden. This maintains local-level performance while physically separating compute and storage (Disaggregation).
The practical meaning of storage offload lies in improving resource Utilization. Traditionally, because servers embed local disks, the scaling units of compute and storage are tied together, causing the waste of having to scale the entire server even when only one side is short. When the DPU projects remote NVMe as if local, Composable Infrastructure that scales compute and storage independently as needed becomes possible, reducing resource stranding and raising overall datacenter utilization. On top of this, since the DPU transparently performs encryption at the storage tier (at-rest/in-transit), there is also the side benefit of meeting data-protection regulations without changing the application.
Security offload is the pinnacle of the DPU's value. By performing line-rate IPsec/TLS encryption, stateful firewalling, microsegmentation, DDoS mitigation, and deep packet inspection (DPI) outside the host, it enforces security policy at a trust boundary that tenant workloads cannot see. This becomes the foundation for implementing the previously discussed Zero Trust and microsegmentation at the hardware level, and makes it impossible for even a compromised host to bypass security controls.
The common principle running through the three offload domains is "Control/Data Plane Separation." Complex but infrequent decisions (policy decisions, exception handling, management) are handled flexibly in software by the DPU's general-purpose cores, while the simple but ultra-high-frequency actual data processing is entrusted at line rate to hardware tables and accelerators. Thanks to this separation structure, the DPU obtains both "flexibility (software) and performance (hardware)," which can be seen as bringing the control/data plane separation philosophy established in SDN down to the level of the server interface. Recently, programmable data planes that describe this data plane in a domain-specific language such as P4 (Programming Protocol-independent Packet Processors) and redefine new protocols and policies like software without replacing circuits are spreading.
4. SmartNIC Types and Component Comparison
SmartNICs/DPUs are divided into three families according to how programmability is implemented. Each approach has clear trade-offs in performance, flexibility, and development difficulty, and the reason this difference arises lies in the design choice of "up to where to harden into fixed circuits and from where to leave to software."
| Category | ASIC-based | FPGA-based | SoC (Arm core)-based |
|---|---|---|---|
| Performance | Highest (fixed circuit) | High (reconfigurable circuit) | High (cores + accelerators) |
| Flexibility | Low | Very high | High (SW programming) |
| Development difficulty | Low (fixed functions) | High (HDL design) | Medium (C/Linux) |
| Representative examples | Early offload NICs | Xilinx/Intel FPGA SmartNIC | NVIDIA BlueField, Intel IPU, AMD Pensando |
The ASIC approach implements specific offloads with fixed circuits for the highest power/performance efficiency, but cannot respond when new protocols or policies appear. The FPGA approach can reconfigure circuits for maximized flexibility, but requires HDL (Verilog/VHDL)-based development, raising the entry barrier and burdening unit cost and power. The SoC approach, the mainstream of today's DPUs, combines Arm multicore with dedicated accelerators so that the data plane is processed at line rate by accelerators and the control plane is programmed in the familiar C/Linux environment, balancing performance and development productivity. NVIDIA BlueField, Intel IPU (Mount Evans/E2000), and AMD Pensando are representative products of this family.
The reason the superiority among the three approaches has shifted over time lies in "the speed of change of the required flexibility." In the era when offload targets were stable like TCP checksum, ASIC fixed circuits were reasonable, but as the cloud constantly changed overlay and security policies, circuits could no longer be hardened. FPGAs provided that flexibility but hit the wall of development productivity and power efficiency, and ultimately the SoC approach that splits "frequently changing policy into core software, stable high-frequency operations into embedded accelerators" converged as the compromise. This evolutionary path shows that the control/data plane separation principle seen earlier is carried through even into the choice of hardware form.
The point where the DPU is distinguished from the CPU/GPU/NPU lies in "what it is specialized for." The table below organizes the role division of the four processors, and the prose that follows explains what implications that difference has in practice.
| Processor | Specialized area | Parallelism form | Role in the datacenter |
|---|---|---|---|
| CPU | General-purpose control/serial logic | Low-to-medium parallelism (tens of cores) | Applications/orchestration |
| GPU | Massively parallel floating-point compute | Ultra-large (thousands of cores) | AI training/graphics/HPC |
| NPU | Neural-network multiply-accumulate (MAC) | Regular array parallelism | AI inference acceleration |
| DPU | Data movement/infrastructure processing | Packet/flow parallelism | Networking/storage/security offload |
If the CPU specializes in general-purpose control/serial logic, the GPU in graphics/massively parallel floating-point compute, and the NPU in neural-network multiply-accumulate, the DPU specializes in data movement and infrastructure processing. These are complementary, not competitive, and in the datacenter the GPU handles AI computation while the DPU takes on the role of supplying data to that GPU safely and at high speed. In fact, in large-scale AI training clusters the DPU handles ultra-low-latency communication between GPUs (RDMA/RoCE), raising GPU utilization.
The point where the difference leads to practical implications is "the location of the bottleneck." In AI training, as models and data grow, what determines overall throughput is not the computation itself but how uninterruptedly data is supplied to the GPU, and if this supply path is left to CPU software the CPU becomes the bottleneck. The DPU processes this supply path in hardware on the CPU's behalf, reducing GPU idle time, thereby directly improving the return on investment (ROI) of expensive GPUs. In other words, the perspective that the DPU is not a competitor of the GPU but an auxiliary device that protects GPU investment should be the starting point of practical design.
5. Deep Dive: Recent Trends and Practical Cases
DPU technology is evolving rapidly along three branches. First, integration with AI infrastructure. In large-scale LLM training, when thousands of GPUs are connected the communication bottleneck governs the overall training time, and the DPU accelerates inter-GPU data exchange through GPUDirect, RDMA, and collective-communication offload, and reduces the communication burden by processing part of collective operations such as All-Reduce at the switch/NIC via In-Network Computing (e.g., SHARP). In the AI datacenter, the DPU is now settling in as a basic component rather than an option.
Second, hyperscalers' in-housing of their own silicon. AWS, with its Nitro System, offloaded virtualization, networking, storage, and security onto dedicated cards, effectively converting the EC2 host into a pure compute resource, and on top of this structure provides bare-metal-level performance and strong isolation simultaneously. Microsoft Azure develops FPGA-based SmartNICs (the past Catapult lineage) and its own DPU, and Google too has adopted an IPU (E2000) developed in cooperation with Intel; thus major clouds are vertically integrating infrastructure offload into their own hardware. This is a strategic choice not only for performance but also in terms of supply-chain and security sovereignty.
Third, convergence with confidential computing and Zero Trust. As the DPU provides a Root of Trust separated from the host, it is expanding into an integrated security platform that backs the previously discussed confidential computing, microsegmentation, and Zero Trust Network Access in hardware. For example, an architecture is being realized in which, even in a compromised hypervisor environment, the DPU verifies and encrypts tenant traffic and, by placing the management plane outside the host, blocks even administrator-privilege abuse. On the standardization front, the Linux Foundation's OPI (Open Programmable Infrastructure) project continues its effort to standardize the APIs and provisioning of DPUs/IPUs to reduce vendor lock-in.
The telecom field is also a strong application area for the DPU. The User Plane Function (UPF) of the 5G core network must encapsulate and route enormous packets per second through GTP-U tunnels, and processing this on a general-purpose CPU consumes many cores and causes jittery latency. Offloading the UPF data plane to a DPU/SmartNIC secures line-rate processing and deterministic low latency, stabilizing the quality of URLLC/edge services that demand ultra-low latency. This is a hardware-accelerated realization of the previously discussed SDN/NFV and is emerging as a key element supporting the cloud-native transformation of telecom infrastructure.
In concrete numbers, the host CPU cores reclaimed by infrastructure offload range from a few to several tens depending on the workload, which translates into an increase in sellable vCPUs per server and savings in power and floor space. Also, in a 400GbE environment, hardware line-rate IPsec greatly lowers latency and CPU usage compared with software encryption, so the higher the bandwidth, the greater the relative advantage of the DPU. However, since such figures vary widely by vendor and configuration, the principle at adoption time is to verify by measuring (PoC) with one's own workload.
To sketch a practical application, suppose that in a multi-tenant virtualization cluster of about 100 servers, each server consumes on average 6 physical cores on OVS software switching, overlay encapsulation, and tenant encryption. Offloading this to a DPU reclaims compute resources equivalent to 6 cores per server, about 600 cores across the whole cluster, and that much can be converted into sellable vCPUs. Added to this, the latency/jitter lowered by removing software switching improves the quality of latency-sensitive workloads (real-time trading, telecom core networks), and there is even the side benefit of meeting tenant-isolation audit requirements through management-plane separation. Of course, this benefit must be offset against the added cost of DPU card price, power, and operations staff, and since it may not exceed the break-even point in clusters with a low proportion of infrastructure processing, scale and traffic character must be weighed together.
6. Considerations and Implications
The DPU is a technology that changes the game of datacenter architecture, but it is not a panacea. The key is the balance between "the efficiency/security benefit obtained by detaching infrastructure processing from the CPU" and "the cost/operations/lock-in burden that new hardware adds," and from a professional engineer's perspective the adoption strategy must be carefully designed along the following four axes.
Adoption strategy (break-even from a TCO perspective): Since the DPU adds card unit cost and power, economics hold only when "the reclaimed CPU cores, power savings, and increased vCPU sales" exceed a threshold of scale that surpasses that cost. Therefore, multi-tenant clouds, AI clusters, and high-bandwidth storage environments with a large infrastructure tax are the priority targets, while the benefit is limited in small-scale, single workloads with a low proportion of infrastructure processing. One must always compute the break-even after a PoC with one's own traffic profile.
Trade-off (performance vs. operational complexity): Adopting a DPU adds yet another computer to manage (its own OS, firmware, agents) outside the host. This entails new operational burdens: firmware updates, vulnerability management, securing observability, and fault diagnosis. The performance gain and the increase in operational complexity must be evaluated together, and a lifecycle-management system must be prepared in advance.
Standardization and vendor lock-in response: Currently, DPUs have low portability because vendor-specific SDKs (NVIDIA DOCA, Intel IPDK, etc.) and programming models differ. One should adopt open standards such as OPI and DASH and P4-based programming, and design so that offload logic is wrapped in an abstraction layer and not directly coupled to a specific vendor's API, lowering the lock-in risk.
Redesign of the security/trust boundary and outlook: The DPU is a new Root of Trust and, at the same time, a new attack surface. Unless the DPU's own secure boot, firmware integrity, and privilege separation are strictly managed, the privileged control point may instead become a point of risk. Going forward, the DPU is expected to settle in as the datacenter's infrastructure-control-plane standard, combining with CXL-based memory pooling/composable infrastructure, In-Network Computing, and confidential computing, so one must draw an integration roadmap together with related technologies (CXL, SR-IOV, RDMA, Zero Trust).
In summary, the DPU/SmartNIC is not a "high-performance NIC" but an architectural shift that fundamentally redefines how the datacenter processes infrastructure, secures it, and places resources. Rather than individual product specs, the professional engineer must read integratively the design principles of offload, acceleration, and isolation and the resulting changes in organization, operations, and economics, and possess the discernment to judge the timing and scope of adoption based on the size of the infrastructure tax in one's own workloads.
References
- NVIDIA, "What Is a DPU?" (BlueField DPU technical overview) — https://blogs.nvidia.com/blog/whats-a-dpu-data-processing-unit/
- AWS, "AWS Nitro System" — https://aws.amazon.com/ec2/nitro/
- Open Programmable Infrastructure (OPI) Project, Linux Foundation — https://opiproject.org/
- Intel, "Infrastructure Processing Unit (IPU)" — https://www.intel.com/content/www/us/en/products/details/network-io/ipu.html
In one line: The DPU/SmartNIC offloads, accelerates, and isolates infrastructure processing such as networking, storage, and security from the CPU onto a dedicated programmable chip, simultaneously raising the datacenter's compute-resource efficiency and Zero Trust security as the third pillar of computing.