Skip to main content

When processing packets at very high rates, performance is rarely determined by one big optimization. Instead, it comes from removing small sources of overhead from the packet-processing path. At 5×9 Networks, we have achieved around 400 million packets per second (Mpps) on a single x86 server. With SR-IOV, a NIC can expose Virtual Functions (VFs) that can be assigned directly to virtual machines, allowing packets to bypass much of the host networking stack. But delivering the packet directly to the VM is only part of the story. Once it gets there, it still needs to be processed — and this is where DPDK comes in.

Bypassing the guest networking stack

In a traditional Linux application, the guest NIC driver receives a packet, the operating system processes it through the networking stack, and eventually the application receives the data through a socket. That model works well for many applications, but when forwarding millions of packets per second, every additional operation matters.

DPDK — the Data Plane Development Kit — takes a different approach. Instead of allowing the Linux networking stack to own the network interface, DPDK can bind an SR-IOV Virtual Function to a userspace driver such as vfio-pci. The interface is removed from the guest kernel networking stack, and the application communicates directly with the NIC using a DPDK Poll Mode Driver (PMD).  At a high level, the application continuously receives bursts of packets from NIC queues, processes them and sends them back out.

rte_eal_init();
configure_ports_and_queues();
create_mempool();

while (running) {
    rx_burst();
    process_packets();
    tx_burst();
}

The simplicity of this loop hides several important optimizations that make DPDK fast.

Polling instead of interrupts

One of the first things people notice about DPDK is that a CPU core running a polling loop can sit close to 100% utilization. Traditional networking relies heavily on interrupts: a packet arrives, the NIC generates an interrupt, the CPU switches to the interrupt handler, the kernel and NIC driver process the packet, and eventually the application is notified that data is available. At low packet rates this is efficient because the CPU can do something else while there is no traffic, but at millions of packets per second repeatedly handling interrupts and switching execution contexts introduces significant overhead.

DPDK’s Poll Mode Drivers instead continuously ask the NIC whether packets are available. There are no packet-arrival wakeups and no interrupt-driven processing loop. Because the CPU repeatedly executes the same code and accesses the same data, its caches can remain warm, reducing memory-access latency. The trade-off is that a CPU core remains busy even when traffic is low, in exchange for predictable processing and high throughput. Think of a supermarket cashier: instead of waiting in the break room until every customer rings a bell, the cashier stays at the register, ready to serve the next customer immediately.

Hugepages reduce memory translation overhead

Packet processing is extremely dependent on memory performance. Modern CPUs use virtual memory, meaning virtual addresses have to be translated into physical addresses. CPUs cache these translations in the Translation Lookaside Buffer (TLB), but when a translation isn’t available there, a page-table walk is required, introducing additional memory accesses and latency.

Standard Linux memory pages are typically 4 KB, while DPDK commonly uses 2 MB or 1 GB hugepages. A 1 GB memory region requires more than 260,000 standard 4 KB pages, but only 512 2 MB hugepages. By allocating packet-processing memory from hugepages, DPDK reduces the number of translations needed to cover the same amount of memory and therefore reduces TLB misses. At hundreds of millions of packets per second, even small reductions in memory-access latency can result in measurable throughput gains.

Reusing packet buffers

Memory allocation itself can become expensive when repeated millions of times per second, so DPDK doesn’t normally allocate a new memory object every time a packet arrives. Instead, every packet is represented by an rte_mbuf, a lightweight structure containing information about the packet and a reference to its data. These buffers are preallocated and stored in an rte_mempool. When a packet arrives, DPDK takes a free buffer from the pool, uses it to receive and process the packet, and returns it when it is no longer needed. This avoids millions of malloc() and free() calls.

A supermarket provides a useful analogy: customers don’t get a newly manufactured shopping basket when they enter the store. Thousands of baskets already exist; customers take one, use it, return it, and the next customer reuses it. DPDK does essentially the same thing with packet buffers.

Processing packets in bursts

DPDK also typically processes packets in bursts rather than one at a time. Every interaction with the NIC has some fixed processing overhead. If the application receives one packet at a time, it pays that cost for every packet. If it receives 32 packets in a single burst, the fixed cost is shared across all 32.  Batch processing also improves cache locality: once the CPU has loaded packet-processing instructions and related data into its caches, additional packets can reuse them instead of repeatedly fetching the same information.

Think of unloading a delivery truck with a forklift. You could drive to the truck, pick up one box, drive back to the warehouse and repeat, or load an entire pallet and move dozens of boxes in one trip. The distance is the same, but each trip moves much more cargo. Burst processing applies the same principle to packets.

Keeping packet processing NUMA-local

High-performance x86 servers often contain multiple CPU sockets, each with its own memory controller and locally attached RAM — an architecture called Non-Uniform Memory Access (NUMA).  If a DPDK worker running on a core belonging to CPU 0 accesses packet buffers stored in memory attached to CPU 1, those accesses have to cross the inter-socket connection. This introduces additional latency into the processing of every packet.

The solution is to keep the processing path local. Packets received from a NIC connected to a particular NUMA node should use buffers allocated from memory on that node and be processed by CPU cores belonging to the same node. This is also why DPDK applications commonly pin worker threads to dedicated CPU cores: moving a worker to another core or NUMA node can destroy locality and hurt performance. Think of a forklift that has to drive to another warehouse every time it needs a pallet — keeping the pallets in the same warehouse eliminates that unnecessary travel.

Small optimizations multiplied by millions of packets

There is no single trick that makes DPDK fast. Its performance comes from combining several techniques: bypassing the kernel networking stack, polling instead of relying on packet interrupts, using hugepage-backed memory, preallocating and reusing packet buffers, processing packets in bursts, and keeping packet processing NUMA-local.

Each optimization removes a small amount of work from the packet-processing path. A few CPU cycles saved, one less allocation, fewer address translations or fewer NIC interactions may seem insignificant for a single packet. At hundreds of millions of packets per second, however, those small savings become a substantial amount of compute capacity.

That is the core idea behind high-performance packet processing on general-purpose x86 hardware: make the work performed for each packet as small, predictable and local as possible.