Skip to content
INTERNET INFRASTRUCTURE & NETWORKS

Presto Solves the Data Centre "CPU Tax" by Bringing Full TCP Stacks to Programmable Hardware

In modern data centres, the Transport Control Protocol (TCP) has quietly become one of the most resource-intensive tasks managed by a Central Processing Unit (CPU). While network hardware has scaled forward at an impressive pace while maintaining reasonable power footprints—such as single NVIDIA ConnectX-7 Network Interface Controllers capable of handling 400Gbps and 300 million packets per second at roughly 25 watts—software network stacks have failed to keep up.

Today, these software stacks represent the dominant bottleneck in data centre operations, inhibiting the ability of systems to translate raw hardware gains into tangible application performance. Even when utilizing the most advanced kernel-bypass stacks, communication-intensive applications can burn up to 74% of their CPU cycles solely within the transport layer, rather than executing actual application logic. As link speeds steadily climb toward terabits, this compounding "CPU tax" manifests as dozens of wasted cores, hundreds of watts of excess power consumption, stranded network bandwidth, and severely inflated tail latencies.

For years, data centre operators have been forced to live with an uncomfortable compromise. Traditional software TCP offers flexibility, interoperability, and robustness, but it demands heavy CPU execution. Conversely, hardware-implemented transports such as Remote Direct Memory Access (RDMA) and TCP offload engines deliver exceptional speed and efficiency, but they remain notoriously rigid and brittle to operate at cloud scale.

RDMA serves as a primary example of this limitation. Originally designed for tightly controlled, lossless fabrics at a significantly smaller scale, its retrofit into large-scale data centre networks has proven error-prone and slow, gated entirely by rigid hardware iteration cycles. When a network transport is permanently fixed in silicon, operators lose the vital ability to debug, trace, manage, and adapt the stack quickly to meet emerging application demands or shifting deployment requirements. Consequently, hardware-implemented transports have largely remained confined to specialized silos like storage clusters and high-performance computing environments, leaving software stacks as the default choice everywhere else.

Presto: A match-action TCP stack for the terabit era | APNIC Blog

A new open-source project called Presto aims to break this long-standing tension. Developed through a collaborative effort between the University of Washington and the Max Planck Institute for Software Systems (MPI-SWS) by researchers Rajath Shashidhara, Antoine Kaufmann, and Simon Peter, Presto is built directly upon the Reconfigurable Match-action Table (RMT) architecture. This is the exact design that underlies modern programmable network switches like the Intel Tofino, alongside a growing segment of SmartNICs such as AMD’s Pensando product line.

The RMT architecture is uniquely capable of sustaining deterministic, line-rate packet processing at billions of packets per second, achieving sub-microsecond latency and Application-Specific Integrated Circuit (ASIC)-class power efficiency, while remaining fully software programmable through Programming Protocol-Independent Packet Processors (P4).

However, leveraging this architecture for TCP has historically presented a fundamental design conflict. TCP’s complex, dynamic state machine is notoriously ill-suited to RMT’s strictly unidirectional execution model.

Challenges for TCP on RMT

To understand the core technical mismatch, one must examine how the RMT architecture processes network traffic. The system parses an incoming packet into a header vector and subsequently pushes that vector through a fixed sequence of match-action stages. Each stage performs simple Arithmetic Logic Unit operations against stage-local memory in absolute lock-step.

Every single packet traverses the exact same path in the exact same number of clock cycles, and each individual stage carries only a strictly limited budget of match-action logic. While this architectural rigidity is the foundational source of RMT’s determinism and efficiency, it simultaneously creates significant programming difficulties.

Presto: A match-action TCP stack for the terabit era | APNIC Blog

TCP, by contrast, relies on complex and highly interdependent state updates. Window boundaries, sequence numbers, and out-of-order packet bookkeeping frequently need to be updated multiple times and in arbitrary orders. An RMT pipeline permits none of this natively, as state is strictly stage-local, packets can only move forward through the pipeline, and packets cannot revisit earlier stages. Direct implementations of TCP on RMT historically force engineers into problematic compromises that either break the protocol logic or exhaust hardware resource limits.

Reconciling TCP State with RMT Constraints

To overcome these structural limitations, Presto introduces a series of specialized design techniques tailored to handle TCP state within strict RMT constraints. Rather than trying to force a traditional state machine into programmable silicon, the system re-architects how transport logic interacts with the packet pipeline.

The mechanics of these design principles, which successfully adapt complex TCP state management to match-action operations, are detailed in the research paper presented at the ACM SIGCOMM conference. By systematically resolving these hardware-software mismatches, Presto unlocks ASIC-level performance and power efficiency while preserving the operational flexibility traditionally associated with software-based stacks.

A Modular Data-Path, and Why That Matters

Unlike traditional fixed-function hardware offloads, Presto features a data path that is programmable end-to-end. To make this architectural approach tractable in practice, the RMT data path executes core transport logic—including window management, packet reassembly, acknowledgment generation, sequence-to-address translation for Direct Memory Access, and rate control—not as a monolithic block, but as a sequence of decoupled, independently programmable functional blocks.

Each of these functional blocks spans several match-action stages within the pipeline. This deliberate modularity allows network operators to evolve transport semantics, alter congestion-control algorithms, or integrate application-specific co-designs simply by editing a few lines of P4 code, entirely avoiding the need to redesign underlying silicon pipelines.

Presto: A match-action TCP stack for the terabit era | APNIC Blog

Presto Redefines the Software TCP Design Space

The research team built a fully functional prototype of Presto utilizing an Intel Tofino 2 switch hosted on a Netberg Aurora 810 hardware platform. They subsequently compared its performance against standard Linux networking stacks, the TCP Acceleration Service (TAS) kernel-bypass stack, and traditional RDMA implementations.

The evaluation demonstrated that Presto successfully dissolves the historic trade-off between performance, efficiency, and flexibility. By maintaining standard POSIX socket interfaces, unmodified applications running on standard operating systems can leverage the high-performance transport layer without requiring recompilation or source code modifications. This compatibility extends to common data centre workloads and applications such as memcached, NGINX, and the Storage Performance Development Kit.

Conclusion

Presto allows enterprise operators to retain the transport protocols they already deploy and thoroughly understand—standard TCP coupled with a familiar POSIX interface—while requiring zero application-level modifications. The system achieves ASIC-class efficiency in both CPU utilization and power consumption, delivering terabit-scale performance alongside microsecond-level tail latencies. Because the entire data path is programmable in P4, transport logic can evolve organically on a standard software release cycle rather than being chained to a hardware vendor’s multi-year silicon manufacturing schedule.

More broadly, Presto successfully redefines the design space for high-performance TCP stacks by demonstrating that comprehensive transport functionality can fit comfortably within strict RMT hardware constraints. While initially rooted in TCP, these foundational principles are not restricted to a single protocol. Any reliable network transport shares a nearly identical structural pattern: tracking in-flight data, recovering efficiently from packet loss, and regulating send rates. Successfully mapping stateful protocols onto programmable hardware under tight timing and resource budgets addresses a universal challenge across modern data planes.

The complete Presto framework has been released as open-source software, built specifically for the Intel Tofino 2 and validated on Netberg Aurora 810 hardware paired with ConnectX network interface cards. The public repository includes an extensive experimentation suite supporting remote procedure call scalability, packet loss testing, incast scenarios, performance isolation benchmarks, key-value stores, shared logs, and NVMe-oF testing. Because it speaks standard TCP natively, operators can run Presto directly alongside existing infrastructure to evaluate workloads side-by-side.

Leave a Reply

Your email address will not be published. Required fields are marked *