Modern data centres are facing a severe architectural bottleneck that has little to do with raw network hardware and everything to do with the software stacks running on central processing units. While modern network interface cards, such as the NVIDIA ConnectX-7, scale impressively to handle massive throughputs and packet rates at remarkably low power consumption, software network stacks have failed to keep pace. As a result, communication-intensive applications are losing a staggering proportion of their processing power to the transport layer rather than focusing on actual application logic.
Even when utilizing the most advanced kernel-bypass stacks, performance-heavy environments routinely burn up to 74% of their CPU cycles simply managing transport operations. As network link speeds steadily climb into the terabit era, this compounding "CPU tax" manifests as dozens of wasted processor cores, hundreds of watts of excess power consumption, stranded network bandwidth, and heavily inflated tail latencies. Cloud operators and data centre engineers have long been forced into an uncomfortable compromise between software flexibility and hardware rigidity.
Software-based Transmission Control Protocol implementations are flexible, robust, and universally interoperable, but they place a heavy burden on the CPU. Conversely, hardware-implemented transport mechanisms like Remote Direct Memory Access and dedicated TCP offload engines offer extreme speed and efficiency, but they are notoriously rigid and brittle to operate at cloud scale. RDMA serves as a prime example of this limitation, having been originally designed for tightly controlled, lossless fabrics at a much smaller scale.
Retrofitting RDMA into massive data centre networks has proven consistently error-prone and slow, gated entirely by rigid hardware iteration cycles. When a transport protocol is permanently fixed in silicon, operators lose the vital ability to debug, trace, manage, and quickly adapt the network stack to meet emerging application demands. Consequently, hardware transports remain largely confined to specialized silos such as high-performance computing clusters and dedicated storage systems, leaving software stacks as the default option everywhere else.

A collaborative research team has introduced Presto to dismantle this long-standing tension in data centre architecture. Developed jointly by researchers at the University of Washington and the Max Planck Institute for Software Systems—with contributions from Rajath Shashidhara, Antoine Kaufmann, and Simon Peter—Presto is built upon the Reconfigurable Match-action Table architecture. RMT serves as the foundational design underlying programmable network switches like the Intel Tofino and a growing ecosystem of SmartNICs, including AMD’s Pensando line.
The RMT architecture is uniquely capable of sustaining deterministic, line-rate packet processing at billions of packets per second. It delivers sub-microsecond latency alongside application-specific integrated circuit class power efficiency, all while remaining fully software programmable through Programming Protocol-Independent Packet Processors, commonly known as P4. However, integrating standard TCP into this environment presents a fundamental design conflict because TCP relies on a complex, highly interdependent state machine that is fundamentally ill-suited to RMT’s strictly unidirectional execution model.
Challenges for TCP on the RMT Architecture
To understand the core engineering hurdle, one must examine how the RMT architecture processes network traffic. The hardware parses incoming packets into a header vector and pushes that vector through a fixed, sequential series of match-action stages. Each stage performs simple arithmetic logic unit operations against local memory in absolute lock-step. Every single packet must follow the exact same path in the exact same number of clock cycles, and each stage is restricted to a very small, rigid budget of match-action logic.
This strict rigidity is the secret behind RMT’s remarkable determinism and efficiency, but it is precisely what makes programming complex protocols so difficult. TCP requires intricate, interdependent state updates where window boundaries, sequence numbers, and out-of-order packet bookkeeping might need to be modified repeatedly and in arbitrary sequences. An RMT pipeline permits none of this flexibility, as state is strictly stage-local, packets can only ever move forward through the pipeline, and they cannot revisit earlier processing stages. Directly mapping TCP onto such an environment forces severe design compromises.
Reconciling TCP State with RMT Constraints
To overcome these structural limitations, Presto introduces a series of dedicated design techniques that directly address each constraint imposed by the RMT architecture. Rather than treating the transport protocol as an indivisible monolith, Presto breaks the data path down into a modular, decoupled sequence of functional blocks.

The core transport logic—encompassing window management, packet reassembly, acknowledgment generation, sequence-to-address translation for direct memory access, and sophisticated rate control—is distributed across several independent match-action stages. This modularity means that network operators can evolve transport semantics, alter congestion-control algorithms, or integrate application co-designs by modifying a few lines of P4 code rather than undertaking a complete hardware redesign.
Redefining the Software TCP Design Space
To validate their approach, the researchers built a full working prototype deployed on an Intel Tofino 2 switch using a Netberg Aurora 810 hardware platform. They then compared Presto directly against standard Linux networking, the TCP Acceleration Service kernel-bypass stack, and traditional RDMA implementations. The results indicate that Presto successfully dissolves the historical trade-off between performance, efficiency, and operational flexibility.
By keeping the application interface entirely standard, Presto allows operators to maintain the transport protocols they already deploy and understand. Unmodified applications running over a standard POSIX sockets interface can leverage the new architecture without requiring recompilation or source code modifications. This means established workloads such as NGINX, Memcached, and storage performance development kit applications can run seamlessly on top of Presto while delivering ASIC-class efficiency in both CPU utilization and power consumption.
Presto achieves terabit-scale performance coupled with microsecond-level tail latencies. Furthermore, because the underlying data path is fully programmable using the P4 language, transport logic can evolve on a standard software release cycle rather than being chained to a hardware vendor’s lengthy silicon manufacturing timeline.
The underlying principles developed for Presto extend beyond standard TCP. Any reliable transport protocol shares the same fundamental structure of tracking in-flight data, recovering from packet loss, and regulating transmission rates. Mapping such stateful protocols onto programmable hardware under tight timing and resource budgets addresses a universal challenge in modern programmable data planes.

The complete Presto project has been made available as open source software, complete with an experiment framework that supports remote procedure call scalability, packet loss testing, incast scenarios, performance isolation, and key-value store benchmarks. The open source repository provides automated analysis and plotting tools, enabling researchers and network engineers to deploy the architecture in their own clusters and evaluate custom workloads.
Further details regarding the research and technical implementation are documented in academic papers presented at the ACM SIGCOMM conference. The project continues to invite feedback and experimental results from the broader systems and networking community as data centres increasingly seek alternatives to traditional software transport stacks.
Leave a Reply