As artificial intelligence models rapidly evolve into complex, multi-turn agentic systems, the infrastructure supporting them is facing an unprecedented bottleneck. Industry leaders widely view the traditional "memory wall"—the latency and bandwidth limitations that occur when memory cannot keep pace with high-speed processors—as the single greatest hurdle to scaling AI deployments efficiently. To combat this physical and architectural constraint, moving memory closer to compute has emerged as the industry-standard solution.
Astera Labs is addressing this challenge head-on with a sweeping update to its product portfolio. The company has introduced its new Leo X-Series fabric-attached memory controller alongside the next-generation Leo 2 E-Series and P-Series controllers. Together, these advanced silicon solutions are engineered to dramatically optimize memory performance, reduce latency, and alleviate the punishing data bottlenecks that plague modern large language model (LLM) infrastructures.
During a recent briefing with EE Times, Thad Omura, senior vice president of Astera Labs’ compute connectivity group, emphasized that memory has firmly cemented itself as the central infrastructure challenge of the current AI boom. As the industry races to keep pace with explosive advancements in GPU clustering, the old structural limitations have reemerged with a vengeance.
"All the bottlenecks around memory are back and driving new architectures, and we’re trying to deal with the memory constraints," Omura explained, highlighting the urgent need for innovative silicon that can bridge the widening gap between processor capability and memory bandwidth.

Expanding CXL Capabilities and Redeveloping Aging Hardware
At the heart of the company’s CPU-attached memory expansion and pooling strategy are the new Leo 2 E-Series and P-Series controllers, which continue to leverage the Compute Express Link (CXL) standard. These versatile devices support both modern DDR5 memory and repurposed DDR4 memory, providing data center operators with flexible configuration options. Specifically, the controllers support up to 4 terabytes (TB) of DDR5 capacity per controller, alongside up to 768 gigabytes (GB) of repurposed DDR4 capacity.
Beyond raw capacity expansion, Astera Labs has integrated sophisticated diagnostic and management features directly into the new Leo 2 family. The controllers now incorporate advanced memory testing, telemetry, predictive failure analysis, and automated repair capabilities. These tools are specifically designed to assist hyperscale data center operators in identifying potential hardware degradation early, allowing them to safely redeploy aging dual in-line memory modules (DIMMs) rather than prematurely discarding them.
This focus on hardware longevity arrives at a critical economic juncture for the industry. In the same media briefing, Ahmad Danesh, associate vice president of product management at Astera Labs, pointed out that DDR5 pricing has risen sharply as major memory manufacturers shift their manufacturing focus toward high-bandwidth memory (HBM) production to satisfy soaring GPU demand.
"There’s this big supply crunch and wafer supply crunch because of AI," Danesh noted, underscoring why operational efficiency and hardware recycling have become paramount for cloud providers.

The dynamic pooling and sharing features built into the Leo 2 family are engineered to directly mitigate this supply crunch by reducing stranded memory—idle capacity trapped inside underutilized servers—and allowing it to be dynamically allocated across the data center network as workloads demand.
"The expanded Leo family turns idle, stranded capacity into memory an AI agent can actually use, on hardware operators already own," Omura said. "That stranded DRAM is the industry’s next scale-up resource."
Unlocking Inference Performance with the Leo X-Series and Scorpio Switches
While CPU-attached memory expansion addresses general resource pooling, inference workloads present a different set of obstacles, particularly when key-value (KV) cache data threatens to overwhelm GPU high-bandwidth memory. To solve this specific problem, Astera Labs developed the Leo X-Series fabric-attached memory controller.
Rather than forcing critical cache data to route inefficiently through CPU memory, the Leo X-Series works in tandem with Astera Labs’ Scorpio fabric switches. This combination establishes a dedicated, high-speed memory tier attached directly to the PCIe fabric, bypassing traditional system bottlenecks.

The Scorpio fabric switches play a vital role in optimizing connectivity to AI accelerators while simultaneously improving performance per watt—a metric of mounting concern for data center operators grappling with escalating power consumption. Astera Labs announced its latest Scorpio fabric switch earlier this year, with Omura stressing that intelligent fabric design represents the fine line between peak computational efficiency and wasted resources.
The Scorpio X-Series 320-Lane Smart Fabric Switch reflects the company’s core mission of delivering purpose-built connectivity designed specifically for rack-scale AI architectures. Because modern data center footprints are increasingly measured in gigawatts, scalability and resilience have become critical success factors, especially as generative AI inference scales to unprecedented enterprise levels.
Danesh explained that the latest Scorpio switch was architected with massive AI clusters in mind, achieving lower latency by enabling single-hop communication between GPUs. To deliver these performance gains without inflating power budgets, Astera Labs integrated dedicated hardware acceleration engines directly into the switch silicon.
"Those engines are built to maximize the token economics further and get better performance per watt," Danesh said, noting that these hardware-accelerated and in-network compute engines can boost collective operations by up to two times. By offloading specific calculations from the GPU accelerator into the network silicon itself, the system spends less time transmitting data across cables and more time executing primary computations, resulting in substantially higher overall GPU utilization.
Real-World Benchmarks and the Shift Toward Active Interconnects

The real-world impact of these architectural improvements was demonstrated in benchmark testing conducted by Astera Labs. Utilizing a configuration of four Nvidia H200 GPUs running the Qwen2.5-32B model under demanding multi-turn agentic workloads, the company reported a striking 62% reduction in time-to-first-token (TTFT) and a 22% increase in tokens per second when compared against conventional CPU-DRAM-based KV cache offloading methods.
To extract maximum value from this hardware, Astera Labs pairs its silicon with Cosmos software, adding an intelligent orchestration layer to the infrastructure stack.
"At the scale at which things are being deployed, it’s that software that really provides a lot of value," Danesh observed.
This combination of advanced hardware accelerators and intelligent software illustrates a fundamental paradigm shift in modern data center design. Interconnects are no longer treated as passive copper plumbing; they are now active system components that directly dictate GPU utilization, service quality, resilience, and the economic return on constrained data center power budgets.
The broader semiconductor industry is rapidly moving in a similar direction, treating CXL as an indispensable tool for unlocking hidden memory capacity. Five years ago, Astera Labs helped catalyze this movement by releasing its Leo CXL Memory Accelerator Platform to help CPUs efficiently manage CXL-attached DRAM and persistent memory.

Other industry players are also aggressively targeting data movement to alleviate memory bottlenecks. South Korean startup Xcena recently unveiled a CXL Type 3 device combining up to 2 TB of DDR5, SSD-backed InfiniteMemory capacity, and more than 1,000 custom RISC-V cores to push compute directly into memory. Meanwhile, major infrastructure suppliers like Synopsys have updated their CXL intellectual property portfolios for the AI era, and Marvell has introduced a comprehensive suite of memory disaggregation products aimed at helping cloud providers optimize inference token efficiency.
As the ecosystem matures, innovations in memory controllers, fabric switches, and software orchestration are steadily transforming how data centers handle the relentless computational demands of next-generation artificial intelligence.
Leave a Reply