Processing data closer to memory is a time-tested concept in computer architecture, but South Korean startup Xcena believes it has engineered a vastly superior approach by leveraging the emerging Compute Express Link (CXL) protocol. During Hot Chips 2026, the company pulled back the curtain on its flagship MX1 device, a CXL Type 3 hardware accelerator designed to tackle the crippling memory bottlenecks that plague modern artificial intelligence and data-intensive workloads.
The MX1 represents a dense integration of hardware resources, combining up to 2 TB of high-speed DDR5 memory, SSD-backed InfiniteMemory capacity, and more than 1,000 custom RISC-V processor cores onto a single board. According to Xcena Chief Product Officer Harry Kim, the MX1 can deliver up to 4.7 times the throughput and an astonishing 18.7 times the energy efficiency of a traditional host CPU processing data over a standard CXL connection on selected data analytics benchmarks.
In a briefing with EE Times, Kim emphasized that the industry’s persistent memory bottleneck cannot be solved simply by stacking more raw memory capacity onto a system board. Instead, existing memory resources must be utilized with far greater efficiency. "If we can reduce the data movement, then it will be really helpful," Kim explained, pointing to the fundamental inefficiency of hauling raw data back and forth across system buses to be processed by distant host CPUs or GPUs.

Targeting Memory-Intensive AI and Database Workloads
The MX1 architecture is tailored specifically for memory-intensive AI training, inference, and database operations where only a minor subset of the stored data ultimately needs to be processed by the primary computing engines. In applications such as vector search, key-value (KV) cache operations, and complex database filtering, performing computation directly adjacent to the memory pool drastically reduces latency, power consumption, and interconnect bandwidth congestion.
"For the database example, we can reduce data movement by filtering out the old data you don’t need to see from the CPU side," Kim said. In the context of modern artificial intelligence workloads, this near-memory processing model allows systems to pull only the relevant, active slice of a KV cache rather than forcing the entire cache structure across the system bus. This capability is becoming increasingly critical as large language models scale to sizes that routinely saturate conventional memory bandwidth channels.
Bringing together disparate technologies like DDR5 memory, flash-based solid-state storage, and thousands of RISC-V cores required Xcena to resolve formidable software and memory management hurdles. The primary design goal was to ensure the MX1 remained frictionless and intuitive for systems engineers to program. "We want to keep the memory model as similar as what they want to do on the CPU side," Kim noted.

One of the most complex engineering challenges involved developing a unified virtual memory system capable of seamlessly maintaining complex address mappings between the host processor’s memory, the accelerator’s local device memory, and the extended SSD-backed capacity layers.
A Specialized Near-Memory Accelerator
Xcena positioned the MX1 above all as a robust memory expander, yet it fundamentally differentiates itself from competing CXL memory expansion products currently on the market by integrating a full-featured memory controller alongside specialized near-memory processing hardware. While alternative computational CXL devices rely on general-purpose processors, Xcena’s architecture is explicitly optimized for massively parallel data processing that fully saturates the internal bandwidth of the hardware.
The computational backbone of the device relies on a many-core RISC-V architecture. Kim noted that this choice was heavily influenced by the maturity of the open-source RISC-V ecosystem. To make the hardware accessible to software developers, Xcena modified an LLVM-based compiler to support custom processing instructions and developed a CUDA-like software development kit (SDK), allowing developers to write applications in familiar languages such as C, C++, and Rust.

Because managing thousands of individual processor cores concurrently remains a notoriously difficult software engineering problem, Xcena engineered a proprietary framework inspired by distributed database systems. This framework simplifies application programming and workload orchestration, shielding developers from the raw complexity of the underlying hardware topology.
Xcena is currently sampling its MX1 prototypes to major memory manufacturers, CPU vendors, and hyperscale cloud providers. The company has laid out an aggressive commercial timeline, targeting mass production by the end of 2026 and anticipating initial commercial customer revenue to materialize in 2027.
Navigating the CXL Ecosystem and Industry Adoption
A crucial dependency for Xcena’s commercial strategy is the broader industry transition toward CXL 3 running natively over high-speed PCIe 6.0 interfaces. Although market adoption of the CXL standard has arguably progressed more slowly than many early industry forecasts anticipated despite rapid protocol evolution, Kim remains optimistic. He argued that the technology itself is not at fault, pointing out that all major CPU and GPU manufacturers, including market leader Nvidia, have already integrated CXL physical and logical layers directly into their silicon roadmaps.

"CXL is ready," Kim said, emphasizing that the primary remaining hurdle is proving its tangible value to enterprise buyers. Successful, widespread adoption demands a mature supporting ecosystem. "We need to prove CXL is good for everyone for AI infrastructure with CXL 3."
Independent industry analysts view Xcena’s hardware strategy through the lens of established storage paradigms. In a separate briefing with EE Times, Jim Handy, principal analyst at Objective Analysis, characterized the MX1 architecture as an intriguing combination of memory expansion and computational storage. "It’s an interesting product," Handy remarked.
A recent market research report from Objective Analysis titled CXL Market Now Taking Shape highlights how CXL is rapidly gaining traction within the hyperscale computing community, positioning the broader technology sector for significant financial and operational growth.

According to Handy, while memory expansion remains the most commercially compelling capability of CXL hardware today, innovative enterprises are also finding value in repurposing older memory infrastructure—such as older DDR4 modules rather than purchasing costly new DDR5 stock at current market prices. He highlighted that Meta, the parent company of Facebook, has successfully deployed CXL-based memory expansion across millions of servers, actively recycling DDR4 modules pulled from decommissioned hardware to extend server utility rather than retiring them prematurely.
Furthermore, Handy noted that the advanced memory pooling capabilities enabled by CXL have yet to be fully exploited across the industry, despite offering immense theoretical promise. "There’s probably fire under the feet of the people who are developing pooling," Handy observed. "It improves the percent utilization that you get out of the memory chips you’ve already got, and if memory is really expensive, then you want to make sure that you’re not having any of it sitting around idle."
Leave a Reply