XCENA Unveils MX1: Advancing Computational CXL Memory for AI Infrastructure

XCENA, a leader in semiconductor innovation, has introduced the architecture and measured performance of its MX1 computational CXL memory solution at the Hot Chips 2026 Memory session. Presented by Chief Product Officer Harry Kim, the session highlighted how MX1 addresses the escalating memory demands of AI infrastructure by combining large-scale memory expansion, SSD-backed capacity, and programmable near-memory computing within a single CXL Type 3 device.

Addressing AI Memory Bottlenecks with MX1

As artificial intelligence models continue to scale, system performance is increasingly limited by memory capacity, data movement, and the cost of high-bandwidth memory, rather than compute power alone. MX1 is engineered to complement CPUs, GPUs, and AI accelerators by expanding memory resources and offloading memory-bound operations such as vector search, KV-cache retrieval, and data preprocessing. This allows host processors to concentrate on compute-intensive inference tasks.

Key Architectural Features of MX1

  • Memory Expansion: Supports up to 2 TB of DDR5 memory across four channels, connected via CXL 3.2 over PCIe 6.0, enabling significant memory scaling for demanding AI workloads.
  • SSD-Backed Capacity: XCENA’s InfiniteMemory technology exposes SSD storage as byte-addressable CXL memory, with transparent caching of frequently accessed 64 KB pages in DRAM for optimal performance.
  • Near-Memory Computing: Over 1,000 custom RISC-V cores, organized into Memory Acceleration Units, execute highly parallel, memory-intensive workloads close to the data. Integrated vector engines deliver approximately 3 TFLOPS of FP32/FP16 dot-product throughput, accelerating tasks such as vector search and KV-cache scoring.

Measured Performance on Data-Processing Workloads

XCENA evaluated MX1 on six representative data-processing kernels commonly used in analytics and AI pipelines: compression, decompression, Parquet decoding, less-than filtering, LIKE filtering, and aggregation.

Compared to a host CPU processing data over CXL, a single MX1 device achieved up to 4.7x higher throughput and 18.7x greater energy efficiency. When compared to the same host CPU accessing local DDR5 memory, MX1 delivered up to 2.0x higher throughput and 6.2x greater energy efficiency. These results underscore the benefits of executing parallel, memory-bound workloads closer to the data, reducing unnecessary data movement and allowing CPU resources to focus on system services and orchestration.

Developer Ecosystem and Software Integration

MX1 is programmable using standard C/C++ or Rust languages through XCENA’s LLVM-based software toolchain. The PXL runtime manages workload scheduling and synchronization across the device’s RISC-V cores, providing a shared virtual address space to simplify memory management for complex data structures and multi-application environments.

At the framework level, XCENA’s XFLARE analytics library integrates with SQL engines and FAISS-based vector search. The software stack is being extended to support widely adopted AI and data frameworks, including Apache Arrow, PyTorch, and vector databases, with additional SDK integrations planned.

Production Timeline

XCENA anticipates beginning mass production of MX1 by the end of 2026, with initial customer deployments targeted for 2027.