Memory has become one of the main constraints on modern AI and cloud infrastructure. Processors and accelerators continue to become faster, yet adding enough memory to feed them efficiently is increasingly difficult and expensive. Compute Express Link, better known as CXL, addresses this problem by providing a coherent connection between processors, accelerators and external memory. CXL 4.0, released by the CXL Consortium in November 2025, raises the maximum data rate from 64 GT/s to 128 GT/s while retaining the memory pooling and sharing capabilities developed through earlier CXL generations. By 2026, that combination is particularly relevant to AI inference, large cloud servers and rack-scale computing. It does not make conventional DDR memory or high-bandwidth memory obsolete, nor does it instantly turn every rack into one enormous memory system. What it does provide is a practical route towards treating memory as a more flexible resource rather than a fixed quantity permanently attached to one processor.
CXL was created because the traditional relationship between processors and memory was becoming too restrictive for data-centre workloads. A conventional server receives a fixed amount of local memory through its CPU memory channels. That arrangement works well when workloads are predictable, but cloud and AI systems rarely remain predictable for long. One machine may need several terabytes of memory for a temporary workload while another machine in the same rack has large amounts of unused DRAM. Installing enough local memory in every server to cover its possible peak demand leaves expensive capacity idle for much of the time. CXL changes this relationship by allowing compatible memory devices to sit outside the processor’s normal DIMM arrangement while still being accessed with memory-style load and store operations. Earlier CXL releases established expansion, switching, pooling and sharing. CXL 4.0 concentrates heavily on increasing the amount of data that can travel through those connections.
The headline change is the move to 128 GT/s, twice the 64 GT/s maximum data rate of CXL 3.x. CXL 4.0 achieves this by building on the physical signalling defined for PCI Express 7.0, whose final 1.0 specification was released in June 2025. GT/s means giga-transfers per second and should not be confused with gigabytes per second. The figure describes the signalling rate of each lane, while real usable bandwidth depends on factors including lane count, protocol overhead and the device configuration. PCI Express 7.0, for comparison, can provide up to 512 GB/s of bidirectional bandwidth with a sixteen-lane connection. CXL uses the same high-speed physical foundation but adds the coherency and memory behaviour needed for processors, accelerators and memory devices to work together. The result is substantially more link capacity without forcing software to treat attached memory like ordinary storage or a conventional network resource.
This higher data rate matters because memory expansion is only useful when the path to that memory is fast enough for the workload. Local DRAM remains the preferred location for latency-sensitive data, while HBM remains essential when GPUs and other accelerators need extremely high bandwidth close to their compute engines. CXL memory sits in a different part of the hierarchy. It can provide considerably more capacity than can economically be installed next to every processor or accelerator, with better access characteristics than moving the same information to SSD storage. For AI infrastructure, this creates another usable memory tier between scarce high-speed local memory and much slower storage. For cloud operators, it also provides a way to match memory capacity more closely to changing workloads instead of buying every server for its theoretical worst case. The value of CXL 4.0 therefore comes as much from flexibility as from its raw transfer rate.
Doubling a link from 64 GT/s to 128 GT/s does not mean that an application automatically runs twice as fast. Many applications are limited by processor performance, accelerator throughput, software behaviour or memory latency rather than by the CXL link itself. The extra bandwidth becomes important when large amounts of data need to move between processors, accelerators and expanded memory at the same time. A server performing large-model inference, analytics or an in-memory database workload may have many workers reading and writing data concurrently. In those situations, a slower interconnect can become a shared bottleneck even when plenty of memory capacity is available behind it. CXL 4.0 gives system designers more headroom for such traffic. The CXL Consortium also retained the fixed-size transfer structure and error protection developed for the 64 GT/s generation, allowing the higher rate to be introduced without simply accepting a proportional increase in protocol latency.
CXL 4.0 also introduces bundled ports. In straightforward terms, several physical CXL connections can be treated as one logical connection where the device and host design support it. That is useful when a single link cannot provide enough bandwidth for a high-performance accelerator or another demanding component. Instead of forcing the rest of the system to view each connection as a completely separate device path, bundled ports provide a defined method for aggregating them. The specification also supports native x2 links, which can help designers connect a larger number of devices when maximum bandwidth is not required on every connection, and it allows as many as four retimers to extend channel reach. These changes are relevant to dense servers and rack designs where components cannot always sit immediately beside the processor. They give designers more options for balancing link width, device count, distance and bandwidth.
Reliability is equally important when memory moves beyond the motherboard. A failed conventional DIMM generally affects one server, whereas pooled or shared memory may support several machines or important workloads. CXL 4.0 therefore includes additional memory reliability, availability and serviceability capabilities intended to improve error visibility and maintenance. This does not eliminate failures, and operators still need redundancy, monitoring and sensible workload placement. It does make high-capacity CXL memory easier to manage as infrastructure rather than as an unusual peripheral. Full backward compatibility is another practical consideration. CXL 4.0 systems are designed to work with earlier CXL generations where supported, so organisations do not have to replace an entire CXL estate in one step. That matters in 2026 because the industry is simultaneously deploying CXL 2.0 and 3.x equipment while beginning design and validation work for 128 GT/s CXL 4.0 hardware.
Memory pooling predates CXL 4.0. The feature was introduced with CXL 2.0, alongside switching and a standardised Fabric Manager model. CXL 3.0 then expanded the concept with larger fabrics, multi-level switching and coherent memory sharing. The distinction is important. Memory pooling means that a quantity of CXL-attached memory can be allocated to different hosts as demand changes. A portion assigned to one server can later be released and assigned elsewhere. Memory sharing goes further by allowing supported memory regions to be available to more than one host while coherency mechanisms keep their view of the data consistent. CXL 4.0 retains these capabilities while giving the fabric substantially more bandwidth. The practical result is not a new form of pooling invented in 2026, but a faster interconnect that makes existing pooling and sharing models more attractive for larger systems and heavier workloads.
A simple example shows why this matters. Imagine several cloud servers, each with enough local DRAM for ordinary operation, connected to an additional bank of CXL memory. One server may suddenly need hundreds of gigabytes more memory for analytics, an AI inference service or a large database. Instead of requiring that capacity to have been installed permanently inside that particular machine, part of the CXL pool can be assigned to it. When demand falls, the capacity can be returned and made available elsewhere. A Fabric Manager coordinates the relevant resources, while operating-system and orchestration software determine how the extra memory should be used. The exact implementation differs between vendors, and there is still a cost in latency compared with local DRAM, but the underlying economic idea is straightforward: memory that would otherwise sit unused behind one CPU can become capacity available to other workloads.
This approach targets the problem known as stranded memory. Cloud servers are often configured for peak requirements rather than average requirements because running out of memory can severely degrade a workload or prevent it from running altogether. As a result, a server can have free DRAM at the same time that a neighbouring system lacks capacity. Pooling reduces the need for every machine to carry the same large safety margin. It can also make upgrades less disruptive because additional memory can be installed in a shared CXL subsystem rather than only by populating local processor memory channels. The economics still depend on workload behaviour, switch costs, power consumption, software support and the performance difference between local and CXL-attached memory. CXL should therefore be treated as another memory tier, not as evidence that local DRAM is no longer necessary. The most efficient designs generally use several memory types for different jobs.
AI inference is one of the clearest reasons for interest in CXL memory during 2026. Serving large language models involves more than storing model weights. Each active request can also require temporary state, including the key-value cache used by transformer models to avoid recalculating previous attention information. As context windows become longer and servers handle more simultaneous users or AI agents, that cache can occupy a substantial amount of memory. Keeping every byte in expensive accelerator HBM can restrict the number of sessions that a server handles even when the GPU still has unused compute capacity. Moving all excess data to SSDs, on the other hand, can introduce much larger access delays. CXL memory offers an intermediate tier where selected information can remain memory-addressable without consuming the accelerator’s most valuable local capacity.
This does not mean that an AI operator should simply move an entire model from HBM into CXL memory. The hottest data still benefits from being as close as possible to the accelerator, and the bandwidth available from HBM is considerably higher than that of an external CXL memory tier. A more realistic design keeps latency-sensitive model data in HBM, retains appropriate working data in local system DRAM and uses CXL capacity for information that needs to remain readily available but does not require maximum local bandwidth at every moment. KV-cache tiering and offload are prominent examples discussed by the CXL industry in 2026. When software can identify which data belongs in each tier, the server may support larger contexts or more concurrent inference requests without adding an equivalent amount of costly accelerator memory. The benefit is therefore capacity efficiency rather than a claim that CXL is inherently faster than HBM.
The same principle applies beyond large language models. Recommendation systems can maintain very large embedding tables, data-processing jobs can work with datasets larger than local DRAM, and CPU-based inference may require more memory than a conventional server configuration offers economically. CXL can extend the usable memory footprint for these workloads while preserving standard memory access semantics. Pooling adds another level of flexibility because capacity can follow demand rather than being permanently dedicated to one machine. This is particularly attractive in shared AI infrastructure where different services peak at different times. A batch job may need substantial memory overnight, while an inference service may need the same capacity during business hours. Dynamic allocation cannot remove every operational constraint, and moving capacity between hosts still requires management by the operating system and infrastructure software, but it gives operators more choices than a fixed-memory server design.

The long-term change created by CXL is architectural rather than simply numerical. Traditional servers are designed around resources that belong to one machine: its processors, DIMMs and accelerators are installed for that server and remain there even when they are lightly used. CXL allows some of those resources, particularly memory, to be treated more independently. A rack can contain conventional servers alongside CXL switches, memory expansion devices and pooled memory systems, with capacity assigned according to workload requirements. CXL 3.x already provides much of the fabric behaviour required for this model. CXL 4.0 adds substantially more link bandwidth and new connectivity options, which are important as the number of devices and amount of traffic grow. This does not turn a rack into a single computer, but it weakens the assumption that every useful byte of memory must be physically installed beside the CPU that will eventually use it.
For cloud providers, the most obvious financial benefit is potentially better memory utilisation. DRAM represents a significant part of the cost and power budget of memory-heavy servers. If each server is equipped for rare demand spikes, much of that investment may produce little useful work during normal operation. A shared memory pool can allow operators to purchase capacity for aggregate demand rather than giving every machine enough memory for its individual theoretical maximum. The saving is not automatic. CXL switches, controllers, memory enclosures and management software also cost money and consume power, while workloads that require consistently low memory latency may still justify large amounts of local DRAM. The strongest business case therefore appears where memory demand varies significantly between hosts, where capacity limits are more important than minimum latency, or where extra memory allows expensive processors and accelerators to remain productive instead of waiting for data to be moved from storage.
Operational flexibility may prove just as important as hardware cost. A cloud service can change considerably during the lifetime of a server. AI models become larger, databases grow and customers move between instance types. With a fixed-memory machine, an increase in capacity may require moving the workload or replacing hardware. CXL expansion and pooling create the possibility of changing the available memory without rebuilding the compute node. They can also support more specialised server configurations: some machines may contain large local memory capacity, while others depend more heavily on shared capacity for workloads that tolerate the additional latency. Reliability and security remain essential in this model. Memory must be reassigned safely, failures must be isolated, and information belonging to one workload must not become visible to another. As CXL moves towards multi-host and rack-scale use, these management requirements become part of the design rather than optional extras.
The most important reality in 2026 is that a published CXL 4.0 specification does not mean 128 GT/s CXL hardware is already universal in production data centres. Current commercial and demonstration systems span several generations. Samsung’s MD220 CXL memory module, for example, uses CXL 2.0 over PCIe 5.0 and is available in 128 GB and 256 GB capacities, while Samsung’s rack-oriented CMM-B memory pooling system also uses CXL 1.1 and 2.0 technology. SK hynix was still showing CXL 2.0-class CMM-DDR5 solutions during 2026, and CXL Consortium events during the year included live CXL 3.2 links and controllers. At the same time, semiconductor design suppliers are already providing CXL 4.0 controller, security and verification technology for chips intended to operate at 128 GT/s. In other words, the transition is underway, but today’s mature products and tomorrow’s 4.0 designs coexist.
That gradual adoption is normal for a server interconnect. A specification must be followed by controller designs, switches, retimers, processors, memory devices, firmware, operating-system support, compliance testing and multi-vendor interoperability work before large operators can deploy it confidently. CXL 4.0 has an advantage because it continues the architecture of earlier CXL releases and remains backward compatible, allowing manufacturers to build on existing software and design experience. The first reason to adopt it will not necessarily be a need for the number 128 GT/s itself. Operators are more likely to care about whether the additional bandwidth allows more memory devices, more simultaneous traffic or more accelerator activity without creating a bottleneck. As CXL fabrics grow from simple memory expansion towards pooled and shared rack-level resources, the extra bandwidth becomes increasingly useful because more workloads compete for the same interconnect capacity.
For AI and cloud infrastructure, CXL 4.0 should therefore be viewed as part of a broader change in memory design rather than a single-generation speed upgrade. Local DDR memory will continue to provide relatively fast CPU memory, HBM will continue to serve accelerators that require extreme bandwidth, and SSDs will continue to provide economical persistent capacity. CXL adds a flexible layer between those resources, allowing servers to reach larger pools of coherent memory and allowing capacity to be assigned with fewer physical constraints. The 128 GT/s generation gives that layer more room to scale. In the near term, many production systems will still use CXL 2.0 or 3.x while new silicon moves towards CXL 4.0. Over the following server generations, the more significant change is likely to be the shift from asking how much memory is installed in one server to asking how much suitable memory can be made available to a workload when it actually needs it.