Skip to content

Marvell Expands AI Memory and Storage Infrastructure with Bravera SC6

Networking now determines how effectively memory reaches GPUs

Introduction

Memory has become a critical constraint in AI infrastructure. As models scale, solutions involve a mixture of experts, and context windows expand, the way memory is used must change. If we don’t change, GPUs stall waiting for data rather than computing. Networking now determines how effectively memory reaches those GPUs. CXL-based architectures solve this by treating memory as a networked resource that scales independently of compute. Marvell’s latest portfolio announcements at FMS 2026 demonstrate exactly how this shift works in practice for hyperscalers and cloud providers.

The Importance of Networking for Memory

Traditional server-attached memory creates hard limits. HBM sits close to the GPU but costs too much to expand in proportion to model size. Local DRAM and NVMe introduce latency that kills token throughput. Networking changes the equation. CXL turns the PCIe physical layer into a coherent memory fabric. Memory controllers and switches move capacity and bandwidth across servers and racks with near-local access times. This disaggregation keeps GPUs fed and utilization high. Without strong networking between memory tiers, every additional parameter or longer context window increases idle time and raises cost per token.

CXL and Memory Controllers Unlock Better AI Performance in GPUs

CXL memory controllers and switches deliver the bandwidth and low latency GPUs demand. Marvell’s Structera X devices expand memory capacity and enable pooling across multiple hosts. The new Structera S CXL switches add rack-scale pooling so operators allocate memory dynamically to the GPUs that need it. Combined with Alaska P retimers, these solutions keep data movement efficient. The result is fewer stalls, higher FLOPs utilization, and faster job completion. Cloud operators already attach these controllers to both CPUs and XPUs. Performance gains compound because the network fabric itself becomes part of the memory hierarchy rather than a bottleneck outside it.

CXL and Memory Create Larger Context Windows

Longer context windows and growing KV cache drive the need for terabytes of fast memory. HBM alone cannot scale economically. CXL memory expansion and pooling let operators keep warm KV cache closer to the accelerators. Photonic Fabrics will extend this further by creating an optical shared memory tier across multiple racks. The platform supports up to 32 TB of warm KV cache offload with high bandwidth and near NUMA latency. Token throughput rises two to three times inside existing power and space envelopes. Larger context windows become practical because the memory network scales with the model instead of forcing operators to overprovision expensive on-package memory.

Cloud Providers Architect for Future GPU Designs

Hyperscalers already design around these principles. They treat memory, storage, and compute as independent resources connected by high-performance fabrics. Memory controllers doubled performance over the previous generation and move more KV cache from HBM to high-capacity storage while preserving endurance. New designs give providers sourcing flexibility. Photonic Fabric and Structera solutions prepare the infrastructure for next-generation GPUs that will demand even larger shared memory pools. Operators build the memory network in conjunction with their GPU/XPU roadmap, so future GPU generations plug in and immediately deliver higher utilization and lower cost per token.

Conclusion

Networking for memory is no longer optional. CXL controllers, switches, and optical fabrics remove the memory wall that limits GPU performance and context length. Cloud providers that architect this way today position themselves for the next wave of agentic AI workloads. Customers that treat memory as a networked resource will extract more value from every GPU they deploy.