Flash innovations in the AI era: From Storage to an Active Memory Tier
AI is redefining infrastructure requirements as model sizes, token volumes, context windows, multimodal inputs, and reasoning workloads grow faster than improvements in accelerator memory capacity, bandwidth, power efficiency, and cost. As a result, the next performance bottleneck is increasingly shaped by how efficiently data, model parameters, activations, and inference state can be stored, moved, and reused.
Flash already plays a critical role across the AI data lifecycle, from raw data archives, preparation, checkpointing, and fine-tuning to vector databases, retrieval-augmented generation, and generated-content storage. Its role is now moving closer to compute, supporting model loading, data staging, GPU feeding, parameter access, and persistent KV-cache management.
This keynote will examine how storage-centric and compute-centric flash architectures, across both direct-attached and network-attached implementations, can extend the AI memory hierarchy. It will also explore high-bandwidth flash, intelligent data placement, near-data processing, and deeper silicon, system, and software co-design.
The central thesis is that flash is no longer only where AI data resides. It can evolve into an active data and memory platform that reduces re-computation and data movement, improves accelerator utilization, and enables scalable, power-efficient AI inference.
Key Technologies Covered
- AI pipelines and Innovation ideas in Enterprise SSD usage model in each stages
- Storage-centric and compute-centric flash architectures
- HBF with intelligent data placement