News

/
News

Micron is exploring GPU-accelerated near-memory NAND flash memory for running larger-scale large language models.

It is reported that Micron is exploring a plan to develop high durability NAND flash memory modules: placing the flash memory closer to the GPU instead 

of in remote storage pools and through multiple-layer protocol forwarding access. The company is researching the migration of NAND flash memory to the

 GPU side for processing those business loads that do not require the low latency and high bandwidth capabilities of traditional HBM or ordinary DRAM.


This architecture is currently referred to as GPU Near-Cache NAND (near-GPU NAND), aiming for a compromise between storage density and durability. 

Under this concept, the NAND flash memory has a lower storage density than traditional TLC or QLC NAND, but focuses on optimizing I/O speed, bandwidth,

 and read latency.


For example, this solution enables the GPU to directly access hundreds of GB storage pools located inside the GPU package or on the PCB board, similar to 

DRAM such as HBM, GDDR7, and LPDDR6. This storage level is between GPU video memory and ordinary system storage, acting as a high-speed cache, and

 is highly suitable for large-scale language model (LLM) inference tasks.


After adopting this solution, large language models will no longer be constrained by video memory capacity and can run on systems with fewer GPUs, relying

 on a multi-level storage/memory architecture to meet the capacity requirements for inference. Moreover, this type of NAND flash memory has a cost much 

lower than HBM and ordinary DRAM, and can be flexibly configured in terms of capacity according to system specifications. The following illustration is a 

presentation slide of the HBF technology from Hynix and SanDisk.


Not only Micron is researching this technology. For instance, SK Hynix and SanDisk have already launched the High Bandwidth Flash (HBF) concept, by 

moving NAND closer to the GPU and using more durable NAND media, expanding the HBM memory space of the GPU. All companies entering this direction

 need to overcome the shortcomings of bandwidth and durability to make HBF or GPU Near-Cache NAND a valuable supplementary solution to HBM 

memory.