Updated
Updated · Forbes · Aug 21
SK hynix, Sandisk Unveil 512GB HBF Spec With Up to 3.0TB/s for AI Inference
Updated
Updated · Forbes · Aug 21

SK hynix, Sandisk Unveil 512GB HBF Spec With Up to 3.0TB/s for AI Inference

2 articles · Updated · Forbes · Aug 21

Summary

  • SK hynix and Sandisk released the first High Bandwidth Flash specification at FMS 2026, setting initial HBF packages at up to 512GB with three bandwidth grades spanning about 0.4TB/s to 3.0TB/s.
  • The format targets AI inference workloads that need fast access to persistent models, KV cache and agent state, creating a new memory tier between HBM and pooled memory to improve xPU performance and cost efficiency.
  • UCIe will link HBF packages to processors, and SK hynix outlined a roadmap from the 0.7 standard announced now to a full specification in early 2027, samples in early 2028 and production after that.
  • Sandisk said HBF can supplement or replace HBM in some AI systems; in one comparison, a 4TB HBF-only setup with 4 GPUs delivered nearly the same tokens per second as 192GB HBM with 8 GPUs, doubling GPU efficiency.

Insights

As tech giants back High Bandwidth Flash, is the era of relying solely on expensive HBM for AI dominance finally over?
With AI memory tiers expanding, will the hidden software complexity of managing HBF outweigh its massive cost benefits?

High Bandwidth Flash (HBF) Emerges: Open Standard Redefines AI Memory Hierarchy and Slashes Inference Costs by 2026

Overview

Modern AI inference faces a severe memory wall, as fast but limited HBM and slow, high-capacity SSDs cannot keep up with the demands of large models. In response, Sandisk and SK hynix released the open High Bandwidth Flash (HBF) specification in 2026, creating a new memory tier that bridges this gap. HBF leverages the low cost and high capacity of NAND flash, delivering 8 to 16 times the capacity of HBM at similar cost, while using open UCIe interfaces to avoid vendor lock-in. Early hybrid architectures combining HBM and HBF show dramatic improvements in throughput and efficiency, enabling complex workloads to run on far fewer GPUs and reducing both costs and power consumption.

...