Samsung Electronics shares reportedly fell 3.5% and SK Hynix 2.2% in South Korea on September 11 after reports that DeepSeek’s AI models require less high-bandwidth memory — a real market jolt that does not prove global HBM demand has permanently cracked, NeoTeo reported.
DeepSeek’s Multi-Head Latent Attention (MLA) compresses key/value attention state in the KV cache, and Engram describes conditional memory that can sit in an offloaded hierarchy. Secondary accounts attributed figures such as DeepSeek V4 using about 10% of DeepSeek V3.2’s KV-cache requirement at a one-million-token context and 27% of its single-token inference FLOPs — specific workload metrics, not measurements of HBM shipments or chip revenue.
Model weights, activations, multi-accelerator communication, and total request volume still drive memory needs. Cheaper inference can even expand total demand if providers run more queries. NeoTeo’s takeaway: DeepSeek may ease accelerator-memory pressure in some inference cases, but the September 11 selloff is not evidence that HBM demand has permanently fallen.
Sources: NeoTeo


Leave a Reply