Deep|DeepSeek V4: The Inflection Point for Large-Scale NAND-Based KV Cache
In our previous article we discussed DeepSeek V4’s architectural customization on non-NVIDIA hardware and the first round of API price cuts at 75% off. This article focuses on V4’s second round of cuts: DeepSeek separately took the input cache-hit tier further down to 1/10 of list, stacked on top of the 75% off from the previous round, with the floor at ¥0.025 per million tokens. This widens the cache hit / cache miss spread from 1/12 to 1/120 (cache hit ¥0.025 vs. cache miss ¥3). DeepSeek V4’s real-world cache hit rate in agent settings has reached 95%+, and based on our research, DeepSeek’s current SSD configuration and utilization have stepped up materially versus before. Behind this is V4 compressing KV cache size to 10% of V3.2’s, plus DeepSeek’s accumulated engineering work on SSD-based KV cache, which together migrate KV cache from expensive, capacity-limited DRAM / HBM onto larger and cheaper SSD at scale. We believe DeepSeek V4’s cache-hit repricing implies upside for SSD, with NAND demand set to grow exponentially.
Deep|DeepSeek V4: The First Model Custom-Built for Non-NVIDIA Chips, Optimized for Cost
The most important thing about DeepSeek V4 is not just the capability improvements, and not just the longer context support. What stands out to us is that it appears to have been deeply customized for non-NVIDIA AI accelerators and ASICs — most notably Huawei Ascend. V4’s architectural choices are clearly suited to running on non-NVIDIA hardware and cluster environments.


