Loading...
Loading...
reply Why: AI inference is a (high spede) storage problem. A 405B model leaves ~0.5 MB of KV cache behind per token of context. 32 long-context sessions = 2 TB of state vs ~1.5 TB of high bandwidth memory (HBM) on an 8-GPU node. It spills to flash. SanDisk sells the flash, and sizes this one workload at 75-100 exabytes of 2027 demand.
Source:https://x.com/SynapseProtocol/status/2089603240639840343
Impact Score