Loading...
Loading...
quote: My bro @zephyr_z9 called it Was gonna write an 10000-word easay today to call out $SNDK marketing BS but got distracted by Funny AF with Kevin Hart Admittedly the investor day has better comedy material With this example, it's reasonable to assume that Sandisk sized up KV cache TAM based on BF/FP16 too so divide their numbers by 2x (or 4x even) | They are deliberately misrepresenting HBM performance For some reason, they fix the bandwidth of both HBM & HBF stacks at 1.6TB/s (12.8/8) Secondly, nobody serves models on bf16 anymore Most models are served on fp4 or fp8 So a Qwen 480B should occupy 240GB to 480GB depending on the quantization They also fix the HBM capacity per GPU to 192GB Even though a 16Hi HBM4E can achieve 512GB with 8 stacks and have a bandwidth of 32TB/s (3x higher than what they report here)
Impact Score