Hanami

Loading...

reply Why: AI inference is a (high spede) storage problem. A 405B model leaves ~0.5 MB of KV cache behind per token of | Hanami