Loading...
Loading...
Google has open-sourced their TPU Raiden inference optimization library. This is the equivalent layer of the stack to NVIDIA NIXL, where it provides KVCache transfer between prefill & decode instances & has primitives for KVCache offloading movements! It is great to see Google open-source & externalize more and more of their TPU stack!
Source:https://x.com/SemiAnalysis_/status/2086241160243118556
Impact Score