Loading...
Loading...
Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node — ROCm Blogs.
Explore how one 8-GPU AMD MI355X node serves Kimi Linear from 1K to 64M tokens under vLLM—and the TTFT and decode throughput behind the run.
Source:https://rocm.blogs.amd.com/artificial-intelligence/long-context-serving/README.html
Impact Score