Loading...
Loading...
GLM-5.3 on KernelBench-Mega. Kimi-Linear Decode for RTX PRO 6000 at 21.4x the PyTorch baseline. GLM-5.2 was 11.1x and failed the single-launch gate. One CUDA `load_inline` launch. Persistent 512-thread CTAs. 30 grid barriers. Int4 dequant stays inside the GEMV. MLA is absorbed so the fat KV is not materialized. kernelbench.com/mega solution file below!
Impact Score