Workers AI Doubles Kimi Context to 1.37 Million Tokens, Cutting Cost 30%
Updated
Updated · The Cloudflare Blog · Aug 3
Workers AI Doubles Kimi Context to 1.37 Million Tokens, Cutting Cost 30%
1 articles · Updated · The Cloudflare Blog · Aug 3
Summary
Workers AI said new memory optimizations let it serve Moonshot’s Kimi and Z.ai’s GLM more efficiently, with no measurable accuracy loss across its benchmark suite.
FP8 quantization halves Kimi’s KV cache from 16-bit precision, lifting in-memory context from about 686,000 tokens to 1.37 million and pushing peak decode throughput 41% above BF16 while lowering cost per token by roughly 30%.
GLM 5.2 weight compression to INT4 shrank the checkpoint to 421 GB from 705 GB and cut per-GPU memory to about 52 GB from 88 GB, boosting decode speed by 16% to 55% depending on concurrency.
Shared-cache integrity checks were added as denser packing put hundreds of requests on the same GPU; Cloudflare said the protection costs under 1% in throughput and tail latency and aborts mismatched requests.
The company said it is expanding FP8 caches across more of its fleet, testing NVFP4 weights on Blackwell GPUs, and aiming to leave integrity checks enabled broadly at negligible cost.