As reported by QbitAI, Google TPUs run Kimi (Moonshot AI) 57% faster than NVIDIA GPUs — with DeepSeek's open-source inference framework underneath. The significance is not the number itself but the two plates it moves at once: the viability of non-NVIDIA silicon running top models, and the de-facto standard status of domestic open-source inference stacks.
Inference frameworks as the new middleware
The first half of the LLM race was about training compute; the second half's key variable is inference efficiency. DeepSeek's open-sourced inference stack (FlashMLA, DeepEP, 3FS and successors) has been adapted worldwide — AMD, Ascend, domestic GPUs — and "DeepSeek framework + non-NVIDIA hardware" keeps appearing on performance lists. Now TPU runs Kimi on it: the abstraction layer is general enough that cross-architecture migration costs have collapsed.
Three shocks to the compute landscape
- Hardware lock-in loosens: as open frameworks make TPUs and domestic cards efficient hosts for mainstream models, the CUDA moat gets bypassed more often;
- Inference cost curves steepen: framework-level throughput gains translate directly into API price cuts;
- Supply-chain autonomy: for Chinese government and enterprise, "open framework + domestic silicon" lifts the performance ceiling of private deployment.
NineZenith's lightweight deployment on domestic AI chips rides exactly this trend: when inference middleware levels the hardware field, competition returns to the model itself.
(Facts aggregated from public reporting; measurements as originally reported)