AI September 17, 2026 mixed ⇧ 143 pts across 1 thread

Geopolitical AI Infrastructure Is Splitting in Two

GLM published a post about building its own production inference infrastructure on a cluster of more than 100,000 Chinese-made AI accelerators. All production inference for GLM-5.3-Flash runs on this system, no Nvidia hardware involved. The engineering post is technically detailed and credible. The comment thread immediately raised the distillation attack allegation: GLM was accused of routing paying customers' requests through Anthropic's Opus model to steal training signal.

Set the ethics aside for a moment and look at the infrastructure fact alone. A Chinese lab has built a complete, production-scale inference stack on domestic hardware. That is not a prototype. That is operational at scale.

This is the clearest signal yet that the AI hardware and software stack is splitting along geopolitical lines. Nvidia is not available to Chinese labs at scale due to export controls, so they built around it. The question of whether Chinese-made accelerators can match H100 performance is becoming less relevant than the fact that a full production system exists and is running.


So what?

Founders building infrastructure or tooling that assumes Nvidia dominance in AI compute should start planning for a world with at least two parallel hardware ecosystems. If your product works only on CUDA, you are building for half the market, and possibly the half with slower long-term growth if Chinese labs continue scaling on domestic silicon.

Read these