Infrastructure September 4, 2026 bullish ⇧ 601 pts across 1 thread

Cerebras runs Qwen 3 at 1,500 tokens per second

Cerebras is now hosting Qwen 3 27B at 1,500 tokens per second, which commenters called genuinely difficult to keep up with. This is not a marginal speed improvement; it's fast enough that the bottleneck shifts from generation to comprehension. One commenter noted they used the product with GLM4.7 and found it fun, but only as good as the model being hosted.

The broader pattern here is that inference speed is becoming a competitive axis independent of model quality. Cerebras is making a bet that raw token throughput is a product differentiator, not just a backend metric. At 1,500 t/s, you're in a regime where agentic loops that previously took minutes complete in seconds.

The lingering question, raised in the thread: why only host small models? The physics of the wafer-scale chip means interconnect between wafers limits what you can run. That's a real architectural constraint, and it means Cerebras speed advantages are model-size bounded for now.


So what?

For founders building agentic workflows where latency compounds across multiple model calls, Cerebras at this speed is worth a serious look. The cost-per-token math may look different when you factor in the wall-clock time savings on multi-step agent tasks.

Read these