Xiaomi’s MiMo AI Is Now 15x Faster Than ChatGPT and Claude, Without a Single Custom Chip
Xiaomi's MiMo-V2.5-Pro-UltraSpeed has crossed 1,000 tokens per second on a standard 8-GPU server, a milestone companies like Cerebras and Groq built custom chips to reach, achieved here through software alone.

Most people know Xiaomi as a phone brand. Some know it for electric scooters. Almost nobody had it on their list of companies likely to break a major AI inference record. And yet, on Monday morning, that is exactly what happened.
Xiaomi, working in collaboration with inference partner TileRT, has released MiMo-V2.5-Pro-UltraSpeed, breaking the 1,000 tokens per second decode speed barrier on a 1-trillion-parameter model for the first time. Demos show it peaking near 1,200 tokens per second. The kicker: it runs on commodity GPUs, a single standard 8-GPU node, not custom silicon.
That last detail is what makes this announcement genuinely significant.
What 1,000 Tokens Per Second Actually Means
Tokens are the chunks of text an AI model reads and writes, roughly three-quarters of a word each. The speed at which a model generates them determines how fast responses arrive, how many requests a server can handle simultaneously, and whether the model is fast enough for real-time applications.
To put the number in perspective: GPT-5.5, what most ChatGPT users are talking to, sits at 68 tokens per second. Claude Opus 4.6 lands around 71, with the lower-end Haiku model touching 98. Gemini Flash hits 192 tokens per second. MiMo-V2.5-Pro-UltraSpeed does 1,000, on a model that matches Opus on coding benchmarks.
That gap is not incremental. It is a different category of performance entirely.
The Companies That Built Chips for This Problem
The inference speed bottleneck is not a new problem. Two well-funded companies, Cerebras and Groq, built entire businesses around solving it with custom silicon.
Cerebras designed a wafer-scale chip the size of a dinner plate, packing 44GB of on-chip memory to eliminate the bandwidth bottleneck that slows GPU inference. It hit 969 tokens per second on Meta’s Llama 3.1 405B, which is impressive, but that is a 405-billion-parameter model, less than half the size of MiMo-V2.5-Pro. Groq’s custom Language Processing Unit architecture tops out around 300 to 750 tokens per second depending on the model.
Neither runs on hardware you can rent from a standard cloud provider. Xiaomi’s approach does.
How They Actually Did It
The speedup comes from three coordinated techniques across the model and the serving system, what Xiaomi calls extreme model-system codesign.
The first is FP4 quantization. Instead of running the model at full numerical precision, Xiaomi compresses the expert layers, which make up most of the 1 trillion parameters, down to 4-bit. Memory footprint drops, bandwidth pressure drops, and speed increases. Critically, only the expert layers are compressed while everything else stays at full precision, keeping quality loss near zero.
The second is DFlash speculative decoding. Normal speculative decoding has a small draft model guess the next few tokens, then the large model verifies them in parallel. DFlash skips sequential drafting entirely; it fills a whole block of masked positions in a single forward pass. In coding tasks, the large model accepts an average of 6.3 out of 8 proposed tokens per verification round. Six tokens confirmed in one step instead of one.
The third is TileRT itself, the inference engine that keeps the entire compute pipeline continuously resident inside the GPU, eliminating per-operator launch overhead and execution gaps.
Neither technique alone reaches 1,000 tokens per second. Together, they do.
What It Changes in Practice
Raw speed numbers are only meaningful if they unlock something new. To put the generation rate in human terms, MiMo-V2.5-Pro’s earlier Flash model was already generating responses at 150 tokens per second, roughly 110 words per second, faster than the fastest human can read or speak. UltraSpeed pushes that ceiling significantly higher.
At 1,000 tokens per second, applications that were previously impossible become viable. Fraud detection, real-time trading signals, parallel reasoning chains, and live agent loops – all of these have hard latency requirements that 68 tokens per second cannot meet. At 1,000, they can.
Xiaomi MiMo AI: Access, Pricing, and Openness
The API trial runs June 9 to June 23, 2026, on an application basis with priority given to enterprise and professional developers. Pricing is set at 3x the standard MiMo-V2.5-Pro rate for approximately 10x the generation speed. The Token Plan is not supported; this is API access only.
For those who want to test the underlying technology without applying for API access, Xiaomi has open-sourced the MiMo-V2.5-Pro-FP4-DFlash checkpoint on Hugging Face, and TileRT has open-sourced select modules on GitHub.
One caveat worth noting: independent third-party speed verification is not yet public. The numbers come from Xiaomi’s own benchmarks and demos. Community testing through the open-sourced checkpoint will be the real verification process over the coming weeks.
The Bigger Picture
The phone brand has been steadily building an AI capability that most of the industry was not paying attention to. MiMo-V2.5-Pro already matches Claude Opus on coding benchmarks at a fraction of the cost, approximately $0.43 input and $0.87 output per million tokens, compared to Opus at $5 input and $25 output. UltraSpeed does not change the model’s intelligence; it accelerates the same model that was already competing at the frontier.
If the speed claims hold up under independent scrutiny, Xiaomi has done something that required hundreds of millions of dollars in custom silicon investment from Cerebras and Groq, using software running on standard hardware that any developer can access today.
That is the kind of result that reshapes assumptions about where AI capability comes from next.
Mobile Phone Taxes Portal
Find the PTA Taxes on All Phones on a Single Page using our Taxes Portal.
Note: Mobile phone tax rates and calculations fall under the jurisdiction of the Federal Board of Revenue (FBR), not the Pakistan Telecommunication Authority (PTA).
Explore NowFollow us on Google News!