Sorry but mimo v2.5 pro really doesn't cut it

Some people share that Xiaomi’s model can perform long-horizon tasks, executing hundreds of tool calls without errors. I don’t know who initiated this testing direction, but it’s like saying a student can hold their bladder for 100 minutes in an exam room—meaningless. Normal people understand that for a standardized exam, the only thing that matters is the score. To be more lenient, the score accounts for over 95% of the evaluation system…

Some people share that Xiaomi’s model can perform long-horizon tasks, executing hundreds of tool calls without errors. I don’t know who initiated this testing direction, but it’s like saying a student can hold their bladder for 100 minutes in an exam room—meaningless. As normal people understand it, for a standardized exam, the only thing that matters is the score. To be more lenient, the score accounts for over 95% of the evaluation system. Generally, we look at the score first, then other performance; if the score is low, there’s no need to look at anything else. mimo can indeed complete a task, but the problem is it doesn’t do it correctly. The following is a subjective, non-quantitative share about mimo 2.5 pro, comparing it only against industry benchmarks GPT5 series and opus series. Poor instruction following and poor recall. Despite having a 1M context, the value of a large context is limited if recall is low. Even when I explicitly specified the ATDD development process in the agent, mimo did not generate case documents or test cases for me. Poor intent understanding, causing tasks to potentially deviate from the start. GPT5 and opus both gather a lot of information at the beginning of a session—the information-gathering phase at session start is slow, but subsequent phases are fast. The benefit of this slowness is fewer intent-understanding deviations; the downside is significantly higher token consumption. mimo 2.5 pro lacks the willingness to gather information at session start, so I believe this model can only handle ultra-small-scale programming tasks, despite having a 1M context. Because the model has low willingness to gather environment information, mimo’s support for Windows tools is poor, defaulting to Linux commands. Unless necessary, I develop on a Windows Dev Drive by default. Needless to say for GPT5/opus—benchmark models don’t have this problem—and glm5.1 and DeepSeek v4 flash/pro also perform normally. Lacks general knowledge, unable to recognize fairly obvious common-sense errors, and will stubbornly continue. In the end it wastes many tokens producing an absurd solution and modifications; it would be better to strengthen the model’s willingness to build context. To sum up, mimo v2.5 pro has strong work willingness, weak exploration willingness, poor memory, poor general knowledge, poor logic, and poor Windows support—a thoroughly unremarkable model. Although I appreciate the token plan gifted by Mr. Lei, compared to GPT5 the gap is like that between a diligent elementary school student and a PhD, far from the “80% capability” some people claim. My subjective experience is all from programming tasks, about 10 requests, consuming 6% of a 329-yuan pro plan—roughly estimating the 329-yuan plan can only send about 200 requests in my usage scenario. Except for one simple single-page modification that was accepted, all other implementations were reverted. Maybe it’s a model suited for writing articles? In any case, it’s not suited for writing code.

Sorry but mimo v2.5 pro really doesn’t cut it Image 1

Sorry but mimo v2.5 pro really doesn’t cut it Image 2