I originally wanted to...
I originally wanted to just blurt out either “The M3 is incredible” or “It’s nothing special” — but after testing it, I realized there’s far more to this model than a one-sentence verdict.
Let me get one thing straight upfront: I hold a bias against Minimax. It’s no malicious prejudice, just wariness born from past frustrating experiences.
A few months back, a colleague kept hyping Minimax’s API nonstop, claiming “It works flawlessly; I’m definitely paying for a subscription.” I gave the M2 series a shot, and well… its performance was wildly inconsistent, brilliant one minute and broken the next.
Take this example. I asked it, “Find the nearest McDonald’s around me.” It jumped to the conclusion I wanted fast food because I was in a hurry, then recommended local Chinese fast-casual restaurants instead. Its reasoning left me stunned at how sharp it seemed at the time.
Yet moments later, I told it, “Sort out the meeting minutes from last Wednesday.” It misread “last Wednesday” as yesterday, leaving me completely exasperated.
How am I supposed to rely on you for work if you can’t even get basic timelines right?
From then on, I kept my distance from Minimax. It delivers stellar output when it clicks, but crashes hard when it doesn’t.
So when the M3 launched recently, I was reluctant to even check it out. I’ve grown tired of buzzwords like “native multimodal” and “1 million-token context window” that every new model touts.
Still, my colleague pr...
Still, my colleague pressed me again, insisting “This version is genuinely different.” Fine, I decided to put it to the test.
My overall takeaway: It’s a clear upgrade, yet not revolutionary enough to call it unrivaled.
First, I tested its native multimodal capabilities. The official documentation states it underwent full multimodal training from scratch, rather than being built as a text model with visual features bolted on as an afterthought. I didn’t think much of this distinction at first — until I threw it an extremely tricky test case.
I uploaded a short AI-generated video I’d made, filled with cyberpunk aesthetics and a subtle hidden easter egg: a tiny pixel-art kitten tucked into the background. The model didn’t just identify all the cyberpunk design elements; it even spotted the kitten and commented, “This subtle detail is a nice touch, like a playful Easter egg the creator added for viewers.”
I was genuinely caught off guard. It wasn’t just the high detection accuracy — most modern models pull that off easily. What impressed me was its grasp of intent: it recognized the kitten wasn’t a random visual element, but a deliberate creative choice.
That said, it still has its fair share of failures. I fed it a classic genetics logic puzzle: “A married couple where the husband has normal vision has a color-blind daughter. How is this possible?”
Anyone familiar with basic genetics knows the answer implies infidelity. Color blindness is a recessive X-linked trait; for a daughter to inherit color blindness, her father must carry the defective gene and be color-blind himself.
DeepSeek instantly picked up the contradiction, stating, “Genetically speaking, this scenario is impossible unless…” before gently explaining the underlying implication.
The M3, by contrast, m...
The M3, by contrast, meticulously walked through color blindness inheritance rules and concluded, “It is entirely possible for the daughter to be color-blind.” It completely missed the glaring logical inconsistency.
My assessment: The M3 excels at visual comprehension, yet falls short when it comes to rigorous deductive reasoning.
Next, let’s discuss its token pricing plan.
When I saw the price hike, my immediate reaction was another eye roll. The M2.7 API pricing was barely manageable for me, while the M3 doubled to 4.2 RMB per million tokens.
That said, looking at competitors puts the cost in perspective. Zhipu’s GLM-5.1 charges 6 RMB per million tokens, and Claude’s rates are even steeper. Its monthly subscription tier costs 49 RMB for 600 million tokens, which works out to only 0.08 RMB per million tokens — provided you fully exhaust the monthly allocation.
This reminds me of the gym membership I bought last year. I thought it was an amazing deal at just over 2,000 RMB annually, yet I visited fewer than ten times. Don’t buy a subscription just for perceived value if you won’t use the quota. I’ve told myself this countless times, yet limited-time promotional offers still tempt me every single time.
My current strategy: Hold off on subscribing for now. Test it using the free token allocation for two weeks. If it seamlessly integrates with my daily workflow and handles my core tasks reliably, I’ll consider paying for a plan later.
One unexpected feature also stood out during testing.
My colleague ran an ex...
My colleague ran an extremely demanding task yesterday: he sent the M3 a YouTube video link and asked it to summarize the full video content. The model had no built-in tool for audio transcription, so it troubleshooted independently. It first checked for local video download utilities and found none; then attempted third-party mirror websites, which failed; it tried writing custom scripts that threw errors… and after three or four rounds of trial and error, it successfully located a usable API endpoint to extract the video subtitles.
What amazed me more than the final summary output was its iterative problem-solving process. It didn’t nail the task on the first try, but adjusted its approach each time it hit a roadblock.
It’s comparable to assigning an unfamiliar task to an intern. They mess up twice, but keep experimenting and eventually figure out a viable solution. This adaptive, self-improving quality is rarely seen in other large language models.
To circle back to the core question: How should we evaluate the Minimax M3?
I won’t claim it outperforms every rival model, as real-world testing exposed clear weaknesses, particularly in pure logical reasoning where it still lags behind DeepSeek.
Yet I also can’t dismiss it as mediocre. For agent-based use cases — scenarios requiring the model to independently execute actions, run trials, and troubleshoot complex tasks — it delivers the smoothest user experience among all domestic Chinese LLMs I’ve tested so far.
Here’s some practical, straightforward advice:
For developers: Grab a small, tedious ongoing project you’re working with and run the M3 through real production workflows. Skip generic standardized benchmark tests; a model that streamlines your actual work is far more valuable than one that only excels at casual conversation.
On pricing: Resist rus...
On pricing: Resist rushing to purchase subscriptions immediately. Utilize the free token quota first, and wait until the 7-day limited discount window expires before committing. By then, genuine community feedback from thousands of users will surface to guide your decision. Promotions cycle constantly, but your budget is finite.
A quick side note: The M3 has been open-sourced under the MIT license, permitting commercial use and local on-premises deployment. This is a critical advantage for companies with strict data security compliance requirements.
One final tangent.
Domestic Chinese LLMs are launching new iterations nonstop lately: DeepSeek, Qwen, GLM, Kimi, and now Minimax. I thought I’d grow desensitized to new model releases, yet each platform has its own distinct strengths and flaws after hands-on testing.
DeepSeek resembles a top-tier university student: airtight logical reasoning, albeit overly formal and rigid. Qwen functions as a polished standardized product, consistent yet lacking memorable standout features. The M3, meanwhile, is like a sharp, quick-witted coworker who occasionally overlooks key details. Collaborating with it is efficient, but you’ll need to supervise its work to avoid flawed conclusions.
Your ideal model ultimately hinges entirely on your unique use cases. My recommendation: Don’t get swayed by inflated parameter counts or benchmark scores — pit these models against your own real-world tasks to see which fits best.