🔍 Read the full analysis: Mistral Large 4: An Option Beyond The US And China, But Not For Agent Work on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
Mistral Large 4, released as a research public preview, scored 38.4 on Artificial Analysis Intelligence Index v4.3.2, making it a notable advance for the French lab but leaving it behind leading US and Chinese models. The source’s analysis questions its value for agent workflows, citing task costs, output verbosity and reported hands-on hallucinations; Mistral says reinforcement learning is ongoing, and the model’s weights are not yet available.
Mistral released Large 4 as a research public preview, and the model scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The result is a substantial rise from the company’s previous Large model, but the index places it below current leading US and Chinese systems, raising questions about its suitability and cost for multi-step agent work.
Artificial Analysis’ current index lists Large 4 below the leading US models and several Chinese models. Its score of 38.4 trails the listed leaders by more than 19 points; the same table puts China’s GLM-5.3 at 44.8, Kimi K3 at 43.6 and DeepSeek V4.1 Flash at 39.5. Large 4 is ahead of GLM-5.2 and DeepSeek V4 Pro in that comparison. The source report describes it as the highest-scoring model from outside the United States and China, while cautioning that this framing does not make it a peer of the top frontier systems.
The report says Large 4 has one trillion parameters, with 49 billion active, accepts text and images, and has a 512,000-token context window. It is currently available through Mistral’s API as a research preview. Mistral has said the model’s weights are expected at the end of October; until then, buyers access a proprietary model, and the report says its licence has not been published.
Listed API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input priced at $0.14. The source says Mistral is offering a 50% discount for the first two weeks. It also reports that Mistral says reinforcement learning is continuing, so the model’s benchmark results may change. Those details make the current score and price a snapshot of a developing release, not necessarily its final profile.
Mistral Large 4: best outside the US and China — and still not a model to run your agents on
The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.
~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.
Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.
Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.
The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.
AA v4.3.2Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.
AAConfident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.
AUTHOR’S TESTING · not an AA figure- Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
- Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
- Speed: 116 tok/s, 1.46s TTFT — well above median.
- The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
- Jurisdiction: French parent, EU hosting, weights promised end of October.
- Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
- Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
- Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.
The Case for Agent Workflows
The benchmark matters to teams considering models for long-running, multi-step tasks, not just short exchanges. Artificial Analysis’ index incorporates agent-oriented evaluations, including knowledge work, SaaS workflows and coding tasks. A lower score does not translate directly into a fixed failure rate, but the report argues that errors can accumulate over successive steps: a weak answer early in a workflow may become the premise for later actions.
Cost is another practical concern. According to figures in the source report, Large 4 costs about $1.13 per Intelligence Index task. GLM-5.3-Flash is listed at $0.25 per task with a score of 41.8, while DeepSeek V4.1 Flash is listed at $0.27 with a score of 39.5. These comparisons suggest buyers should weigh benchmark performance alongside the workload and actual usage costs, rather than treating the model’s national origin as a proxy for value.
The report also says Large 4 generated 200 million output tokens while completing the index, compared with a median of 81 million for comparable models. If that pattern carries over to a buyer’s tasks, verbosity could add both latency and expense. The report’s author separately describes seeing confident false claims during hands-on testing. That is an attributed observation, not an Artificial Analysis benchmark result, and it is particularly relevant where an agent can carry an incorrect assertion into later steps.
As an affiliate, we earn on qualifying purchases.
From Large 3 to Large 4
The clearest evidence of progress is the change between Mistral’s own releases. On the same version of the Artificial Analysis index, Large 3 scored 9, while Medium 3.5 scored 14; Large 4 now scores 38.4. That is a major reported gain for Mistral, even though the updated model remains behind the leading entries in the broader comparison.
The index table in the source report includes six leading US systems scoring from 51.8 to 57.6, followed by Chinese models such as GLM-5.3, Kimi K3 and GLM-5.3-Flash. Large 4’s score falls below those listed systems. The source’s description of it as an option beyond the US and China is thus best understood as a geographic distinction, not evidence that it matches the top-ranked models on the benchmark.
The report’s comparisons are tied to Artificial Analysis Intelligence Index v4.3.2, which combines several task evaluations. Results should be read as benchmark measurements under that version, not a guarantee of performance on every company’s software, data or agent setup. Mistral’s statement that reinforcement learning is still underway also means the current ranking could move.
“Large 4 shows it has closed a lot of ground — and that it is still not one.”
— ThorstenMeyerAI.com report
Preview Results and Open Questions
Several details remain unsettled. The model is in research preview, reinforcement learning is ongoing, and the source says weights are promised for the end of October. The final release date, final benchmark standing and terms governing those weights are not confirmed in the supplied material. The report says the licence is unpublished, so prospective users cannot yet assess those terms from the information provided.
The source’s concern about hallucinations comes from the author’s own hands-on testing; it does not provide a test protocol, sample size or a measured hallucination rate for Large 4. Its cost and output-token comparisons are tied to the Artificial Analysis evaluation and may not match a particular deployment’s mix of prompts, cached inputs and generated text. The available material also does not establish how the preview performs across different languages or specific customer workloads.
Weights and Updated Benchmark Scores
The next stated milestone is Mistral’s planned release of Large 4 weights at the end of October. That release should clarify access and licensing, though the supplied source does not give a confirmed licence or exact date. Buyers can also watch for new Artificial Analysis results as Mistral continues reinforcement learning and the preview develops.
For now, organisations weighing Large 4 against alternatives will need to test it on their own tasks and monitor total token use, error handling and the consequences of incorrect outputs. The current evidence supports two conclusions at once: Large 4 is a substantial improvement over Mistral’s earlier scores, and the available benchmark and pricing figures do not make it an obvious choice for agent workloads.
Key Questions
What is Mistral Large 4?
Mistral Large 4 is a Mistral model available as a research public preview through the company’s API. The source describes it as a one-trillion-parameter model with 49 billion active parameters, text-and-image input and a 512,000-token context window.
How did Large 4 score against other models?
It scored 38.4 on Artificial Analysis Intelligence Index v4.3.2. The listed leading US models scored from 51.8 to 57.6, while several Chinese models also scored above Large 4 in the source’s comparison.
Is Large 4 available as an open-weights model now?
No. The source says it is currently a proprietary API preview and that Mistral has promised weights for the end of October. It reports that the licence has not been published.
Why does the report question using it for agents?
The report points to its benchmark position, estimated task cost and high output-token use, and its author reports seeing confident false claims in hands-on testing. The hallucination observation is the author’s account, not a published Artificial Analysis rate for Large 4.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
