Mistral Large 4: An AI Contender Beyond The US And China With Agent Limits
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4: An AI Contender Beyond The US And China With Agent Limits on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral has released Large 4 as a research preview, with a 38.4 score on Artificial Analysis Intelligence Index v4.3.2. The source material describes a major improvement over earlier Mistral models but says leading US and Chinese models score higher, while benchmark cost, verbosity and reported hands-on hallucinations raise concerns about agent use.

Mistral released Large 4 as a research preview, and the model scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2. The result marks a sharp improvement over the company’s previous models, but the benchmark places Large 4 below current top US and Chinese systems, leaving open questions about its value for costly, multi-step AI agents.

Artificial Analysis lists Large 4 at 38.4 points. In the same index version, the source’s comparison table gives US models scores as high as 57.6 and Chinese models scores of 44.8, 43.6, 41.8 and 39.5. Large 4 therefore trails those listed systems, although it scores above DeepSeek V4 Pro 0813 at 36.0 and GLM-5.2 at 33.7. Artificial Analysis provides the benchmark figures; the supplied material does not include a response from Mistral about the comparisons.

The model is described as having one trillion total parameters, with 49 billion active, and as accepting text and images while producing text. It has a 512,000-token context window. Mistral’s API offers it in research public preview. The source says the weights are promised for the end of October; until they are released, the model is proprietary, and its licence has not been published.

The listed standard API prices are $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million. The source reports a 50% discount for the first two weeks. Artificial Analysis estimates the model costs $1.13 per Intelligence Index task. That is more than the cited estimates for GLM-5.3-Flash at $0.25 and DeepSeek V4.1 Flash at $0.27; both models score higher in the index. These are benchmark task-cost estimates, not a guarantee of the cost of a particular customer workload.

At a glance
reportWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 as a research preview, posting a 38.4 Artificial Analysis Intelligence Index score while remaining behind leading US and Chinese models.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$84,849▼ 1.1%
Ethereum ETH$2,674▼ 1.4%
Tether USDT$0.9999▼ 0.0%
BNB BNB$773.66▼ 1.3%
XRP XRP$1.49▼ 1.3%
USDC USDC$0.9999▼ 0.0%
Solana SOL$119.56▼ 1.1%
TRON TRX$0.3349▼ 0.4%
Live data · CoinGecko · alternative.me (24h change)
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

The Cost of Running Long Agents

Large 4’s result matters because the index includes work intended to measure agentic performance, such as knowledge-work tasks, SaaS workflows and coding. A model used as an agent must often carry information across several steps, call tools and act on earlier outputs. A mistake or unsupported assertion can affect what the agent does next, so single-score comparisons do not capture every deployment risk.

The source reports that Large 4 used 200 million output tokens to complete the index, compared with a median of 81 million for comparable models. If reproduced under a buyer’s workload, high output volume could add cost and latency alongside the model’s per-token price. The supplied material does not provide enough detail to establish that every real-world task would show the same gap.

The author also says hands-on testing found confident false assertions. That is an attributed personal observation, not a result from the Artificial Analysis index. It adds a reason for teams to test the model against their own tasks, especially where errors could flow into later steps. The available source does not establish a general hallucination rate for Large 4.

Amazon

AI language model API

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Rise From Large 3

The strongest case for Large 4 in the supplied material is its improvement over Mistral’s earlier scores. On the same Artificial Analysis index version, Large 3 scored 9 and Medium 3.5 scored 14, compared with 38.4 for Large 4. That is a substantial jump in the benchmark’s results, even though it does not put the model at the top of the current rankings.

The launch framing highlighted Mistral as the most intelligent model outside the United States and China. That description is tied to the comparison set in the source; it does not mean Large 4 is ahead of leading models from those two countries. The source also notes that this category has few competitors, which limits what can be inferred from the label.

Artificial Analysis v4.3.2 is the stated basis for the score comparisons and task-cost estimates. The source says Mistral reported that reinforcement learning is still running, so results could change. The index score is one evaluation, not a complete measure of performance across every use case, and the supplied material does not give details sufficient to independently assess all test conditions.

“In hands-on testing I saw Large 4 assert things confidently that weren’t true.”

— ThorstenMeyerAI.com author

Amazon

large language model with image input

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Scores and Reliability

Large 4 remains a research public preview, and Mistral reportedly says training work is continuing. It is not yet clear whether the 38.4 score will change, how performance varies across individual index tasks, or how well the benchmark results translate to customers’ workloads. The source does not provide a Large 4 hallucination rate from a standardized evaluation.

The promised end-of-October release of model weights and the licence terms also remain outstanding in the supplied account. Until those details are available, prospective users cannot confirm the conditions under which they may run or adapt the weights. Pricing may also vary from the listed standard rates during the reported initial discount period.

Amazon

AI model cost optimization tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Further Evaluation

The next stated milestone is the planned end-of-October release of Large 4’s weights, along with publication of the licence needed to understand their permitted use. Mistral’s ongoing reinforcement learning could also lead to updated model performance, but the source gives no timetable or promised revised score.

For now, the model is available through Mistral’s API in research preview. Buyers weighing it for agent workflows can compare its price and index results with alternatives, then test reliability, output volume and task completion on their own work before wider deployment. The source does not report a customer rollout or independent evaluation of Large 4 in production.

Amazon

tokens and context window management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What score did Mistral Large 4 receive?

It scored 38.4 on the Artificial Analysis Intelligence Index v4.3.2, according to the supplied source material.

Is Large 4 available as an open-weight model now?

No. The source describes it as a proprietary research preview on Mistral’s API and says weights are promised for the end of October. It says the licence has not yet been published.

How does Large 4 compare with the listed alternatives?

In the cited index table, several US and Chinese models score above Large 4. The source also estimates that GLM-5.3-Flash and DeepSeek V4.1 Flash cost less per index task while scoring higher. Those are benchmark comparisons, not predictions for every workload.

Does the source establish that Large 4 has a high hallucination rate?

No. It reports an author’s hands-on observation of confident false assertions, but provides no standardized hallucination-rate result for Large 4. The observation should not be treated as a measured general rate.

What should buyers watch for next?

The stated next steps are the planned weight release and licence publication, as well as any updated results while Mistral continues reinforcement learning. Buyers can also test the preview against their own task accuracy, token use, cost and reliability requirements.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple Silicon’s Quiet Memory Advantage

Apple Silicon’s unified memory architecture allows running large AI models beyond traditional GPU limits, offering capacity at lower cost and power.

2026 External GPU Trends For AI: 8 Must-See Options

Explore the 8 must-see external GPUs in 2026, highlighting the latest trends, features, and performance options for AI, gaming, and creative work.

How to Reduce Heat and Noise in a High-Power AI Workstation

Effective strategies to lower heat and noise in high-performance AI workstations, focusing on undervolting, cooling, and airflow optimization.

2026 AI Innovations: Elevating Gaming And Everyday Tech

Major AI advancements in 2026 are enhancing gaming experiences and everyday devices, with confirmed new capabilities and ongoing developments.