The Case For Mistral Large 4 Outside The US And China—and Its Agent Limits
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Case For Mistral Large 4 Outside The US And China—and Its Agent Limits on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 recorded 38.4 on Artificial Analysis’s Intelligence Index v4.3.2, a sharp rise from its predecessor and a strong result among models outside the United States and China. The source’s benchmark and pricing comparisons suggest it trails leading US and Chinese models, while concerns about verbosity and observed hallucinations may limit its suitability for long-running agents. Its weights, licence and final benchmark performance remain unsettled.

Mistral has released **Mistral Large 4** as a research preview, with Artificial Analysis scoring it **38.4 on its Intelligence Index v4.3.2**. The result makes it the highest-scoring model from outside the United States and China in the comparisons cited by ThorstenMeyerAI.com, but the same data places it below leading US and Chinese systems and raises questions about its cost and performance on agent tasks.

Mistral describes Large 4 as a **one-trillion-parameter model with 49 billion active parameters**, capable of taking text and images as input and producing text. It has a **512,000-token context window** and is available through Mistral’s API as a research public preview. The source says Mistral plans to release the model weights at the end of October; until then, it is proprietary, and the licence has not been published.

Artificial Analysis’s current Index puts Large 4 at **38.4 points**, up from 9 for Mistral Large 3 on the same Index version. That is a substantial improvement, but the cited table ranks it below the leading US models and several Chinese models, including GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash. The source characterizes the result as a European advance, not evidence that Mistral has reached the top tier.

The listed API price is **$1.36 per million input tokens and $4.18 per million output tokens**, with cached input at $0.14 per million. The source reports a 50% introductory discount for the first two weeks. It also says Artificial Analysis measured 200 million output tokens for Large 4’s Index tasks, compared with an 81 million median for comparable models, making output volume relevant to the overall cost of benchmark work.

At a glance
analysisWhen: Released the day before the source repo…
The developmentMistral has released Large 4 in research preview, prompting a debate over its standing outside the US and China and its suitability for agent workloads.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$85,161▼ 0.5%
Ethereum ETH$2,683▼ 0.9%
Tether USDT$0.9999▼ 0.0%
BNB BNB$775.58▼ 0.8%
XRP XRP$1.49▼ 0.7%
USDC USDC$0.9999▼ 0.0%
Solana SOL$119.75▼ 0.6%
TRON TRX$0.3351▼ 0.3%
Live data · CoinGecko · alternative.me (24h change)
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Benchmark Gains, Agent Trade-offs

The result matters to organizations looking for a capable model from a European provider, particularly those seeking alternatives to US and Chinese suppliers. Large 4’s **jump from 9 to 38.4** marks a clear improvement in Mistral’s measured performance. But the source’s comparisons make clear that geography alone does not establish technical or economic advantage: several Chinese models score higher and are reported to cost much less per benchmark task.

For agent use, a model’s ability to sustain multi-step work matters alongside its headline score. The source argues that **errors can compound across a long workflow**, while verbosity adds output cost and latency at each step. Its report of confident hallucinations comes from hands-on testing, not the Artificial Analysis benchmark, and should be treated as an attributed observation rather than an independently established rate. Still, it highlights a practical risk: an incorrect answer used as an agent’s premise can affect later actions.

These findings do not establish that Large 4 is unsuitable for every deployment. They point instead to a need for task-specific testing, including reliability, token use, latency and total cost. For high-stakes or extended workflows, benchmark rank alone may not capture the operational risk.

Amazon

AI model API pricing

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Sharp Rise From Large 3

The comparison in the source uses **Artificial Analysis Intelligence Index v4.3.2**, which includes evaluations of agentic knowledge work, real-world work tasks, software workflows and coding. That mix makes the score relevant to agent capabilities, though it remains a benchmark result rather than a guarantee of performance in a particular company’s tools or tasks.

Mistral’s earlier Large 3 scored **9** on the same Index version, while Medium 3.5 scored 14. Large 4’s result is therefore a major step within the company’s own lineup. The source also notes that Mistral says reinforcement learning is still underway and that scores may change, so the preview result should not be treated as necessarily final.

The phrase “most intelligent model outside the US and China” depends on which competitors are included. The source says few labs in other regions are competing at this level and describes the comparison as a field with limited entrants. That geographic distinction may matter for procurement, but it is separate from how Large 4 performs against the strongest models overall.

“Reinforcement learning is still running, so scores may move.”

— Mistral

Amazon

large language model with 512k token window

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Results Still May Change

Large 4 is a **research preview**, and Mistral says reinforcement learning is ongoing. The source provides no final benchmark result, and the announced end-of-October timing for weights is a plan rather than a completed release. The model’s licence is also unpublished, leaving a material question for organizations considering self-hosting or redistribution.

The source does not provide a controlled, independently verified hallucination rate for Large 4. Its account of confident errors is explicitly based on hands-on use. Nor does the benchmark alone show how the model will perform on a buyer’s specific agent workflow, where tools, prompts, safeguards and task length can change results.

Pricing comparisons also depend on usage patterns. The source gives API rates and per-task benchmark estimates, but those figures do not settle total deployment cost across different input lengths, caching, retries or agent designs. The introductory discount is time-limited, and the source does not specify its end date beyond saying it lasts two weeks.

Amazon

AI model output token counter

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights, Licensing and Retesting

The next concrete milestone is Mistral’s planned **release of Large 4’s weights at the end of October**. Buyers will be able to assess the licence once it is published and determine whether the release permits their intended uses. Until then, access described in the source is through Mistral’s API preview.

Artificial Analysis scores may also change as Mistral’s reinforcement learning continues. Organizations evaluating the model should compare updated results with their own tests, particularly for long-running agents, and track output-token use alongside accuracy and latency. The source does not identify a confirmed date for a final benchmark or a production release.

Amazon

AI model license and weights

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is Mistral’s **one-trillion-parameter model**, with 49 billion active parameters, text and image input, text output and a 512,000-token context window. The source describes it as available through the API in research preview.

How does Large 4 score against other models?

Artificial Analysis’s Index v4.3.2 gives it **38.4 points**. The cited comparison places it below leading US models and several Chinese models, while above Mistral Large 3’s score of 9 on the same Index version.

Is Large 4 suitable for AI agents?

The available information does not establish a universal answer. The source flags verbosity and reports observed confident hallucinations, concerns that may matter in multi-step work. Teams should test the preview on their own tasks before relying on it.

When will the model weights be available?

Mistral’s reported plan is to release the weights **at the end of October**. The source says the licence has not yet been published, so the terms of use remain unknown.

How much does API access cost?

The reported standard price is **$1.36 per million input tokens** and **$4.18 per million output tokens**, with cached input at $0.14 per million. The source reports a 50% discount for the first two weeks; actual costs depend on usage. Model outputs and API use can carry financial and operational risk, and these figures are not a recommendation to use or avoid the service.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What Is Zero-Knowledge Proof

Keen to understand how zero-knowledge proofs ensure privacy in digital transactions? Discover the fascinating mechanics behind this cutting-edge cryptographic technique.

A Frontier AI Model Just Went Dark for 18 Days. The Kill-Switch Is Real Now.

A leading AI model was forcibly shut off for 18 days by US government order, marking a shift in AI governance and raising questions about future regulation.

Religious Scholars On AI: Inside Their Meeting With Anthropic

A New York Times headline reports a meeting between religious scholars and Anthropic, but available details do not explain what was discussed or what followed.

OpenAI’s Latest Move: Cutting Off The Cursor And Its Developer Fallout

OpenAI plans to end its models’ support for Cursor by November 12, citing trust issues after Cursor’s acquisition by SpaceX, impacting developers relying on the tool.