🔍 Read the full analysis: Is Mistral Large 4 Closing The Gap With The AI Frontier? on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral introduced an API preview of Large 4 on October 6, 2026, with public weights scheduled for later in the month. Artificial Analysis scores the preview at 38, behind leading U.S. models and some Chinese competitors; the benchmark does not by itself establish performance on every workload.
Mistral AI introduced a public API preview of Mistral Large 4 on October 6, but an October 7 comparison from Artificial Analysis places it behind several leading U.S. and Chinese models on the firm’s Intelligence Index. The preview’s score is 38; its weights are scheduled for release later in October, so the current offering is not yet a downloadable open-weight model.
Mistral describes Large 4 as its largest model to date, built as a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. It accepts text and images. The company says it trained the model on its own infrastructure in Europe and is continuing to improve it, according to the launch announcement.
Artificial Analysis’s Intelligence Index gives the preview a score of 38. In the October 7 snapshot cited by ThorstenMeyerAI.com, that is level with OpenAI’s GPT-6 Luna at maximum reasoning effort, below DeepSeek V4.1 Flash at 39, and behind other listed models, including Claude Opus 5.5 at 58, Gemini 4 Argon at 53 and GPT-6.1 Sol at 52. The comparison uses the stated settings, which are not evaluations under identical compute budgets.
Thorsten Meyer, writing on ThorstenMeyerAI.com, said he would not select the current preview for demanding agentic work or long tasks when stronger-scoring alternatives are available. That is the author’s assessment, not a controlled head-to-head result. He also reported encountering hallucinations during his own use, while acknowledging that this was personal experience, not a comparative study.
Model Watch · October 7, 2026
Is Mistral Large 4 Closing the Gap With the AI Frontier?
Mistral’s new API preview adds a major European model to the field. Its early Intelligence Index score is 38, while downloadable weights are still scheduled for later in October.
Public API access introduced
Scheduled for release in October
Reported token capacity
Accepted modalities
01 / Benchmark snapshot
How the preview compares
Artificial Analysis Intelligence Index figures cited in the October 7 comparison. Reasoning settings vary and were not evaluated under identical compute budgets.
Selected models · Intelligence Index score
02 / What the launch means
Access now, weights later
The practical choice at the time of the report was whether to test the preview through an API. The model was not yet a downloadable open-weight release.
Mixture of experts
Mistral describes Large 4 as its largest model to date, with one trillion total parameters and 49 billion active parameters.
Built in Europe
Mistral says the model was trained on its own infrastructure in Europe and is continuing to improve it.
About 512K tokens
Artificial Analysis reports a context capacity of about 512,000 tokens. Capacity does not establish accurate reasoning across all included material.
October 6
Public API preview announced
October 7
Comparison snapshot published
Later in October
Weights scheduled for release
Next step
Test against real workloads
03 / Evidence and limits
Performance beyond the index
The launch adds a European-developed option, but teams still need current results on the tasks they actually run.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
Thorsten Meyer · personal assessment
“The model was trained on Mistral’s own infrastructure in Europe and continues to be improved.”
Mistral AI · launch announcement, as summarized by ThorstenMeyerAI.com
The index is not a direct test of a specific coding, research, or professional workflow.
Reasoning settings do not use identical compute budgets. Cohere’s Command A+ scored 13 in the cited table, below Mistral.
Meyer reported hallucinations during his own use. That account is not a controlled comparison of error rates.
The source does not establish workload-specific costs, sustained tool-use reliability, or a confirmed weights release date.
04 / What to watch
Test the work, then judge the model
A benchmark gap can guide evaluation. It cannot replace it.
Weight release
Check whether the planned October release happens and what access terms apply.
Long-running tasks
Measure sustained tool use, coding, verification, and recovery from early mistakes.
Compare on your workload
Use current reliability and cost data for your own tasks instead of relying on a dated snapshot alone.
05 / Key questions
At a glance
Facts below reflect the cited report as of October 7, 2026.
What did Mistral announce?
A public API preview of Large 4 on October 6. It accepts text and images and is described as a mixture-of-experts model with one trillion total and 49 billion active parameters.
Is it open-weight now?
No. The preview was available through an API; weights were scheduled for later in October and were not publicly downloadable at the report date.
How does it score?
Artificial Analysis gave the preview 38. That is level with GPT-6 Luna at maximum reasoning effort and below DeepSeek V4.1 Flash at 39 and several listed U.S. models.
Does that rule out agentic work?
No. The index is an aggregate result, not proof of failure on a particular task. The author’s caution about demanding, long tasks is an assessment, not a controlled head-to-head result.
Benchmark Gaps Shape Model Choices
The launch adds a European-developed model to a market where companies and developers weigh capability, access, cost and deployment needs. Mistral says Large 4 was trained on its own European infrastructure, a point relevant to organizations looking for alternatives to models from U.S. and Chinese providers. But the benchmark snapshot does not place it alongside the highest-scoring systems listed in the source.
That gap may matter for teams considering models for multi-step agentic workflows, where a system plans, uses tools and carries decisions forward. A mistake early in a long task can affect later steps. Still, an aggregate benchmark score is not a direct test of a particular company’s coding, research or professional workflow. Buyers need workload-specific evaluations rather than treating the ranking as a verdict on every use.
The comparison also has limits. Cohere’s Command A+ scores 13 in the cited table, below Mistral, so the figures do not support a claim that every competing provider scores higher. The locations in the table describe the developers, not where API requests are processed. Scores are a dated snapshot and may change as models and evaluations are updated.
As an affiliate, we earn on qualifying purchases.
Preview Now, Weights Later
The timing separates the current product from a possible later release: on October 6, Mistral made a preview API available, while its weights were scheduled for release later in October. As of the source article’s October 7 publication, the weights were not publicly downloadable. The practical choice for users at that point was therefore whether to test the model through the API, not whether to download its weights.
Artificial Analysis reports a context capacity of about 512,000 tokens. That indicates how much material can fit into a request; it does not establish that the model can accurately reason across all of it. Mistral promotes agentic coding and specialized professional tasks as strengths, but the source material does not provide controlled, task-by-task results confirming those claims.
The Intelligence Index figures cited by Thorsten Meyer compare specific models and reasoning settings as of October 7, 2026. The author says Claude Opus 5.5 leads the preview by 20 index points, Gemini 4 Argon by 15 and GPT-6.1 Sol by 14. These are index-point differences, not percentages or predictions of success on a given task.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, writing on ThorstenMeyerAI.com
As an affiliate, we earn on qualifying purchases.
Performance Beyond the Index
The available comparison does not show how Large 4 performs across specific coding, research or professional tasks, nor does it establish how reliably it completes long agentic workflows. The Index is an aggregate measure, and the cited reasoning settings do not use identical compute budgets. A score of 38 does not prove failure on any individual task.
Meyer’s account of hallucinations is based on his own use, not a controlled comparison of error rates. The source material also does not establish how the model’s costs compare across a clearly specified workload and price basis. Mistral’s weights had not been released as of October 7, and the source does not give a confirmed release date beyond saying they were scheduled for later in October.
AI model performance evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch for Weights and Tests
The next stated milestone is the planned release of Mistral Large 4’s weights later in October. Whether that timetable holds, and what access terms will apply, remain to be confirmed in the material available as of October 7.
Further evaluation will be needed to judge the preview on specific workloads, including sustained tool use, coding and verification of claims. Mistral says it is continuing to improve the model, so later versions may perform differently. Developers weighing it against alternatives will need current results, costs and reliability data for their own tasks rather than relying on this dated benchmark snapshot alone.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did Mistral announce?
Mistral introduced a public API preview of Large 4 on October 6, 2026. The model accepts text and images and is described as a mixture-of-experts system with one trillion total parameters and 49 billion active parameters.
Is Mistral Large 4 open-weight now?
No. As of October 7, the preview was available through an API, while Mistral’s weights were scheduled for release later in October. They were not publicly downloadable at the time of the report.
How does the preview score against competitors?
Artificial Analysis gives it 38 on the Intelligence Index in the cited October 7 snapshot. That is level with GPT-6 Luna at maximum reasoning effort, below DeepSeek V4.1 Flash at 39, and below several listed U.S. models. The settings are not based on identical compute budgets.
Does the score mean Large 4 cannot handle agentic tasks?
No. The score is an aggregate benchmark result, not a direct measure of every workflow. The source author advises against choosing the current preview for demanding, long agentic tasks when higher-scoring options are available, but that is an assessment rather than proof the model will fail a particular task.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
