🔍 Read the full analysis: The AI Frontier Moves Ahead As Mistral Large 4 Lags on ThorstenMeyerAI.com
Get office and shipping supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Mistral introduced Mistral Large 4 as an API preview on October 6, 2026, with public weights expected later in the month. Artificial Analysis gives it an Intelligence Index score of 38, below leading U.S. models and two Chinese models in the cited comparison; the source author also reports hallucinations in personal use, not a controlled study.
Mistral AI introduced Mistral Large 4 in public API preview on October 6, but an October 7 comparison by ThorstenMeyerAI.com places its benchmark score below several leading U.S. and Chinese models. The model is Mistral’s largest to date, while its weights are not yet publicly downloadable; the company has scheduled their release for later in October.
Artificial Analysis assigns Mistral Large 4 Preview an Intelligence Index score of 38. In the comparison published by ThorstenMeyerAI.com, that matches OpenAI’s GPT-6 Luna at maximum reasoning effort, trails DeepSeek V4.1 Flash at maximum effort by one point, and sits below two other Chinese models and three leading U.S. models. The figures are a dated benchmark snapshot, not a direct prediction of success on any individual task.
The listed U.S. models are Anthropic Claude Opus 5.5 at 58, Google Gemini 4 Argon at 53, and OpenAI GPT-6.1 Sol at 52. China’s Z.ai GLM-5.3 scores 45 and Moonshot AI’s Kimi K3 scores 44. Cohere Command A+ of Canada scores 13, below Mistral. The comparison includes different reasoning settings and does not hold compute budgets constant; the source also cautions that developer locations do not indicate where API requests are processed.
Mistral describes Large 4 as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters, accepting text and images. The company says it trained the model on its own European infrastructure and is continuing to improve it. Thorsten Meyer, the source article’s author, recommends against choosing the current preview for demanding agentic tasks, citing the benchmark results and his experience. His reported hallucinations are personal observations, not findings from a controlled comparison.
AI FRONTIER / MODEL WATCH / OCT 07, 2026
The AI Frontier Moves Ahead As Mistral Large 4 Lags
Mistral Large 4 is now available in API preview, with public weights expected later this month. Its cited Intelligence Index score is 38—below several leading U.S. and Chinese models. That benchmark is a dated snapshot, not a verdict on every task.
INDEX SCORE
38Artificial Analysis snapshotTOTAL SCALE
1TParameters, per MistralACTIVE SCALE
49BMixture-of-experts active parametersCONTEXT
~512KTokens reported by Artificial Analysis01 / BENCHMARK VIEW
Where the score sits
Scores below come from the October 7 comparison. Reasoning settings differ, so treat this as a dated reference point rather than a perfectly controlled ranking.
Selected models above Large 4
Artificial Analysis Intelligence Index · score out of 100
A wider spread
Additional points in the cited comparison
02 / RELEASE AT A GLANCE
What arrived—and what did not
The launch brings a new model to Mistral’s API. Public access to downloadable weights is a separate milestone.
RELEASE
Preview now
Mistral introduced Large 4 in public API preview on October 6, 2026. As of October 7, it was not yet available for public weight download.
MODEL DESIGN
Large MoE
Mistral describes a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. It accepts text and images.
INFRASTRUCTURE
Built in Europe
Mistral says it trained the model on its own European infrastructure and is continuing to improve the preview.
03 / DEVELOPER DECISION
What the score means for real work
A benchmark can inform evaluation, but it cannot settle performance on your workflow.
CAPABILITY SIGNAL
A useful screening point
The cited index places Large 4 below several competing models. It does not prove the model will fail at a particular coding, research, or multi-step task.
WHAT IT DOES NOT SHOW
Context is not reliability
A reported context window of about 512,000 tokens describes how much input may fit. It does not establish reliable reasoning over all of that input.
EVIDENCE LIMIT
No controlled task results
The supplied material provides no controlled, workload-specific results for long agentic tasks, coding, or professional work.
COST CLAIM
Not independently quantified
The source says DeepSeek V4.1 Flash has roughly comparable benchmark intelligence at much lower measured cost per task, but provides no cost figures to assess.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
Thorsten Meyer · ThorstenMeyerAI.comThis is the source author’s recommendation, informed by the cited benchmark and personal experience.
“In my own use of this preview, I encountered hallucinations again.”
Thorsten Meyer · ThorstenMeyerAI.comThese are personal observations, not findings from a controlled comparison or a measured hallucination rate.
04 / EVALUATION PATH
From preview to evidence
Agentic work chains planning, tool use, interpretation, and follow-through. Early errors can affect later actions.
05 / KEY QUESTIONS
What developers should know
The current evidence supports a limited conclusion: substantial scale, a new preview, and a benchmark score behind several competitors.
Is Mistral Large 4 publicly available?
It is available as a public API preview. Its weights were not publicly downloadable as of October 7, 2026; Mistral scheduled their release for later in October.
What does a score of 38 mean?
It is an Artificial Analysis Intelligence Index result, not a percentage or a guarantee of performance on a specific task.
Which models scored higher?
Claude Opus 5.5 scored 58, Gemini 4 Argon 53, GPT-6.1 Sol 52, Z.ai GLM-5.3 45, and Moonshot AI’s Kimi K3 44 in the cited comparison.
Does the benchmark prove it will fail at agentic tasks?
No. The index does not directly measure reliability in every multi-step workflow. Developers need task-specific testing before assigning unsupervised work.
What remains uncertain?
Controlled hallucination rates, long-task performance, detailed cost comparisons, and the effect of future preview updates or public weights.
What is the practical takeaway?
Large 4 is a substantial new preview, but the available material does not establish it as a reliable leading choice for demanding autonomous work.
What the Score Means for Developers
The release gives developers a new Mistral model to evaluate, but the available evidence does not establish that it is a leading choice for complex, extended work. Agentic tasks involve multiple steps—planning, using tools, interpreting results and carrying decisions forward. Errors early in that process can shape later actions, so sustained accuracy and verification matter alongside a model’s headline capabilities.
The benchmark is relevant evidence, but it is not a test of every developer’s workflow. A score of 38 does not prove that Large 4 will fail a particular coding or research task. Likewise, a large context window—Artificial Analysis reports about 512,000 tokens—indicates how much input can fit, not whether the model will reason reliably over all of it. Developers considering the preview need task-specific testing before assigning it unsupervised work.
The comparison also has limits: reasoning settings differ, the scores are a snapshot, and the source says DeepSeek V4.1 Flash has approximately comparable benchmark intelligence at a much lower measured cost per task. The supplied material does not include the underlying cost figures, so the cost claim cannot be independently quantified here.
As an affiliate, we earn on qualifying purchases.
Preview Now, Weights Later
The distinction between an API preview and a public-weight release matters. As of October 7, Large 4 can be accessed through a preview API, but its weights are not yet available for public download. Mistral says those weights are due later in October, meaning the current decision for developers concerns the preview rather than a finished downloadable release.
Mistral’s launch is also a development in European AI capacity: the company says it trained the model on its own infrastructure in Europe. That fact describes where and how Mistral says it built the system; it does not establish that the model matches the strongest alternatives on performance. The cited source says Mistral is continuing to improve the preview, so subsequent versions and the eventual weight release may warrant separate evaluation.
“I would not choose it for demanding agentic work or long tasks when stronger models are available.”
— Thorsten Meyer, ThorstenMeyerAI.com
Performance Beyond the Benchmark
The available material does not provide controlled, workload-specific results showing how Large 4 performs on long agentic tasks, coding, or professional work. Mistral promotes strengths in agentic coding and specialized tasks, but those claims require testing against the workflows developers intend to use. The source author’s hallucination observations also cannot establish how often the model produces unsupported answers across users or tasks.
It is not yet clear whether later preview updates or the scheduled weight release will change the benchmark picture. The Artificial Analysis scores are a snapshot dated October 7, 2026, and the listed models were evaluated with different reasoning settings. The source provides no complete cost table, so its claim that DeepSeek is much cheaper cannot be assessed in detail from the supplied figures.
October Weights and Further Testing
Mistral has scheduled the model weights for release later in October, according to the source material. Developers and evaluators can then assess the released model and any preview updates against their own tasks, including whether it follows constraints, checks evidence, and maintains reliable performance across multiple steps.
Further benchmark results may also change the current comparison. Until those results and workload-specific evaluations are available, the evidence supports a limited conclusion: Large 4 is a new preview with substantial model scale, but the cited index places it behind several competing models, and the current material does not establish its reliability for demanding autonomous work.
Key Questions
Is Mistral Large 4 publicly available?
It is available as a public API preview, according to the source. Its model weights were not publicly downloadable as of October 7, 2026; Mistral has scheduled their release for later in October.
What score did Mistral Large 4 receive?
Artificial Analysis gives Mistral Large 4 Preview an Intelligence Index score of 38 in the comparison cited by ThorstenMeyerAI.com. The score is a benchmark result, not a percentage or a guarantee of performance on a specific task.
Which models scored higher in the cited comparison?
Anthropic Claude Opus 5.5 scored 58, Google Gemini 4 Argon 53, OpenAI GPT-6.1 Sol 52, Z.ai GLM-5.3 45, and Moonshot AI’s Kimi K3 44. The source notes that the comparison uses different reasoning settings.
Does the benchmark prove Large 4 will fail at agentic tasks?
No. The cited index does not directly measure reliability on every coding, research, or multi-step workflow. The source author advises against selecting the current preview for demanding agentic work, but frames that as a judgment informed by benchmarks and personal experience.
What remains to be evaluated?
Developers still need workload-specific tests of the preview and, when available, the public weights. The supplied material does not establish controlled hallucination rates, performance across long tasks, or detailed cost comparisons.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
