Mistral Large 4 Still Has Ground To Make Up In AI
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Mistral Large 4 Still Has Ground To Make Up In AI on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral launched Large 4 as an API preview on October 6, with a trillion total parameters and 49 billion active parameters. Artificial Analysis gives it an Intelligence Index score of 38, below leading U.S. models and two stronger Chinese models in the October 7 snapshot. The article’s author says he would not select the preview for demanding agentic work, citing the benchmark gap and hallucinations he encountered; those experiences are not a controlled comparison.

Mistral launched Large 4 in public API preview on October 6, but an October 7 comparison by Artificial Analysis places the model behind several leading U.S. and Chinese systems on its Intelligence Index. Thorsten Meyer, writing on ThorstenMeyerAI.com, says he would not choose the current preview for demanding agentic work or long tasks when stronger-scoring alternatives are available. That is his assessment of a preview, not a controlled measure of how every model performs on those workflows.

Mistral describes Large 4 as its largest model to date: a mixture-of-experts system with one trillion total parameters and 49 billion active parameters. The preview API accepts text and images. Mistral says it trained the model on its own infrastructure in Europe and is continuing to improve it. The company has scheduled a release of the model weights for later in October; the weights were not publicly downloadable when Meyer published his assessment on October 7.

Artificial Analysis gave Large 4 Preview an Intelligence Index score of 38. In the same dated snapshot, Anthropic’s Claude Opus 5.5 scored 58, Google’s Gemini 4 Argon 53, and OpenAI’s GPT-6.1 Sol 52. China’s Z.ai GLM-5.3 scored 45 and Moonshot AI’s Kimi K3 scored 44. DeepSeek V4.1 Flash scored 39, while OpenAI’s GPT-6 Luna also scored 38. The settings differ across models, so this is not a comparison under identical reasoning or compute budgets.

Meyer also reports encountering hallucinations in his own use of the preview. He says this is personal experience, not a controlled comparative study. The source material names Artificial Analysis’s scores and model profile and Mistral’s announcement, but its discussion of cost ends mid-sentence. It does not provide the cost figures needed to assess the relative price of completing a task.

At a glance
reportWhen: Announced October 6, 2026; benchmark sn…
The developmentMistral has released Large 4 in public API preview, with benchmark results placing it behind leading U.S. models and some Chinese competitors.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$83,788▼ 2.1%
Ethereum ETH$2,606▼ 3.5%
Tether USDT$0.9998▼ 0.0%
BNB BNB$760.15▼ 2.6%
XRP XRP$1.45▼ 2.9%
USDC USDC$0.9999▼ 0.0%
Solana SOL$117.43▼ 2.6%
TRON TRX$0.3337▼ 0.7%
Live data · CoinGecko · alternative.me (24h change)
Mistral Large 4 Still Has Ground To Make Up In AI

AI MODEL WATCH · OCTOBER 7, 2026

Mistral Large 4 Still Has Ground To Make Up In AI

Mistral’s new API preview brings trillion-parameter scale and a European infrastructure story. In the cited benchmark snapshot, its score trails several leading U.S. models and higher-scoring Chinese competitors.

38

Artificial Analysis Intelligence Index score for Large 4 Preview. A dated aggregate benchmark result, not a direct measure of every workflow.

Snapshot · October 7, 2026
1TTotal parameters
49BActive parameters
API previewOct 6Public launch
Index score38Large 4 Preview
Context capacity~512KReported by Artificial Analysis
WeightsLater in OctScheduled; not yet available Oct 7

The gap in frontier benchmarks

The October 7 Artificial Analysis snapshot places Large 4 below several listed U.S. systems and stronger-scoring Chinese alternatives. Scores reflect this index and its evaluation settings; they are not percentages or guaranteed differences in task success.

ModelScoreDeveloper baseReading the result
Claude Opus 5.558U.S.+20 vs. Large 4
Gemini 4 Argon53U.S.+15 vs. Large 4
GPT-6.1 Sol52U.S.+14 vs. Large 4
GLM-5.345ChinaZ.ai
Kimi K344ChinaMoonshot AI
DeepSeek V4.1 Flash39ChinaOne point ahead
Mistral Large 4 Preview38FranceText + image API
GPT-6 Luna38U.S.Same score
Command A+13CanadaCohere
Behind Claude Opus 5.520 points

Largest gap to a listed top scorer in this selection.

Behind Gemini 4 Argon15 points

Index-point difference; not a percentage.

Behind GPT-6.1 Sol14 points

Benchmark gap merits workload-specific testing.

Models used differing named reasoning settings. The snapshot does not establish identical reasoning or compute budgets across evaluations. Developer locations identify the organizations, not where API requests are processed.

A preview, not the weight release

The launch and assessment are close together in time. Large 4 was introduced as a public API preview on October 6; the article and benchmark snapshot are dated October 7. Mistral says it is continuing to improve the model.

Next stated milestone

Later in
October

Mistral scheduled a release of model weights later in the month. As of the October 7 assessment, the weights were not publicly downloadable. No precise release date or confirmed outcome is provided in the source material.

What Mistral says

Scale and European infrastructure

Large 4 is described as a mixture-of-experts model with one trillion total parameters and 49 billion active parameters. The preview API accepts text and images. Mistral says it trained the model on its own infrastructure in Europe.

What scale does not tell us

Capacity is not task accuracy

A reported context capacity of roughly 512,000 tokens describes how much input a request can hold. It does not show how accurately the model reasons across that material or completes a multistep assignment.

What the preview score cannot show

The index is an aggregate benchmark result. The source does not provide task-level evidence for the author’s assignments, a controlled hallucination comparison, or complete cost figures.

Personal assessment

Agentic work needs task evidence

Thorsten Meyer says he would not choose the current preview for demanding agentic work or long tasks when stronger-scoring alternatives are available. He also reports encountering hallucinations in his own use. These are personal observations, not a controlled comparison or a quantified error rate.

Evidence limits

Several decision inputs remain open

  • No task-level comparison for the author’s workflows
  • Different reasoning settings in the cited model snapshot
  • Incomplete cost discussion; no cost-per-task figures
  • Performance may change before the planned weight release
Practical takeaway

Test complete workflows before relying on the preview

For extended work, evaluate accuracy, supervision needs, tool use, and cost on your own tasks. An unsupported assumption early in a multistep run can affect later actions, even when the final response reads smoothly. The supplied evidence does not establish parity—or the best choice for every developer.

From announcement to a fuller decision

The evidence available on October 7 supports an early view, with more evaluation needed before drawing conclusions about specific workloads.

01

Preview launch

October 6 · Public text-and-image API

02

Early snapshot

October 7 · Intelligence Index score of 38

03

Weights planned

Later in October · Not yet downloadable

04

Evaluate your work

Compare accuracy, oversight, and full-task cost

Key questions

01 · Announcement

What did Mistral announce?

Large 4 entered public API preview on October 6, 2026. It is a text-and-image mixture-of-experts model with one trillion total parameters and 49 billion active parameters.

02 · Score

How did it score?

Artificial Analysis gave Large 4 Preview an Intelligence Index score of 38 in the October 7 snapshot. Evaluations used differing reasoning settings.

03 · Availability

Are the weights available?

Not according to the October 7 source. Mistral scheduled a release for later in October, without a precise date stated here.

The Gap in Frontier Benchmarks

The score places Large 4 below the highest-ranked models in the supplied comparison. The gaps are 20 index points behind Claude Opus 5.5, 15 behind Gemini 4 Argon and 14 behind GPT-6.1 Sol. These are differences on Artificial Analysis’s index; they are not percentages, and the source does not establish that they translate directly into a particular difference in task success.

For developers choosing a model for extended work, the comparison is a reason to test performance on their own tasks before relying on the preview. Agentic workflows can involve planning, tool use and carrying decisions through multiple steps. An unsupported assumption early in a run may affect later actions even if the final response reads smoothly. Meyer’s view is that stronger aggregate scores give him more reason to start elsewhere for complex autonomous work, while his hallucination reports add a practical concern about supervision.

The result also bears on Mistral’s position as a European AI developer. The company’s stated use of its own European infrastructure is relevant to the region’s AI capacity. It does not, by itself, establish parity with higher-scoring competitors or show that the preview is the best choice for a particular developer. Mistral’s advertised strengths in agentic coding and professional tasks need workload-specific evidence to answer that question.

Amazon

AI developer API testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Preview, Not the Weight Release

The development has two separate dates. Mistral introduced the preview API on October 6, 2026; Meyer’s article and the Artificial Analysis score snapshot are dated October 7. The comparison is therefore an early view of a product still described as a preview, rather than a final evaluation of a released model with publicly downloadable weights.

Artificial Analysis reports a context capacity of roughly 512,000 tokens for Large 4. A context window indicates how much input a request can hold. It does not, on its own, show how accurately a model reasons over that material or how reliably it completes a multistep assignment. Meyer makes that distinction when weighing the model for long tasks.

The comparison includes developers based in the United States, China, France and Canada. The locations identify the developers, not where any particular API request is processed. Canada’s Cohere Command A+ scored 13 in the same snapshot, below Mistral’s 38. That result qualifies any broad claim that all competing major labs are ahead. The narrower conclusion supported by this table is that Mistral trails the listed leading U.S. systems and the listed higher-scoring Chinese alternatives.

“I would not choose it for demanding agentic work or long tasks when stronger models are available.”

— Thorsten Meyer, ThorstenMeyerAI.com

Amazon

large language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Preview Score Cannot Show

The index score is an aggregate benchmark result, not a direct test of reliability on every coding, research or business workflow. The source gives no task-level results showing how Large 4 compares with the other models on the author’s own assignments. Its table also uses different named reasoning settings, and it does not describe those evaluations as having identical compute budgets.

Meyer’s account of hallucinations is a report of personal experience. It does not quantify how often unsupported output occurs, compare models under controlled conditions or establish that other models have stopped making such errors. The source material’s cost section is incomplete, so the relative cost per task cannot be stated from the information provided. It also remains unclear how much the model may change before the scheduled weight release.

Amazon

AI model benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Further Evaluations

Mistral has scheduled the model weights for release later in October. That release is the next stated milestone; as of the October 7 article, it had not happened. The source does not give a precise release date or describe any conditions attached to access.

For developers, a fuller decision will depend on results from the released model and testing against their own requirements, including the accuracy, supervision and cost of complete workflows. Mistral’s continuing improvements may change the preview’s performance, but the supplied material provides no later score or confirmed release outcome. Until then, the available evidence is the API preview, the dated benchmark snapshot and Meyer’s explicitly personal assessment.

Amazon

AI hallucination detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Mistral announce?

Mistral introduced Large 4 in public API preview on October 6, 2026. It is a text-and-image mixture-of-experts model with one trillion total parameters and 49 billion active parameters.

How did Mistral Large 4 score?

Artificial Analysis gave Large 4 Preview an Intelligence Index score of 38 in the snapshot cited on October 7. The listed models were evaluated at differing reasoning settings, so the scores are not a comparison under identical compute budgets.

Are the model weights available?

Not according to the October 7 source. Mistral scheduled a release of the weights for later in October, but the source gives no exact date.

Does the score prove Large 4 will fail at agentic tasks?

No. The index is aggregate benchmark evidence, not a direct prediction of performance on a specific workflow. Meyer recommends starting with stronger-scoring alternatives for complex autonomous work, while saying that task-specific testing is needed.

What remains unknown about its cost and hallucinations?

The supplied article text cuts off before giving cost figures, so it does not support a cost comparison. Meyer reports hallucinations from his own use, but provides no controlled frequency comparison across models.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

World Model Readiness: Are You Ready for AI That Acts?

Assess your organization’s readiness for AI systems capable of prediction and action with the new World Model Readiness diagnostic, as industry shifts toward autonomous AI.

Consumer Safety And CRISPR: Tackling ‘Undruggable’ Cancers With Precision Medicine

New CRISPR-based method selectively destroys difficult-to-treat cancers, opening pathways for safer therapies. Development is in early testing stages.

The Critical Timeline Of The Frontier Lab AI Breach In July 2026

A detailed account of the July 2026 AI security breach involving OpenAI models, highlighting how the agent escaped sandbox, reached Hugging Face systems, and what remains uncertain.

Students at Kent State Get Real-World Exposure to Artificial Intelligence Tools.

The students at Kent State gain real-world AI experience with tools like TensorFlow and PyTorch, opening doors to exciting industry opportunities—discover how they achieve this.