Why Mistral Large 4 Is Worth A Look Beyond The US And China
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why Mistral Large 4 Is Worth A Look Beyond The US And China on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Mistral Large 4 scored 38.4 on Artificial Analysis’s Intelligence Index, making it a leading model from outside the US and China but leaving it below current US and Chinese flagships in the supplied comparison. Its progress from earlier Mistral models is substantial, yet benchmark performance, per-task costs and still-unpublished weight terms leave its practical value unsettled.

French AI company Mistral released Large 4 in a research public preview, and Artificial Analysis’s Intelligence Index v4.3.2 gives it a score of 38.4. The result makes it a strong contender among models from outside the United States and China, but the cited benchmark puts it below the leading US and Chinese models—and its reported price per benchmark task is higher than that of two Chinese models that score above it.

Large 4 is a one-trillion-parameter model with 49 billion active parameters, according to the source. It accepts text and images, produces text, and supports a 512,000-token context window. Mistral has made it available through its API as a research public preview. The company says it plans to release the model weights at the end of October; until then, the model is proprietary and its licence has not been published, the source reports.

On the cited Artificial Analysis index, Large 4 scored 38.4, compared with 9 for Mistral Large 3 and 14 for Medium 3.5 on the same index version. That marks a sharp improvement for Mistral. But the supplied comparison lists US models at the top, including Anthropic’s Claude Opus 5.5 at 57.6 and OpenAI’s GPT-6 Astra at 52.7. Chinese models also score above Large 4 in the table, including GLM-5.3 at 44.8 and DeepSeek V4.1 Flash at 39.5.

The source gives API prices of $1.36 per million input tokens and $4.18 per million output tokens, with cached input at $0.14 per million tokens. It says Mistral offered a 50% discount for the first two weeks. Artificial Analysis’s task-cost comparison puts Large 4 at $1.13 per benchmark task, against $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash. Those figures apply to the benchmark’s tasks; they do not establish what every customer will pay for a different workload.

At a glance
analysisWhen: Released yesterday, according to the so…
The developmentMistral released Large 4 in a research public preview, with an independent benchmark placing it at 38.4 and below the top US and Chinese models.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$85,161▼ 0.6%
Ethereum ETH$2,683▼ 0.9%
Tether USDT$0.9999▼ 0.0%
BNB BNB$775.58▼ 0.9%
XRP XRP$1.49▼ 0.7%
USDC USDC$0.9999▼ 0.0%
Solana SOL$119.75▼ 0.7%
TRON TRX$0.3351▼ 0.3%
Live data · CoinGecko · alternative.me (24h change)
Mistral Large 4: Not a Frontier Model — Reality Check
AI Dispatch · Reality Check · 7 October 2026

Mistral Large 4: best outside the US and China — and still not a model to run your agents on

The headline is true: France has the most intelligent model outside the US and China. The independent data says the rest: every US and Chinese flagship scores higher, the best by 19 points. It costs 4× more per task than Chinese open models that outscore it, and it’s 2.5× as verbose as the median model.

Artificial Analysis Intelligence Index v4.3.2 — same version, like for like
Claude Opus 5.5 US57.6
Claude Sonnet 5.5 US56.0
Claude Fable 5.1 US53.4
GPT-6 Astra US52.7
Gemini 4 Argon US52.6
GPT-6.1 Sol US51.8
GLM-5.3 CN · open44.8
Kimi K3 CN · open43.6
GLM-5.3-Flash CN · open41.8
DeepSeek V4.1 Flash CN · open39.5
Mistral Large 4 (Preview) FR38.4
GPT-6 Luna US · small model~38
DeepSeek V4 Pro 0813 CN36.0
GLM-5.2 CN33.7
vs US frontier
−19.2 pts

~two-thirds of Opus 5.5. Level with OpenAI’s small model, Luna.

vs China open
8th

Eighth among open models once weights ship — behind seven Chinese ones. Beats GLM-5.2 and V4 Pro, loses to their successors.

vs Canada
n/a

Cohere doesn’t compete at this tier — reported ~14% hallucination at ~9% accuracy, because it declines most questions. A field of one.

The cost problem is worse than the intelligence problem — $ per Index task
Mistral Large 4
$1.13
Index 38.4 · $0.57 launch promo
GLM-5.3-Flash
$0.25
Index 41.8 · 4.5× cheaper
DeepSeek V4.1 Flash
$0.27
Index 39.5 · 4.2× cheaper
Gemini 4 Argon
~$1.99
Index 52.6 · +14 points
Per-token pricing looks competitive ($4.18/M output, well under the $10 median) — but it burns 200M output tokens on the Index vs an 81M median. Cheap tokens × 2.5 as many tokens is not a cheap model.
Why not for agentic or long-running work
The gap compounds
19 pts behind

The Index is now agentic-heavy — Briefcase, GDPval, AutomationBench, Terminal-Bench. Errors multiply across steps: tolerable in chat, fatal over a two-hour run.

AA v4.3.2
Verbosity
200M vs 81M

Output tokens to complete the Index. On an agent, verbosity is cost and latency on every step.

AA
Hallucination is back
observed

Confident false assertions in hands-on use. US frontier has largely moved past this — Gemini 4 Argon: 15%. In fairness Chinese open models are worse (Kimi K3 51%, DeepSeek V4 Pro 94%). In an agent, a fabrication is a wrong premise every later step builds on.

AUTHOR’S TESTING · not an AA figure
✓ What it’s genuinely good at
  • Cyber defence: 50 on the AA Cyber Index; 82% CyberGym-E2E (ahead of Luna’s 78%). Likely top-3 open model on cyber.
  • Documents & images: 19% GDP.pdf (+18 vs Large 3); 100 images per request.
  • Speed: 116 tok/s, 1.46s TTFT — well above median.
  • The jump: Large 3 scored 9 on this Index. 9 → 38 is real progress.
  • Jurisdiction: French parent, EU hosting, weights promised end of October.
▸ Who should actually use it
  • Legally bound buyers (defence, classified, DORA, health data): now the best European option by a wide margin. Wait for the weights, check the licence, pilot on cyber and documents.
  • Everyone else, for agentic or long tasks: don’t. A US frontier model is meaningfully more capable; GLM-5.3-Flash is more capable and 4× cheaper.
  • Note: Preview — Mistral says RL is still running, so scores may move. That changes next month’s decision, not today’s.
The take

Mistral says it has “essentially closed the gap.” It has closed the gap to where the Chinese open-weights field was a few months ago, while that field and the US frontier have both moved on. On every independent measure that matters for agents — intelligence, cost per task, verbosity and factual reliability — Large 4 is not a frontier model. “Most intelligent outside the US and China” is true mainly because almost nobody else outside those two countries is competing. Use it if you have to. Don’t use it because of the headline.

Sources: Artificial Analysis — Mistral Large 4 article & model/provider pages (6 Oct 2026), Index v4.3.2, comparison data; Trending Topics independent-ranking analysis; AA-derived reporting for frontier scores and AA-Omniscience rates (Argon 15%, Kimi K3 51%, DeepSeek V4 Pro 94%); Cohere profile as reported by Suprmind. Mistral Large 4’s AA-Omniscience result isn’t published in text — the hallucination point is the author’s own testing. Preview scores may change. Not investment advice.
thorstenmeyerai.com

Performance Meets Procurement Costs

The result matters because buyers evaluating AI models need to weigh capability, cost and operational fit, not just a model’s national origin or its headline benchmark position. Large 4’s move from earlier Mistral scores suggests the company has made substantial progress. At the same time, the comparison supplied here does not place it at the frontier of the tested models, and the reported task costs complicate the case for choosing it on value alone.

That trade-off may be especially relevant for multi-step agent workflows, where a model must sustain a sequence of decisions rather than answer a single prompt. The source reports that Large 4 generated 200 million output tokens across the Intelligence Index tasks, compared with a median of 81 million for comparable models. If those measurements apply to a buyer’s workload, greater output volume can add cost and time. Benchmark results, however, cannot by themselves predict performance in every company’s systems.

The source author also reports seeing confident false statements in hands-on testing. That is an attributed observation, not a finding from the cited index. It is a reason for prospective users to test reliability and verification requirements in their own environment, rather than assume that a high-level benchmark score answers those questions.

Amazon

AI language model API access

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Fast Jump From Mistral 3

The supplied account compares Large 4 with earlier Mistral models using the same version of Artificial Analysis’s index: Large 3 scored 9 and Medium 3.5 scored 14, while Large 4 scored 38.4. That makes the release a major step up within Mistral’s own lineup. It does not erase the gap shown in the wider table, where the highest-listed US model scores 57.6 and several Chinese models also exceed Large 4.

The index described in the source draws heavily on agentic and work-oriented evaluations, including AA-Briefcase, GDPval-AA, AutomationBench and Terminal-Bench 4.0. Its score is therefore a measure across those listed tests, not a complete assessment of every use case, such as simple chat, image understanding or a specific company’s retrieval system. The source also says Mistral’s reinforcement-learning work is ongoing and that scores may change.

The article’s claim that Large 4 is the most intelligent model outside the US and China should be read in that limited frame. It reflects the rankings and comparisons supplied by the source, not proof that the model is best for every task or that no other developer could compete under a different test.

Amazon

large parameter AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Preview Terms and Reliability

Several practical details remain unsettled in the supplied material. The weights are not yet available, and the model’s licence has not been published. The source says Mistral plans to release the weights at the end of October, but gives no further detail on access conditions or whether that schedule will hold.

The reported benchmark and pricing data also do not settle how Large 4 will perform or cost in a particular deployment. Actual expenses depend on a user’s token volumes, prompt and output patterns, and the applicable price at the time. Mistral’s ongoing reinforcement learning may alter benchmark results, according to the company as quoted in the source.

Finally, the source author’s account of hallucinations comes from hands-on testing, without details here about the number of tests or the testing protocol. Artificial Analysis’s index score is separate from that observation. The material provided does not establish a broad, independently measured hallucination rate for Large 4.

Amazon

AI model cost comparison

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Weights and Further Testing

The next stated milestone is Mistral’s planned release of Large 4’s weights at the end of October. Their publication, alongside licence terms, should make it clearer whether developers can run or adapt the model outside Mistral’s API and under what conditions. The timing and terms remain subject to confirmation.

In the meantime, Mistral’s research preview is available through its API, and the company says reinforcement learning is continuing. Buyers considering the model can compare updated benchmark figures and test their own workloads, including output volume, accuracy and performance over multi-step tasks. The supplied source does not identify a later benchmark date or a confirmed production release schedule.

Amazon

text and image AI processing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Mistral Large 4?

It is a Mistral model available in a research public preview through the company’s API. The source describes it as a one-trillion-parameter, natively multimodal text-and-image input model with a 512,000-token context window.

How does Large 4 compare on the cited benchmark?

Artificial Analysis’s Intelligence Index v4.3.2 scores it at 38.4. The supplied table lists several US and Chinese models above it, including Claude Opus 5.5 at 57.6 and GLM-5.3 at 44.8.

Is Large 4 open-weight now?

No. The source says it is currently a proprietary preview. Mistral plans to release the weights at the end of October, but the licence has not been published in the supplied material.

How much does the API cost?

The source lists prices of $1.36 per million input tokens, $4.18 per million output tokens and $0.14 per million cached input tokens. It also reports a 50% discount for the first two weeks; customers should check current pricing because the discount period and rates may change.

Do the benchmark results prove Large 4 is unreliable?

No. The source author reports seeing confident false statements during hands-on testing, but the provided account does not include a test protocol or a broad, independently measured hallucination rate for Large 4. Buyers should treat that report as an attributed observation and evaluate the model on their own tasks.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The clause. How a contractual definition of AGI met the capital built on top of it.

An analysis of how the contractual definition of AGI in the Microsoft-OpenAI agreement was renegotiated, shifting from a doomsday trigger to a verification process.

Artificial Intelligence 101: Understanding the Basics of AI

Incredible insights into Artificial Intelligence 101 reveal how machines learn and adapt, but understanding its core concepts is essential to grasp its true potential.

Encourage Trust in AI to Improve Long-Term Significance

AIThis post was created with the assistance of artificial intelligence (AI).In our…

The gigawatt gap. Why China is structurally positioned for AI power and the US is engineering around its grid.

China’s centralized infrastructure and renewable buildout give it a structural edge in AI power deployment, challenging US dominance at the physical energy layer.