Short answer: On 6 October 2026 Mistral launched a public preview of Mistral Large 4, nicknamed le Chonk. The preview is an API (a service you call over the network). The Mistral Large 4 open weights (the file of numbers that is the model) are not public yet. The launch post says 1 trillion parameters, 49 billion of them active, and “weights drop end of this month.” The docs card, read on 7 October, says 1.05 trillion total, 52 billion active, a 1.6 billion vision encoder, and a 1 million token context. A legal PDF from the same day says the on-premise license is confidential, and that the document will be updated “in case of” an open-weights release. There is no public file to download yet.

The number to remember is not 1 trillion. It is the gap between a badge that says Open and a file you can put on a disk.
What shipped on October 6?

A preview API, a nickname, and a promise. Not a download.
A model, here, is a huge list of numbers plus a rule for turning text, and sometimes a picture, into the next chunk of text. A parameter is one of those numbers. A weight is the same thing, said as a file: the stored numbers. An API is a way to send a request to someone else’s computer and get the answer back. You do not receive the numbers.
Mistral’s launch post, dated 6 October 2026, says they are launching a public preview. The unofficial short name is ML4. The official nickname is le Chonk. The post says you can try the preview API today on Mistral Studio, and that weights drop at the end of this month. The same post says the model “pushes the frontier of open-weight performance” and, later, that the point of the work is “delivered through open weights.” Open-weight means the number file is published so you can run it yourself. It does not, by itself, name the license (the rules for copying and changing it).
On 7 October the Hacker News front page had this story second, past 1,600 points, with close to a thousand comments. The item is 49977979. That is attention. It is not a file.
Posts on X the same morning already said the model runs on a desktop. That sentence is false today, because there is nothing to run. The memory section below is why it stays false for one consumer graphics card after the file exists.
| Question | Answer on 7 October | Where it is stated |
|---|---|---|
| Can I call it? | Yes, as a preview | Launch post, and the docs card |
| Can I download the weights? | No | Launch post: end of this month |
| Is the license named? | No | Docs card has no license link |
| Model id on the card | mistral-large-4 | Docs card, version 26.10 |

What is a parameter, and why do the pages disagree?
A parameter is one stored number. “Active” means the numbers used for this chunk of text. The pages do not agree on the count.
Start from the small thing. The model stores numbers. To answer, it does arithmetic on some of them. Total parameters means every stored number. Active parameters means the ones used for this token. A token is a chunk of text, often smaller than a word. “Hello” might be one token. A long word might be two or three.
The launch post says: “ML4 is a 1 trillion-parameter natively multimodal model with 49 billion active parameters.” Natively multimodal means pictures go in as pictures, not as a separate model glued on later. The post also says it was “trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs” in Mistral’s own European data centers, and that the public preview is served on that same hardware. A GPU is a chip that does this kind of arithmetic quickly. This post uses 3,800, the blog’s number. Some news reports round it to about 4,000.
The docs card, read on 7 October, says something close and not the same: “52B active parameters and 1.05T total parameters, and a 1.6B vision encoder.” A vision encoder is the part that turns a picture into numbers the rest of the model can use. The card lists context as 1M. Context, or the context window, is how many tokens one request can look at.
Forty-nine plus 1.6 is 50.6, not 52. Adding the vision encoder to the blog’s active count does not land on the card’s active count. There is no point inventing a reason. Mistral’s own post says further details on the architecture will be shared as they work toward the weights. Until that note exists, 49 billion and 52 billion are two sentences, not one fact.
Artificial Analysis, in its 6 October article, repeats the blog’s pair: 1 trillion parameters, 49 billion active. It lists the context window as 512k tokens. The docs card says 1 million. Those are two published figures, not a measurement.
| Spec | Launch post | Docs card | Anyone else |
|---|---|---|---|
| Total parameters | 1 trillion | 1.05 trillion | Legal PDF: greater than 1 trillion |
| Active parameters | 49 billion | 52 billion | Artificial Analysis: 49 billion |
| Vision encoder | Not in that sentence | 1.6 billion | Not stated |
| Context | Not stated | 1 million tokens | Artificial Analysis: 512k tokens |
| Training GPUs | 3,800 Grace Blackwell | Not on the docs card | Press often writes about 4,000 |
What is a mixture of experts?
A library of specialists. Only a few open for this token. All of them still have to sit in memory.
A dense model uses every parameter for every token. A mixture of experts, an MoE, replaces one big block with many smaller blocks called experts. A router, a small function that picks, chooses a few experts for this token. Those run. The others wait. They wait in memory, because the next token might need them.
So “1 trillion parameters” is the size of the library. “49 billion active,” or 52 billion, is the size of the committee that votes on this word. Forty-nine divided by 1,000 is 4.9 percent. Fifty-two divided by 1,050 is about 5.0 percent. The compute, the work per token, tracks the committee. The file tracks the library. That is why an API can be priced like a mid-size model while the download, once it exists, is enormous.
The docs card calls the architecture a granular mixture of experts. Granular, here, means the experts are sliced finer than the older “a handful of giant experts” designs. Mistral has not published an architecture diagram yet. Their post says the architecture details come later.

Why a laptop still cannot hold it
Active parameters set the work. Total parameters set the file. A 24 GB graphics card is the wrong shelf.
Here is arithmetic, labeled as ours, not as a Mistral spec. Store each parameter as a 16-bit number, which is 2 bytes. A byte is 8 bits, the usual small unit of storage. 1.05 trillion times 2 is 2.10 trillion bytes, about 1.91 tebibytes. People say “about 2 TB.” Quantizing means storing each number in fewer bits, which cuts the file. Mistral has not published the serving format, so there is no smaller number here to call theirs.
Mistral’s own card for the older Large 3, a 675 billion parameter model, lists a GPU memory range of 1,800 to 360 GB. Read that as the scale of a data-center machine, not a desktop. Large 4 is bigger. A post that says le Chonk runs on your rig is selling a file that is not out, at a size a single consumer GPU does not have.

What does “open” mean on this card?
Large 3 is a worked example of the word. Large 4, today, is the word without the file.
Open source, in the usual software sense, means a license that lets you use, change, and share, including for money, with a short list of duties. Apache 2.0 is one such license. Open-weight is narrower: the numbers are published. The license can still say no.
Mistral Large 3 is the comparison that makes the badge honest. The December 2025 launch post says Large 3 has 41 billion active parameters and 675 billion total, and that all of those Mistral 3 models were released under Apache 2.0. The docs card still says Apache 2.0, links the weights, lists a 256k context, and prices the API at $0.50 per million input tokens and $1.50 per million output tokens. That is a file you can download, under a license you can read.
Large 4’s docs card says “Public Preview” and “Open” and uses the words open-weight. The card has no license link and no weights link of the Large 3 kind. The launch post does not name a license.
The 6 October technical PDF for downstream providers, version 1, is the third page. It lists the architecture as a mixture of experts and the parameter range as greater than 1 trillion. It says on-premise use, meaning on the customer’s own machines, sits under a confidential self-deployment agreement. It says the document “will be updated in case of open weights release.” “In case of” is weaker than “we will.” The blog’s sentence is the stronger one: weights at the end of this month. Neither sentence cancels the other. The operational fact today is that no public weight file is posted.
VentureBeat, citing company materials, reports a plan to publish the weights on 27 October, and says the weights are expected under a custom Mistral license. That date is not in the launch post. The blog says “end of this month.” Treat 27 October as a press report of the company’s plan, and treat the license as unnamed until a license file exists.
One cell in that same PDF is a trap if you take it as the product. The modalities table lists a text maximum of 250 tokens and an image maximum of 8. That is not a 250-token cap. The docs card says 1 million tokens. If those PDF cells are a template that was not filled in, they should be fixed. Do not use that PDF as your context-window source.
The training-data note in the PDF is worth keeping, because “open” is also a question about what went in. It says the mix includes public internet text and images, non-public data licensed from others, synthetic data made inside Mistral, and user input and output from products such as Vibe and Mistral AI Studio, with an opt-out. Synthetic means made by a program, not copied from a person. This is not a grade of that mix. The point is that the PDF talks about it, and the launch post does not lead with it.
Why does the cyber score look like a win over Claude?
Because a refusal can be scored as a zero. Mistral says so, in its own post. The independent index is a different number.
Cybersecurity, here, means finding and fixing flaws in software. A benchmark is a fixed test with a score. A refusal is the model answering “I won’t do that” instead of attempting the task.
Mistral’s post says that on the Artificial Analysis Cyber Index, an independent test of finding and fixing flaws in real software, ML4 “ranks among the top five models globally and leads open-weight models developed outside China by a wide margin.” On one test in that index, reproduce a real vulnerability in open-source software and then patch it, “ML4 scores 82%, the highest of any model.” It “solves 93% of the challenges in Cybench,” 40 exercises from security competitions.
Then the sentence that changes how you read the 82. “Several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the same test because they refuse to perform the task.” (For the GPT-6 Astra backstory, see why OpenAI killed GPT-6.1 Astra.) A zero from a refusal is not a measurement of skill. It is a measurement of a policy. Mistral’s pitch is that a lab can turn a closed model off, and that defending software often starts by proving a flaw is real.
Artificial Analysis’s own 6 October article puts a second number next to that pitch. ML4 scores 50 on the Cyber Index, level with GLM-5.3-Flash at 50, and behind MiMo-V2.6-Pro at 56. The same article says the 82% result, on the test they call CyberGym-E2E-AA, is ahead of MiMo-V2.6-Pro at 79% and GPT-6 Luna (max) at 78%. So “highest of any model” can be true of one subtest, and “first on the index” can be false, in the same week, without anyone lying. Read the noun. Subtest is not index.
There is a second refusal sentence, and it points the other way. Despite the cyber scores, “the average refusal rate of the model on cyber prompts from JailbreakBench, StrongREJECT, and AgentHarm is higher than all OSS models.” OSS means open-source software, used loosely here for open models. A high exploit score and a high refusal rate on malicious prompts are both in the post. They are not the same claim.
And the public API is not the less-moderated build. The post says that until the weights drop, Mistral is red-teaming (asking a system to try to break itself) with cybersecurity leaders, vetted partners, and state authorities, “who will access the same model with reduced moderation and expanded cyber capabilities.” If you sign up for Studio, you get the public preview, not that build.

Where does 38 sit?
It is a large jump from Mistral’s last models. It is not the top of the independent chart.
An index, here, is one number built by adding several tests. Artificial Analysis’s 6 October article says Mistral Large 4 scores 38 on its Intelligence Index, in a research public preview, comparable to GPT-6 Luna (max) at 38 and DeepSeek V4.1 Flash (max) at 39. Because the weights are not out, that article treats the preview as proprietary (closed, in the sense that you cannot take the file). “Most intelligent model from outside the US and China” is their framing of that 38. It is not “most intelligent model.”
Secondary writeups of the same index, the same week, put Claude Opus 5.5 (Max) at 58, and put Chinese open models ahead of 38: MiMo-V2.6-Pro around 46, GLM-5.3 around 45. The same writeups put Mistral’s own previous steps far below: Medium 3.5 around 14, Large 3 around 9. The drawing uses those rounded figures and labels them as second-hand. If you only want numbers from Artificial Analysis’s own sentences, keep 38, the Luna tie at 38, the Flash comparison at 39, the Cyber Index at 50, and MiMo’s 56 on that cyber index.
Mistral’s own coding numbers, which are not that index, are still worth having as vendor claims. DeepSWE v1.1: 61.7%. SWE-Atlas-QnA: 59.4%. Terminal-Bench 4: 28.3%. A combined Coding Agent Index of 49.8%, which they say is ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max. A blind human review by Surge AI, identities hidden, scored ML4 second of five models at 3.74 out of 5, behind Claude Opus 5 at 4.22, and ahead of Kimi K3 at 3.59, GLM-5.3 at 3.60, and GLM-5.2 at 3.40. Read that last set as Mistral telling you it did not win its own coding taste test against Claude Opus 5. Opus 5 and Opus 5.5 are different names in their post. They are not merged here.
On pictures, they say ML4 surpasses GPT-6-Astra on a test called Dense 200, 42% versus 41%. A one-point edge on one visual test is not “better than GPT-6 at seeing.”

What should you do this week?
Call the API if you have a real task. Do not budget a cluster for a file you cannot see. Recount the spec the day the file appears.
The price on the card and on the pricing page, per million tokens: input $1.36, cached input $0.14, output $4.18. A cache is a stored copy of a repeated prefix, so the provider does not recompute it. The same pages show a sale at half: $0.68, $0.07, and $2.09. Artificial Analysis describes that half price as the first two weeks. The pricing page shows the sale but does not say “two weeks,” so that duration is Artificial Analysis’s description, not a line from the price table.
A million tokens in and a million tokens out is $5.54 at list, and $2.77 at the sale. That is the bill for volume, not a promise about quality. Large 3 remains $0.50 and $1.50, and you can download it. It is the older, smaller model. Cheaper and owned is sometimes the right trade. It is a different trade from “the new one is open.”
On the Hacker News thread, Simon Willison reported that the preview only offers reasoning effort “none” or “high,” and that high did not clearly spend more output in a quick check. Reasoning effort is a switch for how long the model works before the answer. If you try the API, look at that switch first, and do not assume “high” means a long hidden scratchpad.

- Do call
mistral-large-4on a task you already know how to grade. A graded task is one where you can tell a wrong answer without asking the model. - Do keep two cyber numbers in your notes: 82 on the reproduce-and-patch subtest, 50 on the Cyber Index.
- Do wait for a license that has a name. Apache, a custom license, and no file are three products.
- Do not download it. There is no public weight file.
- Do not size a laptop for it. About 5 percent of the numbers work on a given token. About 100 percent of them have to be stored.
- Do not treat Studio as the reduced-moderation cyber build. That build is for vetted partners and state authorities, until the weights, and maybe under a different rule after.
If you read this after the weights land, recount the docs card, the license file, and the context the API actually accepts. The 49 versus 52 gap, and the 1 million versus 512k gap, are facts about the pages on 7 October. They are not a permanent spec.
Common questions about Mistral Large 4
Can I download the Mistral Large 4 open weights?
Not yet. On 7 October 2026 there is no public weight file. The launch post says weights drop at the end of this month, and VentureBeat reports a planned 27 October release under a custom Mistral license.
What license will Mistral Large 4 use?
None is named. The docs card has no license link, and the downstream-provider PDF calls the on-premise agreement confidential. Large 3, by contrast, ships under Apache 2.0.
How big is Mistral Large 4?
The launch post says 1 trillion parameters with 49 billion active. The docs card says 1.05 trillion with 52 billion active, plus a 1.6 billion vision encoder. At 16 bits per parameter, 1.05 trillion parameters is about 2 TB of weights.
How much does the Mistral Large 4 API cost?
List price per million tokens is $1.36 input, $0.14 cached input, and $4.18 output. The launch sale halves that to $0.68, $0.07, and $2.09.