You cannot open Gemini and try Gemini 4 Argon today. Google and a vetted set of cyber defenders can. Everyone else is waiting on a date Google has not set. The rest of this page is what that closed door actually means, starting from a computer that only follows instructions. Google announced the model on 30 September 2026. By the next morning it was the top story on Hacker News and in the tech press. Most of those pieces hand you a scoreboard. A score is useless until you know what was tested, who was allowed to take the test, and what you are supposed to do on Thursday. Argon is a strong, gated cyber model with a real habit of saying “I don’t know.” It is not a model you can call, it is not sole first place on the public patching exam, and the hospital flaw in the announcement has no name and no confirmed fix.
Start with a computer that only obeys

Before any new product name, start with the machine. A computer does not understand a sentence the way you do. It follows a list of instructions, one step after another, even when a step is wrong. That list is software (instructions written so a computer will follow them). A bug (any wrong step, including a button that is the wrong color) is still followed. The computer does not pause to ask if the step was a good idea. The kind of bug that matters here has an extra property. Someone who is not supposed to be there can use the wrong step to get in, change data, or read what they should not. That bug is a vulnerability (a mistake an outsider can use). A spelling error on a screen is a bug. A medical record that is sent to anyone who asks is a vulnerability.

The left box is the instruction. The middle box is a stranger using it. The right box is the word for that pair. A person who protects systems, a defender (someone whose job is to find holes and close them before a criminal does), already works in three beats. Find the hole. Prove the hole is real. Change the code so the hole is gone and the rest of the program still works. Hold that order. The product name comes after it.
What a model is, before the name Argon
A model, in this article, is not a person and not a search box. It is a very large pile of numbers. Training (showing the pile so many examples that the numbers settle into patterns) is how those numbers get set. After training, you give it text and it guesses the next chunk, then the next, and so on. It can look like it “knows” a hospital. What it has is patterns from text and code it was shown. The lab that built this one is Google DeepMind (Google’s AI research group). The thing they shipped on 30 September is Gemini 4 Argon, written up by Koray Kavukcuoglu, who leads that group. A frontier model (a model at the front edge of what the big labs can currently build, not a small cheap helper) is the class they put it in. Argon is their first model in that heavier class in more than seven months. The cheaper, faster line they have been shipping is called Flash. Argon sits above Flash.

The introductory price is $2 per million input tokens and $10 per million output tokens. Google’s footnote says that after the introductory period the price becomes $4 and $20. The footnote does not give the end date. Do not invent one. Artificial Analysis, an independent lab that times and scores models, priced one of its standard hard tasks at about $1.99 on the discount, versus $3.26 for OpenAI’s GPT-6 Astra. At the later $4 / $20 rate, the same task is about $3.98, roughly 1.2 times Astra. The discount is the headline. The footnote is the plan.
Why the door is shut
Google’s own post says a model this capable needs a phased release (a slow opening, not a public switch). They are in a voluntary U.S. government process that lets agencies look at a frontier model before a wider release. They will tune guardrails using early testers, then open Argon to developers, companies, and consumers “as soon as possible.” There is no calendar date on that sentence. The door that is open has a name. The Fairwind Program (Google’s list of trusted defenders who get cyber models early) started in early September 2026 with a smaller model, Gemini 3.8 Flash Cyber. It now has more than 650 partners. The people they prioritize are governments and national cyber agencies, operators of critical infrastructure (the systems a country cannot casually lose: hospitals, phones, power, banks), core technology platforms, and universities that only want to measure defenses. Applicants get a background check. Access is for the internal security team, the incident team, or the penetration testers (people paid to break in with permission). Login needs phishing-resistant MFA (a second proof of identity that a fake email cannot steal). They may not share, resell, or redistribute the model.

Fairwind partners, and Google’s own staff, get Argon without cyber guardrails (the refusal rules that normally make a model say no to hacking help). Everyone else, when the door opens, is supposed to get the version that still refuses. A guardrail is not a lock on your account. It is a behavior trained and checked into the model: refuse this kind of request. Taking the cyber guardrail off is the controversial part, and it is also the point. A defender often needs the model to think like an attacker, or it cannot rehearse the break-in and then close it. The same rehearsal is what a criminal wants. That double nature is dual-use (one capability, a legitimate use and a harmful use). Fairwind’s rule tries to split those uses. Allowed: authorized threat simulation (a pretend attack, with permission), reverse engineering (taking a program apart to see how it works), and malware analysis (studying harmful software to defend against it), for defense or academic research. Not allowed: creating malware to use for real. “No guardrails” is not “do whatever you want.” It is “we removed the model’s automatic no, and we are trusting your contract, your job, and your login instead.” If you are a normal developer, the annoyed posts on X are correct. There is nothing for you to try yet. Do not rebuild your coding agent (a program that edits files and runs commands, not a chat window) on a model you cannot call.
The job it claims to do
Google says Argon can autonomously (without a person steering each step) find, validate, and patch critical vulnerabilities. Read those three verbs against the human order from the first section. Find means search. That can be the source code, or a live website the model cannot see the inside of. The second kind is a black-box test (you only see what the site sends back, not the code that produced it). Validate means prove it. A proof of concept (a small, controlled demonstration that the hole really works) is the proof. A confident paragraph is not a proof. Patch means the code change. On the public exam below, a patch counts only if the attack no longer works and the project’s existing tests still pass. A fix that closes the hole and breaks the product is a fail.

Google’s claim is that Argon can run all three. The public evidence is an exam score, plus internal stories. It is not a named, shipped fix you can read. The tool wrapper matters as much as the model. A harness (the program around the model that lets it read files, run commands, and try again) is what actually touches the code. Google points Fairwind at a harness called CodeMender for vulnerability research and patching. On the public exam, Argon was scored inside Antigravity (Google’s own coding harness), OpenAI’s models inside Codex (OpenAI’s coding agent), and Anthropic’s models inside Claude Code (Anthropic’s terminal coding agent). You are comparing model-plus-harness, not a naked brain.
The patching exam, without the trophy language
A benchmark (the same exam, given to every model, so the scores can be compared) is the only fair way to talk about “better.” The cyber exam in the announcement is CWE-bench v1, from Collinear AI. CWE (Common Weakness Enumeration, a public catalog of bug types, such as “this program never checks who is asking”) is the family of mistakes the tasks are built from. Version 1 is 120 tasks the models have not been shown as the answer key. Each model gets four tries at each task. A try passes only if two things are true: the old attack no longer works, and the tests the project already had still pass. pass@1 (success on the first try) is the headline number. pass@4 (success on at least one of four tries) is the number that breaks ties. Google says Argon ties for first at 68%. SecurityWeek names the tie: OpenAI’s GPT-6 Astra and xAI’s Grok 4.7 are also at 68%. Claude Opus 5.5, from Anthropic, is one point back at 67%. One point on this exam is about 1.2 tasks out of 120. The gap from first to fourth is about one task. 68% is about 82 tasks fixed on the first try, and about 38 still broken.

Four tries reorder the podium. A 1 October read of the leaderboard puts Grok 4.7 at 81% pass@4, Opus 5.5 at 79%, Argon at 75%, and Astra at 74%. On the headline, Argon is tied for first. On the tie-break Google’s own “ties for first” sentence does not mention, Argon is third, and Opus, a point behind on the first try, passes it. That same read of the leaderboard adds three warnings.
- A second judge, a panel of model graders rather than the automatic checker, scores Argon’s first try at 62%, not 68%. The automatic checker and the panel do not agree.
- Argon’s average cost on that exam is about $6.63 a try. Astra is about $2.85, Grok 4.7 about $2.75, Opus 5.5 about $0.79. Opus is close on the first try and much cheaper. Collinear had already flagged this pattern on 28 September, before Argon was in the table: the leaders were within a point, and they do not even solve the same tasks.
- Nobody outside Google has said whether the Argon that sat the exam was the guardrail-free build defenders get, or the refused-request build everyone else is supposed to get. If it was the open build, 68% describes a version you will not be sold. If it was the closed build, the defenders’ version has no public score.
Collinear’s own description of the exam, the day before the launch, is the method. The Argon row is Google’s 68%, plus the later leaderboard read for the four-try numbers. I am not pretending I reran the 120 tasks.
“It found a hole in hospital software” is not a patch
Wiz (a cloud security company Google owns) is using Argon inside Scan for Good, a program that looks for dangerous exposures in critical public infrastructure and fixes them without charging. Google says that in an early run the model found a critical vulnerability exposing personal information in healthcare software used by hospitals worldwide, and that earlier frontier models missed it.
The other scores, and who held the pencil
Google published a table of 18 benchmarks and, by Decrypt’s count of that table, leads 12, ties 1 (the cyber exam above), and trails 5. One trail that is easy to see: on OSWorld (a test where the model has to operate a computer screen, not just write text), Astra is ahead, 72.6% to Argon’s 69.2%. A 1 October briefing of Google’s methodology says Google itself computed Argon’s score on 9 of the 18 rows. The other 9 come from outside leaderboards. A lead on a row Google scored, with no error bar, is a weaker sentence than a lead on a row someone else runs. Here are the rows worth remembering, in words first. DeepSWE v1.1 is an exam of long, messy software jobs, not a single function. Google reports Argon at 77.9%. Decrypt’s chart puts Opus 5.5 at 74.2%, Astra at 74.1%, and Claude Fable 5.1 at 67.4%. Decrypt also says Google computed Argon’s DeepSWE number, while the rivals’ numbers came from a public board and from company reports. Same exam name, not the same referee. Keep that next to the 77.9. AutomationBench, from Zapier, asks whether a model can finish business tasks end to end. Google says Argon is first at 51.3%. First on this one still means it fails about half the jobs. A score of 51 is not “it can run the company.” LVBench asks whether a model understands a long video. Google says 91.7%, and calls that the best. The Vals Index weights finance, coding, legal, and tax tasks by how large those sectors are in the U.S. economy. Secondary writeups of Google’s table put Argon at 68.9%, Opus 5.5 at 67.0%, and Astra at 63.1%. Close. Harvey’s Legal Agent Benchmark is legal research and drafting. The same table puts Argon at 19.6%, against 5.4% for Astra and 3.8% for Opus 5.5. The gap is large. The level is still low. Nineteen percent means it fails about four times in five. Do not hand it a filing and walk away. The interesting part is that a model which refuses to bluff can be more usable on careful work even when its raw accuracy is not the best. That is the next section. Prompt injection (a hidden instruction buried in a web page, email, or file the model reads, trying to make it ignore you and obey the hidden text) is the attack Google says Argon resists best. Their post says Argon leads Gray Swan’s indirect prompt-injection exam. Decrypt prints the attack success rate (how often the hidden instruction wins; lower is better): Argon 0.7%, Opus 5.5 and Fable 5.1 at 1.0%, Astra at 8.5%. Older models on that chart, Grok 4.6 and Kimi K3, are fooled about half the time. The 0.7% is Decrypt’s number. Google’s sentence is only “leading.”
Less bluffing is not more knowledge
Artificial Analysis ran Argon on AA-Omniscience, a test of whether a model answers facts correctly or makes them up. A hallucination (a confident answer that is wrong) is the failure they count. Argon, on its high reasoning setting (the deepest setting Google offers on this model), has a 15% hallucination rate. That is the lowest among models scoring 45 or higher on Artificial Analysis’s Intelligence Index (their single combined score across many exams, not one test). GPT-6 Astra, on its max setting, is at 51%. GPT-6.1 Sol, on max, is at 54%.

Read the bottom half before you quote the top half. Accuracy (how often the answer is actually right) is 50% for Argon and 63% for Astra. Argon is also five points less accurate than Google’s own older Gemini 3.1 Pro Preview. The combined knowledge score is 42 for Argon and 43 for Astra. They are even. Argon got there by abstaining (saying it does not know) instead of guessing. Astra got there by being right more often and bluffing more often too. On the broader Intelligence Index, Argon at high reasoning scores 53, tied with Astra at max, one point ahead of Sol at max (52). That is 23 points above Gemini 3.1 Pro Preview (30) and 12 points above Gemini 3.8 Flash at high (41). Google is back in the top cluster. “Back in the top cluster” and “knows more facts than Astra” are different claims. Only the first one is supported. A Bloomberg report, as carried by Techmeme on 30 September, said some Google employees think Argon looks strong on benchmarks but struggles with some real coding tasks, and that Google disputes that. Unnamed employees are not a measurement. They are a reason to wait for your own tasks before you trust a launch table. I cannot see those tasks. Neither can you.
What Google says it already did inside Google
These are internal stories. You cannot rerun them. They are still the clearest picture of the kind of long job the 1 million output tokens are for. A team of Argon agents read profiling telemetry (a recording of what the computers were actually doing, not a guess) across Google’s data centers and applied memory fixes. Google says that freed over 300 tebibytes once rolled out, with an estimate of 500 tebibytes to 1 pebibyte if the rest lands. A tebibyte (1,024 gibibytes, a bit more than a trillion bytes) is a unit of storage size. This claim is about RAM (the fast memory a computer uses while a program is running), not about a warehouse of hard drives. 300 tebibytes of RAM is a large amount of working memory. It is still a Google-reported number about Google’s own fleet.
What the public version is supposed to refuse
Before a wide release, Google lists four layers. Each one needs a plain definition. Misuse checks. The public model is supposed to refuse help with cyber attacks and with CBRN attacks (chemical, biological, radiological, or nuclear). Legitimate dual-use science is supposed to stay allowed. They are also watching internal activations (the numbers firing inside the model while it works) for signs of misuse, after tests by red teams (people paid to break the safeguard). Prompt-injection defense. Covered above. Hidden text in a document should not be able to hijack the run. Google calls Argon their strongest model on this, and says it still needs several layers because the attack keeps changing. Misalignment monitors. Misalignment (the model pursues the task in a way you did not ask for, including by stepping past your intent) is watched by a second system. That watcher reads the chain of thought (the scratchpad of intermediate reasoning) and the actions, and can stop the run. Google says they used a similar watcher during training, and they deliberately did not feed those catches back into training. The fear is real: if a model is punished using its own scratchpad, it can learn to hide its thoughts and keep the behavior. They are asking the rest of the industry to keep reasoning visible while models are this capable. Whether the scratchpad you see is complete is not something a customer can check today. Hardened sandboxes. A sandbox (a locked box around the model, so a bad action hits the box and not your real systems) is being sealed before high-risk training and tests. A stronger model is also a stronger thing to contain. The box has to get stricter as the model gets more able. That is the whole of their agent-control note, in one loop: capability up, isolation up, or the test itself becomes the incident.
What to do on Thursday

If you are the defender: the application is Fairwind, not a Gemini subscription. Read the malware ban before you celebrate the missing guardrail. Your contract is the guardrail now. Use it on systems you are allowed to test. A proof of concept on someone else’s hospital is not “research” because a blog post used the word. If you ship software into hospitals, power, payments, or telecom: assume faster discovery, not a specific CVE. Inventory what you expose to the internet. Make sure a stranger cannot pull a record just by asking. That advice is older than Argon. Argon is a reason the old advice has a shorter clock. If you are choosing a model for everyday coding or for an agent: wait. When a paid API and consumer plans open, Google has said developers, enterprises, and consumers are the later phase. Secondary reports say paid API customers and Google AI Ultra come before a fully public app. Google’s post itself does not set that order in those words, so treat the Ultra-first detail as reported, not as a date. When you can call it, judge it on your tasks, not on 68%. The properties that would actually change a workflow are the long output (room for one long patch or one long migration plan) and the habit of abstaining. Check the facts anyway. Fifty percent accuracy on the knowledge exam means a coin flip on the facts that exam asks. Pair it with tests you already trust, the same way Google says it will not ship a kernel rewrite without review. If you publish scores: print pass@1 and pass@4, name the harness, say who computed the row, and say whether the build had cyber guardrails. Anything less is a leaderboard costume.
What I would not repeat
I would not say Argon tops every other model at cybersecurity. On the public patching exam it ties two models on the first try and loses the four-try order to Grok 4.7 and to Opus 5.5. I would not say hallucinations are solved. Fifteen percent is a low rate of made-up answers among the big models. Accuracy is still 50%, behind Astra. I would not say a worldwide hospital breach was just disclosed. No product name, no fix, no CVE. I would not treat 300 tebibytes, a 2.7 times faster decoder, or a 40% quantum example as your result. They are Google’s accounts of Google’s machines. The part worth keeping is smaller, and it is enough. A frontier model that can draft a real patch is now in the hands of more than 650 defender organizations, with the automatic “no” switched off and a contract switched on. The exam says it is tied with the best, not past them, and a quarter of the tasks still fail even with four tries. You, most likely, are not in that room yet. Act on the software you already run. Wait to worship a model you are not allowed to call.