TL;DR: Strands Decider 2B, released 1 October 2026 by Strands Labs, is not a chatbot. A chatbot writes the next piece of text. This model cannot. It only scores answers you already listed, such as billing versus sales, and it attaches a confidence (a number for how sure it is). The weights (the stored numbers that make the model), the training data, and the training scripts are public under Apache-2.0 (a license that lets you use, change, and ship the code, including in a product, if you keep the notices). Install it only if your agent (a program that loops: look, decide, act) keeps asking a big model questions that have a fixed menu. Skip it if you needed words, code, or a summary. Their own sample ticket comes back at confidence 0.77. Their README says that on short classification tasks, only scores at or above 0.9 were right about 95% of the time. So the demo itself is not an automatic route.

What actually shipped?

A small open model that picks, plus the recipe to rebuild it.
On 1 October 2026 the Strands blog introduced Strands Decider 2B. The same files are on GitHub at strands-labs/strands-decider and on Hugging Face as StrandsAgents/strands-decider-2B-hobson-v19. When I opened the repo on 2 October 2026 it showed about 137 stars (a public bookmark count). That number will move. The README says a later experiment, v20, did not replace v19. The id to download is still hobson-v19. Hobson here is a checkpoint name (the saved snapshot), not a second product.
The same week, Google announced Gemini 4 Argon, a frontier model (the largest, most capable class) that most developers still cannot call. This post is the opposite object. Argon is built to write a very long answer. Decider is built to refuse that job.
What is a model, before the new word matters?
A model is a pile of stored numbers that turns text into scores.
Those numbers are parameters (one parameter is one stored number). Training is the process that chose the numbers. A language model (often called an LLM, a large language model) uses the numbers to guess the next token (a token is a small chunk of text, often a word or part of a word). It guesses, appends that chunk, and guesses again. That loop is what you experience as writing.
You pay a vendor per token when the model runs on their computers. A local model runs on your computer. You pay electricity and whatever graphics card you already own. There is no per-token invoice.

What did they cut off?
They kept the torso and threw away the part that speaks.
The starting point is Qwen/Qwen3.5-2B-Base. Qwen is an open model family from Alibaba. Base means this copy was not then tuned to chat. 2B means about 2 billion parameters. The README describes this decider as 1.9 billion parameters and still calls it 2B. Close enough to treat as the same size class. It is not a 100-billion-parameter chat model.
A transformer (the common shape of these models) has a torso and a head. The torso reads tokens and, at each position, emits a hidden state (a list of numbers that stands for “what this spot in the text means so far”). The language-model head is the last layer that turns a hidden state into a probability for every token in the vocabulary (the dictionary of chunks the model is allowed to emit). That head is how prose gets born.
Strands removed that head. They put on a pointer head of about a million parameters. The blog says the pointer head scores the hidden state at each option against the hidden state at an <answer> position. You do not need the algebra. The behavior is the product: it cannot invent a sentence. It can only point at choices that are already in the input.

What is the LoRA sticker?
They did not retrain all 1.9 billion numbers.
LoRA (low-rank adaptation: train two thin grids of numbers whose product is a small update, and add that update onto the frozen original) is how the torso was nudged. Rank 16 means each thin grid is only 16 columns wide, so the update stays small. The original torso stays frozen (not rewritten as a whole). The Hugging Face card says what you download on top is that LoRA adapter plus head.safetensors (the pointer head, in a common weight-file format). The base weights come from Hugging Face the first time you run it.
The README says the head runs in fp32 (32-bit floating point, a precise way to store a decimal). That is an implementation detail. The part that changes your install is the size of the torso, not the sticker.

What are the three questions it will answer?
Choice, yes-or-no, or a point on a scale you named. Nothing else.
The README’s sample state is: “Help! My payouts have been failing for 3 days!”
A choice question gives a menu. Their sample menu is billing, sales, retail. The sample output picks billing at 0.845, retail at 0.091, sales at 0.064, and reports confidence 0.768. Those decimals are one example printed in the README, not a law of physics. Run it yourself if you need your own number. I did not.
A noul question is their yes-or-no type. The README does not expand the word noul. Closer to 1 means more yes. Their sample, “Does this convey urgency?”, prints noul = 0.828.
A score question lays down a scale, such as calm, frustrated, depressed. Their sample prints score = 1.10 with frustrated at 0.573 and confidence 0.518. A confidence of 0.52 is a shrug. Do not auto-act on a shrug.
You can ask all three in one command. The README says that is cheaper because the state is loaded once. The HTTP example on the same page reports output_tokens: 1 and latency_ms: 140.03 for a single yes-or-no. One output token is the tell. A writing model would have emitted a paragraph of tokens. This one did not.

Where does this sit in an agent?
In front of the expensive model, and only for the rote fork.
An agent is not the model. The model scores text. The agent is the loop around it: read the repo or the ticket, decide, call a tool (a real action, such as send, edit, or run), look at what happened, repeat. A lot of those decides are menus. Which tool. Which team. Is this allowed. Is this urgent.
The blog’s intended jobs are model routing (which model should take this), tool selection, argument checking, triage, guardrails (rules that block a bad action), evaluations, and the rote half of a hybrid agent. Hybrid means the small model takes the fixed-menu questions and a language model takes the hard ones. Strands also shows an InterventionHandler that can answer Proceed, Deny, Confirm, or Guide before a tool call. I am describing their example, not a setup I deployed.
If the confidence is high and the menu is real, act. If the confidence is low, or the job is writing, hand it to the big model. Coding, summarizing, and “figure out what the options even are” are writing jobs. This model will not grow a vocabulary to save you.

What numbers did they actually publish?
About 167 of 231 public tasks, and a tenth of a second on their cards. Not a trophy I verified.
JevBench is the public test set they score against. Jev is a decision model TypeSafe AI launched earlier in September 2026. The Strands post uses Jev as the reason this category exists, and JevBench as the shared exam. I did not re-grade the exam.
The model card lists 167 of 231 public JevBench tasks right at a 4096-token window, with a Brier score of 0.348 and an ECE of 0.050. The README says that at 4096 tokens, v19 scores 168 of 231. I am not going to hide a one-task gap. Call it 167 or 168 of 231. That is about 72 out of 100, because 167 divided by 231 is about 0.72. The blog’s ranking, not a percent, is third of 33 models in the 2B class on accuracy and calibration together, and first of 30 if you set aside models just over 2B. I did not see that full table, so I am not naming who placed first.
A token window (also called a context window) is the maximum text the model will look at in one go, counted in tokens. The README says the preregistered window is 3072 tokens. Longer than that, you are outside the test they signed up for.
A Brier score is the average of (predicted probability minus the real outcome) squared, where the outcome is 1 if the answer was right and 0 if it was wrong. Zero would be perfect. ECE (expected calibration error) sorts guesses into buckets by confidence and checks whether “80% sure” is right about 80% of the time. Lower is better on both. I am not turning 0.348 into a letter grade. The card does not spell out which variant of the Brier formula they used, and a multi-answer exam is not the same scale as a coin flip.
Calibration (whether the confidence matches reality) is fitted, the card says, as one temperature (a single knob that sharpens or softens the probabilities) per question type, on held-out short classification. Held-out means those examples were not used to train. The confidence bands are established there only. The card tells you to measure your own traffic before you trust a threshold.
Latency (how long one decision takes) is a median (the middle value: half the tasks were faster, half slower) of about 115 milliseconds on an Nvidia RTX 3090, and about 153 milliseconds for small tasks on an M3 Pro, which is an Apple laptop chip. A millisecond is a thousandth of a second. The blog says the wait grows roughly in a straight line as the task gets longer. Their HTTP sample on an unnamed machine was 140 milliseconds. Do not quote 115 milliseconds as your number until you time your machine.

What will fail if you trust it too early?
The question can be ignored. The long task is the weak one. Their threshold is not your threshold. Yes-or-no travels badly.
The model card says questions are read less than documents. With the state and the options fixed, a changed question often gets the same answer. If you only rewrite the question and expect a new decision, test that. Do not assume it.
Long, multi-step documents are the weak spot. The card says JevBench’s hard tier scores far below its easy tier. Do not point this at a research job.
The README’s practical rule is narrower than the marketing sentence “calibrated confidence on every decision.” On short classification tasks, answers with confidence at or above 0.9 were correct about 95% of the time. Below that, it says to confirm or ask a person. Their own billing example is 0.768. Under their rule, that ticket does not auto-route. That is the useful part of the README, and it is easy to skip if you only read the blog.
The card also says score and noul transfer poorly onto rubrics and yes-or-no tasks that do not look like the training mix. Start with choice among labels you actually use. And the training data is public data, so the model inherits those topics and those labeling mistakes.
Their internal sets, which are not the public exam, swing hard. The card lists a held-out short set at 0.641 across 6,000 examples, HotpotQA at 0.717 across 959, and ContractNLI at 0.872 across 1,026. The spread is the lesson. “Third in its size class” is not “right on your tickets.”

How do you try it, and when do you skip it?
One install command. Then a menu you already believe in.
From the README:
pip install strands-decider strands-decider ask StrandsAgents/strands-decider-2B-hobson-v19 \ --state "Help! My payouts have been failing for 3 days! " \ --choice "Which team should handle this?=billing,sales,retail"
pip is the Python package installer. The first run will also pull the Qwen base weights. To keep it awake as a local service:
strands-decider serve StrandsAgents/strands-decider-2B-hobson-v19 --port 8000
Port 8000 is just a door number on your own machine. Nothing there is public unless you make it public. Do not expose it to the internet with a customer ticket in the request until you mean to.
Use it when three things are true. The legal answers are a menu you typed. A wrong label is cheap because low confidence gets a second look. You can hold a 2B model in memory. The license is Apache-2.0, and they shipped the training data and the scripts, so you can see what it studied instead of trusting a slide.
Skip it when any of these are true. You needed sentences. The menu does not exist yet. The note is longer than the 3072-token window. One wrong route is expensive, such as sending a payment or deleting a row. A guardrail that sometimes ignores the question is not a guardrail.
