Laya says it beats Jev. The model most people will load does not.
Twenty thousand GitHub stars in six days will talk you into a bad swap.
I pulled the public GitHub API at 03:06 UTC on 24 September 2026. NandhaKishorM/laya sat at 20,797 stars. The repo was created on 18 September. That is roughly 3,500 stars a day. browser-use/jev-ultrafast was already at 19,171. X is a wall of “10 Jev repos you need” threads. The blogs that rank for jev vs laya mostly reprint one cell from the README: 0.766 versus 0.727.
That cell is real. It is also the wrong row if you are about to rip a TypeSafe key out of a Claude Code router.
On the same typed-decisions table, the two base checkpoints score 0.362 and 0.352. The majority-class guess on that set is 0.461. Random is 0.318. The weights a casual Router() can land on, before you know which checkpoint you loaded, lose to “always pick the common answer.” Laya’s own README says the 0.766 figure comes from the checkpoint fine-tuned on that benchmark’s training split. The listicles skip the sentence.
This is the Jev vs Laya piece I wanted when the screenshots started. Not another star roundup.
Why Jev vs Laya collected a crowd before the footnote did

Jev, from TypeSafe, is a hosted decision model. You send a state and typed questions. You get a label, a probability, and a confidence. No paragraph to parse. The API is closed. Published price on the Laya README is $0.042 per million input tokens, with output free. People who actually measured latency, AbdelStark and nibzard, put p50 around 236 to 276 ms. You cannot download the weights.
That gap is why the open clones sprinted. Laya is Nandakishor M’s Apache-2.0 answer: a ModernBERT-large encoder, 421 million parameters, plus a small head trained with reinforcement learning against strictly proper scoring rules (the README calls this RLCD). One forward pass. Choice, score, or noul. A router picks among an English checkpoint, a 322M multilingual model, and a typed-decisions fine-tune. The project claims about 33 ms for one question on a T4, and 7.2 ms each when you batch ten.
The crowd around that idea is not subtle. Same API pull, same minute. Kev, Jared Palmer’s trainable Qwen3.5 family, had 6,111 stars after six days. laya-mlx had 5,995 after four and a half. fast-jev-compaction, which uses hosted Jev to drop stale tool results inside Claude Code, had 6,599. SemIf, the home-GPU scorer at TheoLeeCJ/SemIf-OpenJev, had 4,115. Von had 592. OpenJev Verdict, a 151M model claiming its own leaderboard win, had 280.
Trendshift’s mention pages this week are the same names. Rising Repo’s 22 September table already had Laya, jev-ultrafast, laya-mlx, and Kev in the gain column. r/LocalLLaMA has been benching them since the 18th. Demand is early. The explainers are mostly the README with a new headline.
If you build agents, you already know the job. The loop in Claude Code still has to write the patch. A decision model is the bouncer: which model, which file, whether this tool result is still worth its tokens. The learn-claude-code teardown is the map of that loop. This post is about not hiring the bouncer because a chart went viral.
What comes back when the model is not allowed to write
A decision model does not finish your sentence. You pass a state and questions whose answers you already listed. Choice picks a label. Score sits on a small ordered scale. Noul returns the probability a statement is true. An LLM can invent a fifth department. A choice head cannot.
Laya’s README is stricter about the traps than the blogs quoting it.
Do not use true, false, yes, or no as choice labels. The checkpoint can follow the word and ignore your description. Noul has the same bug in a lab coat. The default rendering is false: and true:. Issue #156 says that pair, on the English checkpoint, can return a confident “no” on an obviously positive sentence. Override the labels, or ask a two-way choice.
Context is short. English defaults to 512 tokens, 192 of them reserved for options. The other checkpoints default to 1,024, with 256 for the head. A 77-way question gets about three or four tokens per label. That is the Banking77 cliff: Jev on 72 labels is listed at 0.870, Laya on 77 at 0.425. Raise head_max_len, shortlist to 20, or split the question. Jev’s published ceiling is 255 options. The default Laya head is not that product.
Both base checkpoints also ship over-confident. A temperature fit moves mean ECE from 0.466 to 0.081 on the English checkpoint. The multilingual one ships with no fitted temperatures. Khmer, on the English weights, scores 0.000 accuracy at 95.2 percent confidence. The router detects script before the forward pass because the confidence score will not warn you.
Speed is the number you can trust on a warm GPU. Everything else needs a checkpoint name beside it.
The Jev vs Laya headline, and the row under it
Here is the comparison the way the README actually prints it. Jev numbers are third-party published figures. The Laya author says they had no TypeSafe API access, so the samples and prompts are not the same run. Read that twice. The table is not a paired bake-off.
On typed-decisions, 400 cases and 2,000 decisions, routed Laya is listed at 0.766 against Jev 1.13.0 at 0.727. AG News: 0.950 against 0.910. DAIR Emotion: 0.595 against 0.480, and the README adds that Jev put zero probability on the true label for 16 percent of those examples. Latency: 32.8 ms against 236 to 276 ms, about 7 to 8 times faster on the author’s T4 versus those outside measurements. Cost: zero, if you already own the GPU, against a token bill.
Then the same page shows the three checkpoints apart:
laya-typed-decisions is the 0.766 model. Soft accuracy 0.471. Brier 0.062. ECE 0.213 before temperature. Score MAE 0.242.
Plain laya is 0.362 accurate. Soft accuracy 0.332. Brier 0.316.
laya-multilingual is 0.352. Soft accuracy 0.328. Brier 0.463.
Jev’s published row on that set: accuracy 0.727, soft accuracy 0.580, Brier 0.148, ECE 0.144, score MAE 0.391. A teacher self-agreement ceiling sits at 0.735. Majority class is 0.461.
So the fine-tune beats published Jev accuracy by 3.9 points and clears the 0.735 teacher ceiling the author cites. It loses soft accuracy, 0.471 against 0.580. The top pick looks sharper. The probability mass matches the teacher less well. Raw ECE before temperature fitting also trails, 0.213 against 0.144. The 0.081 in the summary table is after a domain fit. Slides love the fitted number. Your logs will show the raw one if you skip the fit.
The sentence I would pin over every “Laya beats Jev” post is already in the README, under Honest limits. The base checkpoints are near chance on this zero-shot task. The 0.766 comes from fine-tuning on that benchmark’s own training split. Laya is a fast base you specialize. On these weights it is not a zero-shot decision engine.
I did not invent the caveat. I scrolled until the README argued with its own headline. The pages ranking for this query on 22 and 23 September, listicles and Medium posts alike, stop at the headline. They will age badly.
Three public scorecards, and they refuse to agree
I did not rent a T4 and I did not spend a TypeSafe budget re-scoring 2,000 decisions. Anyone who tells you they did, last night, in a thread with no raw file, is selling you a mood. What I did was line up three public artifacts that a reader can open.
Scorecard one is the README above. Best case for Laya, on the distribution it was fine-tuned for, with Jev measured elsewhere. Winner, if you only watch argmax: the fine-tune. Winner, if you watch soft accuracy or a 70-label intent set: Jev. Winner on wall clock, once the weights are hot: Laya.
Scorecard two is yibie/laya-jev-lab. The README says the runs happened on an Apple M4 Max, and that every number comes from a script in the repo with raw output under results/. Forty intents. Jev’s API: 31 out of 40, 78 percent, mean confidence 0.88, 588 ms. Local Laya through MLX: 23 out of 40, 57 percent, mean confidence 0.71, 7.6 ms. Clear single-intent items: Jev 100 percent, Laya 75. Ambiguous: 70 against 60. Boundary junk like “test” and “???”: 40 against 20.
That lab’s useful idea is the cascade, not the dunk. If Laya’s confidence is under a threshold, send the question to Jev. At 0.60 they report the same 78 percent as pure Jev, with 55 percent of calls still local, mean latency 327 ms, about 1.8 times faster than calling Jev for everything. Pure Laya is 77 times faster and 21 points worse. You do not get both headlines at once.
Forty rows is a pilot. It is still a paired comparison, which the big README table is not. A Doom toy from the same subreddit on 20 September rhymes with it: Jev averaged 5.63 kills at 117 ms median, Laya English averaged 1.25 kills at 16 ms. Fast is not the same word as right.
Scorecard three landed on r/LocalLLaMA the morning of 24 September, in the thread titled “this time I tested reflex vs semif vs laya vs von.” Different tasks again: commit type (6-way, n=90), file routing (12-way, n=80), “is this a feature” (n=180, AUC), and change breadth (n=90, Spearman). I am quoting the poster’s table, not a run of my own.
Commit type accuracy: Jev 74.2, reflex 63.3, SemIf 55.6, Laya 35.6, Von 25.6, a path-regex baseline 31.1.
File routing: Jev 78.2, SemIf 52.5, reflex 48.7, a keyword baseline 48.8, Laya 36.2, Von 17.5.
Is-a-feature AUC: Jev 0.887, SemIf 0.821, reflex 0.801, Von 0.714, Laya 0.614, keyword 0.543.
Change breadth Spearman: Jev 0.905, reflex 0.824, SemIf 0.817, a word-count baseline 0.222, Laya 0.161, Von negative 0.046.
The poster also says Laya’s own import printed a warning about invalid temperature values, and that accuracy at the default 0.5 threshold on the feature task was 25.0 percent, worse than a coin flip, even though the ranking signal (AUC 0.614) is real. Their best threshold on that set was 0.97. Von lost to the free baselines twice.
Put the three cards on one desk and the story gets boring, which is how you know it is useful. Jev still wins the rows where someone paid for the API. Laya’s fine-tune can win a leaderboard it trained toward, and it can land near a regex when the task is “which commit is this” on someone else’s machine. Speed keeps showing up. Accuracy does not transfer for free.
SemIf and reflex are the names the star charts buried
If your question is “what open source Jev alternative should I try this week,” Laya is the famous answer and not always the strongest one in the only fresh local bake-off I could find.
SemIf is a direct option scorer on open models, built to run at home. The repo that GitHub redirects to is TheoLeeCJ/SemIf-OpenJev, created 16 September, independent of TypeSafe. On that LocalLLaMA table it beats Laya on every task and trades blows with reflex. Reflex, in the poster’s words, wraps stock Qwen3.5-4B, runs each choice forward and reversed, and combines the two. That is a calibration trick on a chat model, not a new architecture. It still beat Laya’s dedicated checkpoint on those four tasks.
I would not crown either of them. Both still sit under Jev, and the poster says nobody in that run is clean on confident nonsense. I would crown the method. Bring your own 80 rows. Include a dumb baseline. A keyword rule at 48.8 percent on file routing makes a 36 percent model look like a demo, not a migration.
Kev is a different sentence. It is a family you can train. I have not run it here, so I will not borrow a tweet for it. If you have a few hundred labeled decisions from your own router, a trainable head is more honest than hoping a viral checkpoint saw your repo. OpenJev Verdict’s 77.10 percent claim can be true on its items and still tell you nothing about commit routing. Until two setups share items, do not average them in your head.
jev-ultrafast is a different product again. Browser Use numbered the controls and let hosted Jev pick the action and the element in one request. Their writeup: a Zurich to London search in 7.1 seconds, median time from 9.45 seconds to 7.09, browser calls from 1,092 to 101. A smaller model types only when something must be typed. Dropping an untested local head into that loop to “save the API” is how a 7 second demo becomes a misclick.
fast-jev-compaction has the same shape, aimed at Claude Code context. It scores tool calls and drops the stale ones instead of asking a big model to rewrite them shorter. If your pain is the bill, that plugin sits closer to the Opus 5.5 prompt audit than Laya does. Shrink the instructions first. The 65-line CLAUDE.md pattern existed to stop extra work. A decision head can refuse a step only after you know which mistakes it makes on your diffs.
When the local model is still the right call
Keep Jev if you need a hosted API this week, your label lists run past twenty, or the state is a long ticket that will not fit in 512 tokens. Log confidence. Send the ugly tail to the coding model, not the easy cases.
Pick Laya if the bytes cannot leave the machine, you will fine-tune or at least fit a temperature, and the questions stay small. Load the typed-decisions checkpoint on purpose. A cold load used to be about 22 seconds and is closer to 2 seconds now. The 33 ms number is the warm forward pass, not the first call after pip.
Use the M4 lab’s cascade if you can spend a little API money. Local when the score is high. Jev when it is not. That is the only pattern in this pile that matched hosted accuracy without shipping every token upstream.
Skip Von until it beats a regex on your task. Skip any “beats Jev and Laya” tweet that does not publish the items. The clerk in an agent stack only classifies, ranks, and abstains. Open weights made that clerk cheap this week. They did not make it automatically right.
Ten minutes, before you delete the key
You can do the honest half of this without a GPU.
Create a virtualenv and install the package. The README’s version check does not load a checkpoint:
Then run the offline router line, the one that does not download weights:
You should get a routing decision, not an answer. That is the point of the exercise. If a blog told you this command proves Laya beats Jev, they confused the language detector with the model.
If you do have the disk and you load a checkpoint, do not ask a toy question and nod. Take 30 real states from your own logs. Write one choice question with fewer than eight semantic labels. Include ten states where the right answer is boring and common, because that is where a pretty accuracy number hides. Score a majority-class baseline on the same 30 before you score the model. If the model does not clear that baseline by a margin you would explain to a teammate, leave the API key where it is.
While you are there, ask one noul with the default true/false wording and the same noul with labels A and B, on a sentence that is obviously positive. If those two disagree, you just reproduced the bug the README already confessed. That is a better afternoon than starring twelve repos.
Common questions about Jev vs Laya
Is Jev open source?
No. Jev is TypeSafe’s hosted System One model. The weights and the training recipe are not public. Laya, SemIf, Kev, Von, and the MLX port are separate projects with their own licenses. Laya’s code and checkpoints are Apache-2.0. Compatible JSON is not the same thing as a TypeSafe release.
Is Laya the same as Jev?
No. Same job, similar question types, different models. Laya can expose a Jev-shaped HTTP server if you install the serve extra. Calling that server “Jev” in a status page will confuse the next person who reads the bill. Log the checkpoint name on every row. The router can switch among three of them inside one process.
Is Laya better than Jev?
On the author’s fine-tune, argmax accuracy on typed-decisions is higher, 0.766 against a published 0.727, and the forward pass is faster. On soft accuracy, on Banking77-scale label sets, and on the two independent checks from this week, Jev is ahead. The base Laya checkpoints lose to a majority-class guess on the suite that produced the viral number. “Better” without a checkpoint name and a task is a tweet, not a result.
Can I run Jev locally instead?
Not from TypeSafe. If the requirement is air-gapped inference, you are choosing an open model and accepting its error bars. Laya is the one with the most eyes on it right now. SemIf and reflex deserved a look in the 24 September bake-off. None of them is “Jev, but on a laptop.”
Which open source Jev alternative should I install first?
If you will fine-tune or you only need a fast reject gate in front of the real API, start with Laya’s typed-decisions checkpoint and measure it against a dumb baseline. If you want the strongest local score on the fresh commit-routing tasks and you have a GPU at home, read the SemIf and reflex writeups before you inherit 20,000 stars of certainty. If you have no labels and no time, keep Jev and spend the afternoon on your prompts instead.
If you try only one thing from this page, make it the baseline. Stars tell you what people shared. A majority-class guess tells you whether the share was earned.