X · Laya
Laya on X
The model itself shipped as a repository and a paper, so the explaining happened in public — and the arguing did too. These are the posts worth a reader's time: the port that carried the model further than anything the publisher wrote, benchmarks run by people with no stake, the sceptics, and the one test that came back with a loss.
22 posts described Likes read 2026-09-24 The projects behind them: builds
- X
The post that carried this model further than anything the vendor published: 14,000 likes for a port to MLX and a Snake game running at 60 decisions a second on a laptop. Read it for the port and the memory ceiling; the speed comparison in the text is against the hosted model's API-side numbers.
- X
A Tetris build against a cloud model, local Laya winning by making decisions eleven times faster on a 16 GB MacBook Air. It is a game, so the result is about the loop rather than the model, and the post does not pretend otherwise.
- X
A thread that starts as a joke about a year-old project nobody noticed and turns into the sharpest take in the corpus: the difference was not marketing but that Jev framed the problem as a category. Worth reading for the argument about framing rather than for anything about the weights.
- X
A patient explainer for readers who have heard the name and cannot say what it does: how a decision model differs from a chat model, why the output has no prose to parse, and what that changes about cost. Posted by someone with no stake in either project, which is why it is near the top of this wall.
-
A Chinese-language roundup of the five open models that appeared within days of Jev, with Laya at the top of the list. It is useful as a map of the week rather than as an evaluation: five names, one line each, no measurements.
- X
The claim-heavy version: ten times faster on published numbers, multilingual, no per-token bill, and a latency comparison of 38 ms local against 400 ms remote. Every figure comes from a benchmark somebody else ran, and the post does not say which.
- X
A benchmark run by someone with no stake, comparing Jev, a 4B open model and Laya on the same tasks — and then inferring that Jev is around 30B parameters from the gap. The inference is labelled as an inference, which is the part we trust.
- X
The second post from the same account, and the more specific one: 7 to 14 ms per decision, under a gigabyte of memory, and a claim that a local assistant setup finished in a single pass. The memory ceiling is the number that decides whether this runs on the laptop you already own.
-
A distribution argument rather than a benchmark: the open weights going into a Linux distribution's hardware integration so that a decision model is always present, offline and unbilled. If the pattern holds, this is the shape it takes when it stops being news.
-
Someone who went looking for what happened a year ago and came back with a list of prior art — GLiClass and others already doing zero-shot classification — and an argument that Jev's contribution is the framing. Disagreement with it is easy to find in the same thread on Hacker News.
- X
A practitioner's post that refuses the comparison everyone else is making: the problem is not which model is better but how you define the decision space and batch the calls, and on that reading the local model and the hosted one came out tied in his own build.
- X
A 'second brain for agents' explainer with the architecture animated, aimed at people building agent loops rather than at people choosing a model. The framing is promotional; the useful part is the split between deciding and explaining.
- X
The licence, the parameter count, the backbone and the per-query latency, in one screen — a rare post in this corpus that puts the Apache-2.0 line and the speed claim in the same paragraph, which is what a reader deciding whether to adopt it actually needs.
- X
An offer to host the open model for free, with the three question types spelled out and a thread promising the rest. Worth a card because self-hosting is the part most readers will not do, and somebody offering to do it for them is a real answer to that.
- X
The hype cut of the comparison, with the funding figure and the launch story attached: a $40M round, a former OpenAI researcher, and an open model going head to head with it days later. Its numbers come from the publisher's own benchmark table, which is worth holding in mind while watching.
- X
A first-person account of getting the MLX port working on Apple silicon and being impressed by what does not happen: no cloud call, no PyTorch, no tokens generated one at a time. Enthusiastic, and specific about the install, which is why it made the wall.
- X
A Chinese-language description of the three question types that treats them as an API rather than as a novelty, which is the right level for this model. Short, technical, and it links the model card instead of a video.
- X
The 33 ms number with the mechanism attached: bidirectional encoders, a calibrated probability instead of tokens, three primitives. A promotional account's summary, but it states the architecture correctly, which is more than most of the summaries here manage.
-
The dissent, kept on purpose: a reply arguing that the fast model is a research project next to the hosted one, with the comparison that one is a useful tool and the other is interesting. Two likes and no evidence, which is what it is — but a wall with no dissent on it is an advertisement.
-
A third-party Japanese model that claims to beat Laya's multilingual checkpoint on a 300-item benchmark built on an unseen schema — and adds an observation worth checking: that the checkpoint never chooses the first-listed score option, in either language.
- X
The most useful post on this wall by a distance, because it is the one that reports a loss: the same author tested both models as a shell-command safety classifier and the faster one let 44 of 111 attacks through against six. Posted as a warning, and it reads like one.
-
A Chinese-language test write-up with a conclusion that cuts against the enthusiasm: on the same machine Jev covered more of the questions asked of it, though the post is careful to say the gap is about coverage rather than about speed, where the local model wins.
How these were chosen
Thirty-four posts were read in full; twenty-two are published and twelve were rejected, each with the reason written down in the queue — a bare shortened link, a joke with no information in it, an automated feed account, a token launch using the model's name, and two replies that were written by a platform's own model rather than by a person. The dissenting posts are here too, including the one whose author found the faster model failing a safety test that the slower one passed.
Source: X's public syndication endpoint No API key, no OAuth, nothing rehosted The Laya topic