Laya × MLX · updated 2026-09-27
Laya with MLX
Laya ships weights; MLX is what makes them run on a laptop. The port we have read keeps the three question types, the formatting and the calibration of the reference model, and adds the thing a hosted API cannot hand you: the decision happens on your machine, with no round trip and no per-call bill.
In short
- For: local typed decisions — 7–14 ms per short decision on an M3 Max, in the author's benchmark, with the weights converted once and no cloud call after that.
- Not for: text. Same as every other Laya page: it returns a typed decision with a probability, not a sentence.
- Read the caveat in the source: the latency figures are a one-question benchmark, not the frame time of the Snake loop in the same README — the author says so himself. Two ports measured on different machines cannot be compared row to row.
Who owns which job
The point of a pairing page is this table, not the two names in the title. Every row is a job that has to happen in the loop, the decision it needs, and which of the two should own it — including the row at the bottom, where the answer is "not the decision model".
| Job | What the decision is | Who should own it | Why |
|---|---|---|---|
| Running the weights | the runtime | MLX — or Core ML, or PyTorch MPS | Three Apple paths shipped within a week of the weights, and they are not the same trade: MLX was measured for latency, Core ML for energy per decision, MPS for memory (a 2.1 GiB fast mode and a 0.74 GiB slower one). |
| Making one decision | a typed decision with a probability | Laya | One pass, no token-by-token generation, and the same three question types as the hosted model — that is what makes the port a port rather than a reimplementation. |
| Keeping a real-time loop honest | a decision per move, sixty times a second in the Snake demo | the loop, with a safety layer | Every demo we read puts something between the model's proposal and the action — a rule that can break a cycle. A decision model that is fast is still a model that can be wrong at 60 Hz. |
| Paying for it | no per-call price | your machine — RAM, energy, and the download | The bill moves from a token price to a memory ceiling. The MLX README claims at most 1 GB resident; the MPS port trades latency for about a third of its footprint. |
What we hold on this pairing
The three Apple runtimes we have written up, plus the post that carried the port further than anything the vendor published — 14,035 likes for a Snake game on a laptop.
-
6,194 stars · Apache-2.0 · read 2026-09-24
mizorewww/laya-mlx
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
-
Source · 14,035 likes · read 2026-09-24
介绍比Jev快50倍,在你设备上跑的laya-mlx! 只在你的设备上占用最高1G内存 Laya是一个开源的类似于Jev的,基于文本输出概率的分类系统 我将其移植到MLX,并且做了一些性…
The post that carried this model further than anything the vendor published: 14,000 likes for a port to MLX and a Snake game running at 60 decisions a second on a laptop. Read it for the port and the memory ceiling; the speed comparison in the text is against the hosted model's API-side numbers.
Links to the source — we have not written this one up yet.
-
1,424 stars · Apache-2.0 · read 2026-09-24
mizorewww/laya-coreml
Local Laya typed decisions on Apple Core ML and Neural Engine. Validated ports, ~5 ms short decisions on M3 Max, reproducible speed and energy benchmarks.
-
19 stars · MIT · read 2026-09-24
afshinm/laya-mps
Run Jev-style typed decisions locally on your Mac with low RAM usage and fast responses
-
Source · 11 likes · read 2026-09-24
THIS IS F**KING GOLD I found Laya, which now runs natively on Apple Silicon Not via a cloud A…
A first-person account of getting the MLX port working on Apple silicon and being impressed by what does not happen: no cloud call, no PyTorch, no tokens generated one at a time. Enthusiastic, and specific about the install, which is why it made the wall.
Links to the source — we have not written this one up yet.
What we have not checked
- The 7.4 ms and 13.4 ms medians are the author's, measured on an M3 Max; our own run of the same weights was on CPU and is not comparable.
- The energy comparison in the Core ML port was measured by its author against one alternative on one machine.
- The 1 GB memory ceiling is a claim in the MLX README; we have not measured resident memory.
repository READMEs read 2026-09-24 like counts read 2026-09-24our own CPU run of the released weights, 2026-09-23Laya × MLX, updated 2026-09-27