HunterAlphaHub
OpenRouter model reference Facts from the public catalogue, dated and labelled
Jev Laya Verified
2026-09-22

Laya × MLX · updated 2026-09-27

Laya with MLX

Laya ships weights; MLX is what makes them run on a laptop. The port we have read keeps the three question types, the formatting and the calibration of the reference model, and adds the thing a hosted API cannot hand you: the decision happens on your machine, with no round trip and no per-call bill.

In short

  • For: local typed decisions — 7–14 ms per short decision on an M3 Max, in the author's benchmark, with the weights converted once and no cloud call after that.
  • Not for: text. Same as every other Laya page: it returns a typed decision with a probability, not a sentence.
  • Read the caveat in the source: the latency figures are a one-question benchmark, not the frame time of the Snake loop in the same README — the author says so himself. Two ports measured on different machines cannot be compared row to row.

Who owns which job

The point of a pairing page is this table, not the two names in the title. Every row is a job that has to happen in the loop, the decision it needs, and which of the two should own it — including the row at the bottom, where the answer is "not the decision model".

Job What the decision is Who should own it Why
Running the weights the runtime MLX — or Core ML, or PyTorch MPS Three Apple paths shipped within a week of the weights, and they are not the same trade: MLX was measured for latency, Core ML for energy per decision, MPS for memory (a 2.1 GiB fast mode and a 0.74 GiB slower one).
Making one decision a typed decision with a probability Laya One pass, no token-by-token generation, and the same three question types as the hosted model — that is what makes the port a port rather than a reimplementation.
Keeping a real-time loop honest a decision per move, sixty times a second in the Snake demo the loop, with a safety layer Every demo we read puts something between the model's proposal and the action — a rule that can break a cycle. A decision model that is fast is still a model that can be wrong at 60 Hz.
Paying for it no per-call price your machine — RAM, energy, and the download The bill moves from a token price to a memory ceiling. The MLX README claims at most 1 GB resident; the MPS port trades latency for about a third of its footprint.

What we hold on this pairing

The three Apple runtimes we have written up, plus the post that carried the port further than anything the vendor published — 14,035 likes for a Snake game on a laptop.

What we have not checked

  • The 7.4 ms and 13.4 ms medians are the author's, measured on an M3 Max; our own run of the same weights was on CPU and is not comparable.
  • The energy comparison in the Core ML port was measured by its author against one alternative on one machine.
  • The 1 GB memory ceiling is a claim in the MLX README; we have not measured resident memory.
Source repository READMEs read 2026-09-24 like counts read 2026-09-24our own CPU run of the released weights, 2026-09-23Laya × MLX, updated 2026-09-27

Send us a link

A project built with Jev, a post, a video, a correction, a tip. We open the link, check it says what you said it says, and write the entry ourselves.

Required a link, and an email to reply to. Optional everything else.

Add context — all optional

One or two sentences about what it does, in your words. We write the entry ourselves.

Cost, latency, a benchmark — anything you measured. We attribute these to you.

We store what you type and email it to ourselves. No IP address, no user agent, no referrer — the same rule as the mailing list.

What happens to what you send →