Laya · kind of work
Serving and API
Standing the model up as something else can call, usually behind Jev's own wire format.
8 records Across the 3 sources that have any 54 of 55 Laya records carry a kind
- GitHub
0xBakeer/arbiter
Serve Laya or a head you trained yourself, on a GPU or a Mac, behind a Jev-compatible API. It arrives with a build log rather than a pitch — a Sunday-morning post about a cat and a local Jev — and the diagram spends its space on where the model sits next to an LLM instead of on benchmark bars.
- GitHub
1Panel-dev/laya-server
Docker-first, with a web interface, from a team that already ships server tooling. The part worth noticing is not the UI: it is that a project with release tags and a Chinese-language README has picked Laya up and packaged it as a service you start with one command.
- GitHub
alvarobartt/sys1
A Rust server for the System One protocol: candle underneath, dynamic batching by token count, CPU, CUDA and Metal backends. The headline number comes with its hardware attached — about 14 ms per query on an RTX Pro 6000 — which is more than most rows in this column manage.
- GitHub
bladedevoff/stuntd
A local proxy that records the typed decisions your application already sends to an LLM, learns them, and answers later ones with a Laya head — so most calls never leave the machine. The interesting half is not the 90% claim, it is that the training data is your own traffic rather than a benchmark.
- GitHub
lkarlslund/laya.cpp
ggml underneath, with CUDA, Vulkan and Core ML backends, a Jev-compatible HTTP server and request batching. It is the first port here that treats Windows as a first-class target — signed binaries rather than a build script and a paragraph of apologies, which for a C++ inference engine is the hard part.
- GitHub
receptron/laya
The Node path: ONNX Runtime, no Python at runtime, and about 1.7 GB of fp32 weights cached on first call. The README claims the output matches the Python reference to four decimal places — which is exactly the kind of claim a reader should re-run on their own inputs before trusting it.
- YouTube
Run a Jev-Style Model on Your Own GPU
Starts from a homely, real problem — a support-ticket classifier that costs money on every request — and runs the model on a local GPU instead. The audience is people deciding whether to keep paying an API bill, which is most of the audience for this model.
- X
They Open sourced a Jev Alternative 6 to 8 times faster than Jev, Laya is a completely open, …
An offer to host the open model for free, with the three question types spelled out and a thread promising the rest. Worth a card because self-hosting is the part most readers will not do, and somebody offering to do it for them is a real answer to that.
Every entry on this page is in one of the columns, with the full list and its filters: All Laya records (55)
The other kinds
Runtime ports (23) · Agents and tools (6) · Browser and computer use (2) · Games and demos (9) · Benchmarks and evaluation (19)
A record can carry more than one kind, so the pages overlap on purpose All kinds