Jev and a small LLM won a WebMCP benchmark at 112× lower model cost
We just ran Jev on our WebMCP benchmark. The result: basically broke the benchmark. Jev + Mercury 2.5 (a fast, low-cost LLM) using WebMCP solved 100% of the tasks at roughly 112× lower model cost than GPT-6 Astra using computer use with code execution. Compared to Astra using…
Why it is here
A WebMCP benchmark where Jev paired with a fast small language model solved every task at roughly 112× lower model cost than a frontier model driving a browser with code execution. The benchmark belongs to the company selling the tooling, so the multiple is their number, not a neutral one.
What we checked
- post read through the public syndication endpoint on 2026-09-24
- text, author, date and like count read from that response on 2026-09-24
- poster image stays on X's CDN; nothing downloaded or rehosted
What we did not check. We did not re-run the thing the post describes, so every figure on this page is the author's own and every claim is theirs — the note above is what we make of it after reading, not a measurement. The like count is a snapshot read on 2026-09-23 and it has moved since; 2,065 is what it said when we looked.