A 2.5-million-user product benchmarked Jev on résumés
Stanford CS PhD here 👋. I rigorously benchmark Jev/Typesafe on my consumer AI app serving 2.5 million monthly active users https://t.co/b52LE8B3yN. At HiringCafe we score user resume x job description relevance. Here's how it does 🧵
Why it is here
A Stanford PhD running a hiring product with 2.5 million monthly users benchmarked the model on the task his company actually pays for: scoring how well a résumé matches a job. Production numbers on a real workload are rare in this topic, and he shows the method.
What we checked
- the post read through X's public syndication endpoint, 2026-09-25
- like count and date read from the same endpoint the same day
What we did not check. We did not re-run the thing the post describes, so every figure on this page is the author's own and every claim is theirs — the note above is what we make of it after reading, not a measurement. The like count is a snapshot read on 2026-09-23 and it has moved since; 809 is what it said when we looked.