A week of testing at Every: probabilities instead of words
we almost never test new foundation models but we've been testing this for ~a week @every and it's pretty wild. the kind of things that will be obviously indispensible in 6-12 months it doesn't produce words as output, it produces probabilities. so it can efficiently act as a…
Why it is here
An editor at a media company on a week of testing the model, making the point the whole column keeps returning to — that the output is a probability rather than a word, which is what lets a decision sit inside a loop. One operator's judgement, no benchmark.
What we checked
- post read through the public syndication endpoint on 2026-09-24
- text, author, date and like count read from that response on 2026-09-24
- poster image stays on X's CDN; nothing downloaded or rehosted
What we did not check. We did not re-run the thing the post describes, so every figure on this page is the author's own and every claim is theirs — the note above is what we make of it after reading, not a measurement. The like count is a snapshot read on 2026-09-23 and it has moved since; 1,818 is what it said when we looked.