Jevmaxxing
I am late to writing about Jev, largely because I've been too busy using Jev. Some notes:
- This is very real. I don't plan to collect and address every point made by Jev skeptics, but I've used Jev as much as anyone I know who didn't get pre-release access, and the only criticisms that are sticking with me are of the form "someone else will do this better soon."
- Jev is fast enough to make some asynchronous things synchronous. Imagine the "Send" button on your email client requiring confirmation if it thinks the email you're sending is ill-advised, or a utility that checks whether something looks like a recipe and adds it to a queue whenever you highlight some text.
- On questions of taste and judgment, Jev is stronger than I expected it to be. It's already more than good enough for, e.g., categorizing flashcards in ways that can't be done with a simple script.
- Jev is also usefully differentiated from more powerful models on taste and judgment questions. Suitable weightings of Jev and Gemini Flash 3.8 do better than Gemini Flash 3.8 on its own. Because Jev is so fast and so cheap, the ROI on adding Jev is large even if the Jev weight is pretty small.
- The Doom video is my favorite element of a launch announcement I can remember (and I'm not a video-games guy).
- Other LLMs are not as good as I would have expected at constructing Jev payloads. I find I have to nudge them to batch requests, ask several questions at once, and so on. This ought to change quickly, but for now, if you're used to LLMs (e.g.) making nice, reasonable OpenRouter requests by default, remember that it might take a bit of extra work to get them using Jev well.