Try the very fast models
Models are getting better. That means that the most powerful models today are much more powerful than the most powerful models of a month ago, but it also means that today's fast models today are a lot more powerful than the fast models of a month ago.
Even if you pay close attention to new models,1 it's easy to focus on power at the frontier and overlook the "93% as good and super fast" model. But that's a mistake for a lot of reasons, some more obvious than others:
- If the fast model does something as well as the full-power model, you can save significant tokens or money by using the fast one.
- Your instincts about what tasks benefit from the full-power model are probably not perfect. Mine certainly aren't, and it seems to me that they're worse than they were a month or two ago. Put another way: "93% as good" is better than it used to be, and I suspect that the shape of that 93% is harder to understand and predict.
- Using different models to check each other's work is an important technique, and combining a full-power model with a fast one can be a great way to do this.
- Saying "fast is different!" is one thing, and (if you're like me in this respect) actually experiencing is quite a different thing. When you fix a bug in 45 seconds instead of 3 minutes, do a data-centric task over thousands of rows in a few seconds, or get professional-quality dissertation feedback2 in 10 to 30 seconds, you see all sorts of new possibilities.
- Many of us formed instincts about what to send to fast models when those models were much, much worse. Lots of day-to-day work is now amenable to fast-model assistance. For me, at least, starting all LLM work with fast models has been a useful way to reset my obsolete3 instincts.
I drafted the first 90% of this blog post last Friday morning and only finished it now. Everything above is significantly more true now than it was when I started writing it. So, this post is itself an example of LLM progress punishing (relatively, at least) normal-speed production.
This is a tricky subject to write about, just on an audience-relationship level. To a first approximation, there are two groups of readers: one group is paying little or no attention to the models powering the AI products they use, and the other is paying very close attention to them.↩
Again, by "professional-quality" I don't mean "as good as expert philosophers within their expertise," I mean "comparably useful to professionally trained philosophers outside their expertise but being careful and doing their homework."↩
That is, approximately two months old.↩