Anthropic’s Opus 5 is the first AI model I have seen where the “how hard should it try” dial is actually exposed to the user, and that is an interesting design choice. I personally haven’t tested out Opus 5 yet as I’m super conscious of my token usage as I have a lot of projects underway, but it is on my radar for something to test.
My way of doing things, honestly, is once I find a model (regardless of provider) that seems to do relatively quality work and that doesn’t just slurp away all of my usage, I tend to stick with that model. While I know that all of the AI companies are playing a game of leapfrog and trying to outdo the competition, I do also feel that you need to cautiously wade into those waters.
Test the new models on a task or project, and use your human side to gauge the quality of the output. Then weigh it against the amount of usage that task or project consumed. If it works for you, go for it. If not, roll back to an earlier model that you are more comfortable with.