Choosing a model
- There is no best model; there is a best model for a task at a price.
- Published benchmark scores are close to useless for deciding what you should use.
- Twenty of your own examples beat any leaderboard.
- Check the rate limits before you commit — they break more prototypes than quality does.
The question is always “best at what”
A ranking of models is a category error. The flagship that writes the best code is overkill for classifying support tickets, and the cheap model that classifies tickets perfectly well will not refactor your codebase.
Start from the task. Use cases does exactly that: it takes a job, works out what one run costs on every model that could do it, and says plainly where the capability data is missing.
Benchmarks are not the evidence they look like
Most published benchmark (A repeatable test set you score a model against, so a change can be shown to help rather than assumed to.) figures are self-reported by the vendor. The test sets leak into training data over time. And a model tuned to score well on a public benchmark is, by definition, tuned for something other than your problem.
This site publishes no benchmark scores for that reason. What it publishes is prices, read off the vendor’s page, with the date and the link.
Build the smallest eval that would change your mind
Twenty examples you actually care about, with a way to tell a good answer from a bad one, beats every leaderboard in existence for the question “which of these two should I ship”.
It does not need to be sophisticated. A spreadsheet with inputs, two columns of outputs, and you marking which is better will settle most model choices in an afternoon, and it keeps settling them each time something new comes out.
The constraint nobody compares
Rate limits (A cap on how much you may send in a window, usually counted in requests and tokens per minute.). Vendors cap how much you can send per minute, usually as requests and tokens together, and exceeding either returns an error rather than a larger bill.
This is invisible in any price comparison, including this one, and it is the thing that most often turns a working prototype into a broken product. Check the limits on the tier you will actually be on before you commit to a model, not the limits on the marketing page.