Name the output first. Chat, code, image, long-context research — the modality decides the shortlist more than any leaderboard.
Compare on usable output. Put two finalists on the compare bench: same sample workload, fee included. Then test both against your actual prompt — retries decide the real price.
Switch without re-funding. The same balance, keys and code work on all 373 models. When a stronger model ships, change one string — the model id — and keep everything else.
Cheap starting pairs. Chat: Flash vs Kimi. Code: Reasoner vs GPT Mini. Frontier: Claude vs GPT. Long context: Gemini vs Llama.