What one run costs
A typical run of this task sends about 80K tokens and gets back 3K, over 25 calls — 2.1M tokens in total. Those are the numbers below, against the published prices.
| Model | Vendor | Per run | Can do what this needs? | Price evidence |
|---|---|---|---|---|
| GPT-5 Nano | OpenAI | $0.1300 | not stated | Vendor page |
| Gemini 2.5 Flash Lite | $0.2300 | not stated | Vendor page | |
| Claude Haiku 5.5 | Anthropic | $0.2375 | confirmed | Vendor page |
| GPT-6 Luna | OpenAI | $0.2375 | not stated | Vendor page |
| GPT-5.6 Luna | OpenAI | $0.4900 | not stated | Vendor page |
| GPT-5.4 Nano | OpenAI | $0.4938 | not stated | Vendor page |
| Gemini 3.1 Flash Lite | $0.6125 | not stated | Vendor page | |
| GPT-5 Mini | OpenAI | $0.6500 | not stated | Vendor page |
This is a reading task that ends in a small edit, which makes it the most context-hungry thing on this list and the one where a large context window (The total tokens a model can consider at once: your prompt, the conversation so far, and the answer it is writing.) earns its price.
Budget for the reading. A model that has not seen the file it is editing will cheerfully invent a function that does not exist, and the failure looks exactly like success until something runs.
Caching matters more here than anywhere else: the codebase goes in once and is read on every turn after that. If your tool supports a one-hour cache, this is the shape that justifies the write cost.
What we would pick
This section is our judgement, not a figure read off a page. Everything above is arithmetic on published prices; this is an opinion, and it is labelled as one.