What to fill in
| Placeholder | What goes there |
|---|---|
| {{TASK}} | classifying inbound support email by urgency |
Why it works
Without an eval (A repeatable test set you score a model against, so a change can be shown to help rather than assumed to.), tuning a prompt is guesswork: you change something, the next three answers look better, and you have learned nothing.
Twenty examples you actually care about, marked by hand, settles most model and prompt choices in an afternoon — and keeps settling them each time something new ships.