Vibedia.
All vendors

Show only the vendors you work with. This applies to every page, and is remembered.

Prompts

Build a test set before tuning anything

For a chat window.

The prompt

I want to compare models or prompts on this task: {{TASK}}

Help me build a small eval:
- 15 to 20 inputs that cover the ordinary case, the edge cases and the ones where I expect
  failure. Give me the actual inputs, not descriptions of them.
- For each, what a good output looks like — specifically enough that someone else could mark
  it the same way.
- A simple scoring scheme I can apply by hand in under fifteen minutes.
- The two or three cases where I should trust my own judgement over any automatic score.

What to fill in

PlaceholderWhat goes there
{{TASK}}classifying inbound support email by urgency

Why it works

Without an eval (A repeatable test set you score a model against, so a change can be shown to help rather than assumed to.), tuning a prompt is guesswork: you change something, the next three answers look better, and you have learned nothing.

Twenty examples you actually care about, marked by hand, settles most model and prompt choices in an afternoon — and keeps settling them each time something new ships.

ShareOpen LinkedIn