PROMPT TEMPLATE

Evaluate a model for a task

evaluateresearch

These templates are starting points for your own work. They are not ranked, scored, or sold. AIScoreboard does not crown a “best prompt.” Money never buys rank, and these pages have no Stripe checkout or paywall.

Prompt

Evaluate whether this API model is a reasonable fit for one task. Do not invent a leaderboard rank.

Model: [id or name]
Task: [task]
Context needs: [tokens / files / tools]
Latency and cost tolerances: [numbers or “unknown”]
Documented facts I have: [pricing, context window, modalities, license — or “unknown”]

Rules:
- Do not invent benchmarks, Elo, or number-one claims.
- Treat missing pricing or limits as unknown, not free or unlimited.
- Separate capability claims from measured evidence.

Return:
1) Task fit hypothesis (2–4 sentences)
2) Required facts checklist (known / unknown)
3) Risks and failure modes for this task
4) A minimal verification plan (what to test, with what inputs)
5) Verdict: proceed / do not proceed / need evidence — with reasons

Reason

Frames model choice around task, context, and documented limits instead of vibes or leaderboard screenshots.

Outcome

A go / no-go / need-more-evidence note tied to one task.

Insight

Require the model to list what it cannot verify. That keeps unscored models unscored.

All prompt templates