Find the model.
Improve the system.
A practical compass for frontier LLMs, classifiers and agent harnesses. Match a goal to the right model and a repeatable way to get the work done.
GOAL
What are you trying to do?
Describe the outcome in your own words. Get a model, a version and a ready-to-run harness—not just a leaderboard.
Frontier, in two flavors.
A curated candidate set across open-weight and hosted models. Use it to shortlist; verify checkpoint, license and live access before committing.
For fixed labels, consider a purpose-built decision model.
Jev and classifier/reranker families can return bounded decisions quickly. Open-weight alternatives need task-specific validation and score calibration. Rerankers are not general classifiers.
Make the harness earn its keep.
The optimization ladder starts where improvement is measurable and cheap: prompts, context and workflow.
OPTIMIZATION
LADDER
12 prompts. 12 ways to run them.
Parameterized prompt + harness pairs for common goals. Fill the fields, preview the prompt and copy it into your stack.
fixed external scorer.
Improve the agent.
Never grade your own homework.
Every proposed change runs against the same frozen task set and scorer. Keep only measured improvements. Revert regressions. Preserve the full lineage.
Start with a harness ↗Signals, not certainty.
Use independent benchmarks and official model cards to validate any shortlist. Refresh cadence and provider free tiers can change; always check current terms.