FRONTIER WATCHLIST / refresh target: 3× daily Snapshot · Sep 26, 2026
MODEL INTELLIGENCE, WITH A SAFETY RATCHET

Find the model.
Improve the system.

A practical compass for frontier LLMs, classifiers and agent harnesses. Type one sentence — we enrich it with the free inference APIs, then match it to a model and a repeatable way to get the work done.

OCG+
24 model candidates · open-weight + closed18 parameterized harnesses
✳YOUR
GOAL
◈OPEN WEIGHTS12 candidates
◉CLOSED MODELS12 candidates
⌘HARNESS FIT18 playbooks
Route by capability, cost & constraints
24FRONTIER CANDIDATES
18PROMPT + HARNESS PAIRS
18STARTER GOALS
✳Every goal is enriched by the free inference APIs, then matched under your cost, speed & privacy constraints.
01 TASK → MODEL

What are you trying to do?

Describe the outcome in your own words. We enrich it with the free inference APIs, then return a model, a version and a ready-to-run harness—not just a leaderboard.

Seed watchlist · model refresh needs Worker setup
0 / 1800Signals appear here as you type
FREE INFERENCEGeminiGroqWorkers AILocal heuristic

✳ Provider keys stay server-side. Your task text is sent to the active inference provider; don’t include secrets or sensitive data.

ⓘRankings vary by benchmark, version, prompt and serving stack. This is a capability-based shortlist—not a universal leaderboard. Enrichment is inference, not fact: every added detail is an editable assumption. Confirm model access, license, cost and data policies before production.
02 STARTER GOALS

Steal a goal. Watch it get enriched.

Every example loads the goal and the constraint preset that fits it, so you can see how routing changes with the same sentence.

03 THE WATCHLIST

Frontier, in two flavors.

A curated candidate set across open-weight and hosted models. Use it to shortlist; verify checkpoint, license and live access before committing.

Snapshot date: Sep 26, 2026 · Not a real-time benchmark rankCompare live benchmarks ↗
DECISION MODELS CLASSIFIERS ARE NOT JUST SMALL LLMS

For fixed labels, consider a purpose-built decision model.

Jev and classifier/reranker families can return bounded decisions quickly. Open-weight alternatives need task-specific validation and score calibration. Rerankers are not general classifiers.

04 THE IMPROVEMENT TOOLKIT

Make the harness earn its keep.

The optimization ladder starts where improvement is measurable and cheap: prompts, context and workflow.

Open the playbook library ↗
THE
OPTIMIZATION
LADDER
L0PromptsOptimize text against a fixed metric
L1ContextCompound playbooks from traces
L2WorkflowSearch agent graphs and tool paths
L3HarnessImprove code in a sandbox
L4–5Meta / weightsResearch tier · needs real compute
05 THE PLAYBOOK LIBRARY

18 prompts. 18 ways to run them.

Parameterized prompt + harness pairs for common goals. Fill the fields, preview the prompt and copy it into your stack.

✳Every loop preserves its
fixed external scorer.
THE RSI LOOP · BUILT-IN GUARDRAILS

Improve the agent.
Never grade your own homework.

Every proposed change runs against the same frozen task set and scorer. Keep only measured improvements. Revert regressions. Preserve the full lineage.

Start with a harness ↗
01 PROPOSE
02 MEASURE
03 KEEP / REVERT
FIXEDSCORERread-only
↗↘↖
✓ equal budgets · sandbox · append-only lineage
01Freeze the signal. Hold eval set, scorer and budget constant.
02Sandbox changes. No network or host access for generated code.
03Keep or revert. Two flat cycles? Stop and diagnose.
BUILT TO BE VERIFIED

Signals, not certainty.

Use independent benchmarks and official model cards to validate any shortlist. Refresh cadence and provider free tiers can change; always check current terms.