whichllm

Stop guessing models.

A public map from coding tasks to the models that actually win at them — curated, cited against Aider / SWE-bench / Artificial Analysis. Your agent asks. You keep your billing.

npx whichllm "debug flaky auth test" Star on GitHub

Task × model scoreboard

Not a proxy. The map.

Claude Code Router already owns the local control-plane niche. whichllm is different: open scoreboard data + a skill/MCP/CLI that tells your harness which model to use for each step.

Works where you already code

  • Cursor skill + MCP
  • Claude Code skill + MCP
  • Codex instruction block
  • Plain CLI via npx