Free tool · no signup · nothing leaves this page

Score the candidates. Keep the receipts.

A model approval that cannot show its inputs is an opinion with a table around it. Pick two to four candidates, weight the criteria your committee actually cares about, and get a scorecard where every number carries its source, every gap is marked unscored rather than zero, and the whole thing exports as a memo you can file.

01 · The scorecard
Candidates

Two to four. The list is the coding-relevant subset of the models this site already prices and cards.

Committee weights 100%
Recommendation
Pick two candidates--

Choose at least two candidates above and the weighted result appears here.

Weighted scorecard. Each bar is the candidate's position between the best and worst scored candidate on that criterion. checked aug 2026

Score provenance. Where every number in the table above came from, and what kind of number it is.

A chip marked independent is a board we can open and re-read. A chip marked list price is a published rate card, which is a fact about billing rather than a measurement of quality. A chip marked vendor would be the lab's own claim; this scorecard scores none, because a vendor benchmark and an independent one do not belong in the same column.

How every number on this page is derived
02 · How to run a model review that survives audit

Write the weights before you meet the candidates.

An auditor is rarely asking whether you picked the best model. They are asking whether you had a method, applied it consistently, and can still show the inputs a year later.

Fix the criteria and the weights first

The failure mode of every model review is weights that get quietly tuned until the preferred candidate wins. The defence is boring and it works: agree the criteria and their weights in a meeting where no candidate names appear, minute the numbers, then open the data. If a weight changes afterwards, that is fine, but it changes on the record with a reason, which is exactly what an audit wants to see. A weighting that shifted from 40 to 15 on cost between two drafts is not a problem. A weighting nobody can date is.

Label every number by what kind of number it is

Three kinds show up in a model review and they are not interchangeable. An independent score comes from a board you can open and re-read: the harness, the split and the task count are published, so somebody else can get the same figure. A vendor score is the lab's own claim, usually on its own harness at its own effort setting, and it is evidence of what the lab wants you to believe. A list price is neither: it is a fact about billing that says nothing about quality. Mixing them into one column is how a review ends up asserting something nobody measured. This page keeps them apart and prints a source URL and a read date against each one.

Treat missing data as missing, not as zero

Most candidates are missing something. The temptation is to score the gap zero so the arithmetic stays tidy, and that single decision quietly does more damage than any weighting argument, because a zero asserts the model performed badly when in fact nobody measured it. Score the gap as unscored, drop that candidate out of that criterion's normalization, and score it on the criteria it does have. Then publish the coverage: what share of the total weight the candidate's evidence actually reached. A winner on 60 percent coverage is a weaker recommendation than a runner-up on full coverage, and a review that hides that difference is not a review.

Normalize inside the comparison, not against the world

Scores from different boards are on different scales, so they need normalizing before a weighted sum means anything. Normalize between the best and worst candidate in this comparison rather than against some absolute ceiling. It keeps the arithmetic honest about what is actually being decided, which is a choice between these candidates and not a ranking of the field. It also means adding a candidate can change the shape of the table, which is correct and worth saying out loud in the memo.

Keep the artifact

The output of a review is not a decision, it is a record: candidates, weights, scores, sources, gaps, and the recommendation that followed from them. Export it, attach it to the approval, and re-run it when the prices or the boards move. Both exports on this page carry the full input set, so re-running next quarter is a diff rather than an argument.

Related: models for coding agents carries the full cards behind every independent chip used here, and enterprise AI governance covers the policy the approval hangs off once the model is in.

03 · What each criterion measures

Six criteria, each with one defined source. Where a criterion has no published figure for a candidate, the scorecard marks it unscored rather than guessing.

CriterionWhat it isSourceKind
Independent coding scorePass@1 on the DeepSWE public split, mini-swe-agent harness, 113 tasks, at the effort level the board names.DeepSWE boardindependent
CostBlended list price per million tokens at four input tokens to one output token, the ratio a coding agent actually runs. Lower wins.Vendor rate cardslist price
SpeedOutput throughput in tokens per second as printed on the model's Artificial Analysis page.Artificial Analysisindependent
ContextPublished input context window in tokens.Vendor docspublished spec
OpennessWhether the weights are published and therefore self-hostable. Binary.Model cardpublished spec
Data termsWhat the vendor publishes about API data: a zero-data-retention path for eligible customers scores highest, a no-training commitment with no retention path scores partial, an unverified vendor is unscored.Vendor policy pagespublished policy

Artificial Analysis Intelligence Index is carried in the provenance panel as context but is deliberately not a scored criterion: it is a general intelligence index across nine evaluations, not a coding-agent result, and averaging it with a harness score would produce a number neither board would stand behind. The same applies to the Vals index.

04 · Questions

What a committee argues about.

The policy side lives in enterprise AI governance.

Four things, and a committee will argue about the model without them. First, the criteria written down before anyone sees a candidate, so the weights are not reverse-engineered from a preferred answer. Second, a weight per criterion that sums to a known total, because an unweighted checklist silently treats a data-retention clause as equal to a benchmark. Third, provenance on every number: whether it came from an independent board, from the vendor, or from a published list price. Fourth, an explicit rule for missing data, because that is the case that decides most reviews.

Six cover most approval decisions: an independent coding score on a real agent harness rather than a vendor benchmark, blended cost per million tokens at a realistic input to output ratio, output throughput, published context window, whether the weights are open and therefore self-hostable, and the vendor's published data terms. Weight them for your own situation. A regulated buyer weights data terms and openness far above a startup shipping a prototype, and both are defensible as long as the weights are written down first.

Mark it unscored and exclude it from that criterion, never score it zero. A zero says the model performed badly, which is a claim about the model. Unscored says nobody published a number, which is a claim about the evidence, and it is the only one of the two you can defend. This scorecard excludes an unscored candidate from that criterion's normalization entirely, scores it on the criteria it does have, and reports how much of the total weight that covers. A winner on 55 percent coverage is a weaker recommendation than a runner-up on 100 percent, and the memo says so.

Write down the criteria and weights before you look at candidates, cite a source URL and a read date next to every number, label which numbers are independent and which came from the vendor, record what was missing rather than quietly filling it in, and keep the artifact. An auditor is rarely asking whether you picked the best model. They are asking whether you had a method, applied it consistently, and can show the inputs. The exported memo and JSON from this page are that artifact.

05 · From approval to enforcement

An approved model list
is only worth what enforces it.

The scorecard settles which models the committee approved. Continuum for organizations is what makes that decision true on a developer's machine: model policy per team and per person, a weekly spend cap behind it, and usage in one view so the next review starts from evidence rather than from another table.

223 cards · every chip sourced · no invented percent