Published 20 July 2026 · Updated 19 August · Runs 17 July–19 August

01 / 09

Can it RunAlfrada?

Sixteen models.
Two real jobs.
Every file audited.

Fifteen of sixteen models delivered useful work. One looked finished but got the decision wrong. Here’s what to use.

16models tested
238judge verdicts
255.1Mtokens processed
$340all-in spend

See the results ↓ Markdown copy Or download the full study for sharing, Notion, or LLMs

01

The answer, at a glance

Pick by quality, cost and risk — not rank alone.

Best quality

Sol · 8.84

Best-checked work. Use it when a wrong answer would be expensive.

Best cost & speed

Luna · $0.92

Scored 8.04 in about ten minutes. Best for fast, lower-stakes work.

Best all-round value

K3 · 8.40

Strong numbers, no decision-changing mistakes, and only $3.54.

The warning

Inkling · 5.60

Complete-looking work, but wrong math and an unreproducible forecast.

Quality versus cost — the useful shortcut

Higher is better. Further left is cheaper. Bigger bubbles took longer. Gold marks the best value frontier.

Read the finding behind this chart

GPT-5.6 Luna is the study's outlier: 91% of the winner's composite at 6% of its cost, each case done in four to six minutes of agent time at $0.057 per quality point. Its artifacts explain the profile — honest framing, correct core arithmetic, disciplined refusal to fabricate — with savings taken from evidence staging and a missing second QA pass. For production workloads where an artifact lint is cheap, that trade wins.

Compare all sixteen models

Sort by quality, cost, time, or value. Click a model to inspect its work.

How scores and costs work

Score / 10 is the quality of the finished work, graded by other AI models that opened every file and re-checked the math. Balanced rank adjusts for judges who grade unusually easy or hard. Costs shown are contestant spend, not judging spend; time is how long the model itself worked.

Why this benchmark matters
Why it matters

Managers buy decisions, not benchmark points. A polished answer can still carry a recommendation-breaking error.

Why this is different

The unit of evaluation is the complete, inspectable work product: sources, calculations, memo, deck, charts and schedules.

What to do

Choose models by quality, decision risk and cost together; require an evidence package; review the load-bearing chain before the prose.

Read the full executive summary

This is a management-risk test, not another isolated model quiz. Sixteen models each completed two long-horizon knowledge-work jobs inside the same production harness — nine in the 17–19 July cycle, then update runs through 19 August (Gemini 3.6 Flash 22 Jul, Claude Opus 5 25 Jul, DeepSeek V4 Flash 1 Aug, Grok 4.6 13 Aug, Gemini 3.7 Flash and Qwen 3.8 2.4T 14 Aug, GLM 5.3 19 Aug). Independent model auditors then opened every deliverable, replayed calculations and challenged the evidence chain. Raw artifacts and every verdict are published below.

Task suites, coding benchmarks and platforms like AgentMark are valuable: they make capabilities comparable and regressions repeatable. This study asks a narrower operational question: can a manager trust the package an autonomous job delivers, and where will expert review still be load-bearing?

The answer is not “all frontier models are good.” Fifteen cleared a practical 6.0 usability floor; Inkling scored 5.60 and failed on core arithmetic, internal consistency and reproducibility. Opus 5 leads the calibrated ranking and Sol the raw composite — a gap well inside run-to-run noise, but the two separate on inspection: Sol carries no catalogued load-bearing defect, Opus 5 the most on the board. Grok 4.6 enters third on calibrated rank with the board's highest single case-1 score (8.87) at a quarter of the leaders' spend. GLM 5.3 debuts fifth — a hair above Fable on raw composite, behind it on calibrated rank — at $2.54, the cheapest top-five package. Luna holds the strongest cost-speed trade; K3 the best quality-integrity balance. Of the two 14 August entrants, Gemini 3.7 Flash posts the board's lowest contestant spend ($0.75 across both cases) but lands mid-pack on audited quality; Qwen 3.8 2.4T debuts with the earlier cycle's strongest reproducibility — three judges independently re-ran its 50,000-trial forecast engine and matched every published statistic.

This edition is a cross-section, not our first benchmark. Over roughly three months we ran the evolving harness hundreds of times during development. The published table is a seventeen-candidate cycle that yielded sixteen attributable entries after one invalid Qwen swarm result was excluded. The pattern matches prior runs — enough for directional confidence, not per-model defect-rate claims.

02What we test

In short: a real AI model, dropped into Alfrada — our actual product — doing a real job. Same setup for every model, so the ranking reflects how they handle the work, not a quiz score. These are long-form jobs: hundreds of steps of research, modelling, writing and delivery — the kind of work that takes at least 45 minutes to complete, not a short quiz.

The model does the thinking; Alfrada gives it the tools: planning, nineteen live tools, a private workspace, the ability to run code, and to create and hand over files. Because every model used the exact same setup, the ranking shows how they handle the same real job — not how “smart” they look on a quiz.

It's not a coding test either. The job mixes research, judgment, number-crunching, writing and delivery all at once. A good setup should even the odds between models: guardrails and checks keep the work complete, on-spec and easy to inspect, so what's left to judge is the quality of the model's own thinking.

03The test

Every model ran in the exact same Alfrada setup: same tools, same instructions, a fresh, empty workspace, and no memory of any other run. No model could peek at another's answer or quietly hand the hard parts to a smarter model. Then we gave each one two long-form jobs — hundreds of steps, at least 45 minutes of work — of the kind a real strategy team might actually order.

Case 1 — Make a business decision

“Should NorthLine, a $12M-a-year logistics software company, expand into the UAE in mid-2026?”

  • Market and competitor research where every claim is backed by a real source — screenshots must actually show the number they're used for
  • A 24-month financial model built from scratch: market size, cost to win a customer, payback time, currency-swing risk, and a clear break-even point
  • A 1,500-word decision memo, a 10-slide board deck, a 40–50-second audio briefing, and a scheduled 30/60/90-day review plan

Case 2 — Make a forecast

“Forecast how much the world will spend building computer chips through 2026–2030 — as a range of outcomes, not a single number.”

  • Five things that drive the forecast, each backed by a named, dated, findable source
  • A simulation of at least 10,000 possible futures, saved so anyone can re-run it — quoting a number the simulation never produced counts as making it up
  • Charts showing the range and what moves it, plus a 1,500-word memo with two “here's what would prove us wrong” scenarios and one testable 18-month prediction
Check one

Did it deliver?

Automatic checks confirm the basics: every required file is there, the word and slide counts are met, the tools ran without errors, and the schedules work.

Check two

Does it hold up?

A panel of judges opens everything, re-runs the math, checks the sources, and scores the finished work from 0 to 10.

The judges are seven other AI models: Kimi K3, GLM-5.2, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Luna, Claude Fable 5 and Claude Opus 4.8. They don't just skim the final answer — they open the files, re-run the math and the simulations, recheck the costs, and even look at the slides and charts. Every model was judged by all seven.

The fine print on the judges

All seven judged each of the nine original contestants. The 22 July Gemini 3.6 Flash and 25 July Opus 5 runs added Gemini 3.6 as a judge — pooled with Gemini 3.5 into one joint Gemini signal — and Opus 5 also judged itself, a verdict published below but carrying no scoring weight. The DeepSeek V4 Flash (1 Aug), Grok 4.6 (13 Aug), Gemini 3.7 Flash and Qwen 3.8 2.4T (14 Aug) and GLM 5.3 (19 Aug) runs used the same expanded panel, with Opus 5's verdicts likewise published but unmapped. Verdict and session totals are recomputed in the published data payload. Qwen, Inkling and DeepSeek Flash are contestants, not judges: Inkling's own work fell below the usability cutoff and lacked the verification discipline an evaluator needs.

04Results

In plain terms: think of 6.0 as the pass mark, and the score as a grade out of 10 — not a percentage. Fifteen models scored 7.05–8.84 and gave a usable first draft, some needing more review than others. Inkling scored 5.60: it looked finished, but the math behind its recommendation was wrong, its key files disagreed with each other, and its forecast couldn't be reproduced. Below six, treat the work as a rough starting point, not a draft to polish.

Two models finished on top: Claude Opus 5 leads the fairness-adjusted ranking, and GPT-5.6 Sol has the highest raw score (8.84). Neither was grading itself — a model's vote on its own work never counts. The gap between them is smaller than the normal wobble you'd see from one run to the next, so treat them as tied; what separates them is the judges' notes: Sol made no mistake big enough to change its recommendation, while Opus 5 made the most on the board. Grok 4.6 comes third on its first try — the single best score on the business-decision job (8.87) at a quarter of the top models' cost, though a weaker forecast pulls its overall to 8.63. Fable is fourth. GLM 5.3 lands fifth (8.55) and is the cheapest of the top five, at $2.54. Then K3 and Qwen 3.8 2.4T — Qwen's forecast could be re-run exactly by three separate judges, but slips on the first job. Luna is the bargain of the group; DeepSeek V4 Flash reaches 7.85 for a tiny fraction of the usual cost. Every model cleared the basic delivery checks — which didn't stop some of the finished work from being unusable.

Who judges whom

Rows are the judges; columns are the models being judged. Click any cell to read what that judge said, in full.

How to read this chart

Cell shading spans scores 3–10 across both cases. The right-hand column is each judge's mean. Sol (5.85) and Luna (6.24) are materially harsher than the joint Gemini row (9.93). Corner-ticked cells contain self-judgments; they are visible but excluded from calibrated rank.

The matrix is where this study earns its keep. Raw means reward generous judges and duplicate architecture families. Calibrated rank instead z-scores each judge within each case, averages Sol/Luna into one GPT-5.6 signal and Fable/Opus into one Claude signal, and gives K3 and GLM one signal each. Self-votes are excluded. The joint Gemini row stays fully visible but gets zero ranking weight — its near-uniform marks add almost no discrimination.

How models grade their own work

Each judge's score for its own work, next to what everyone else gave it (both jobs combined). Inkling and Qwen aren't shown — they weren't judges.

Read the finding behind this chart

Both Geminis rated their own work far above their peers' view. Luna is the opposite: it graded its own case-2 submission 4.5/10 while six peers averaged 7.6. Sol also self-graded below its peer mean. Self-votes are diagnostic only and do not affect rank.

05How to read the scores

Treat 6.0 as the pass mark, and the score as a grade out of 10. From 6–7, it's a draft that needs supervision; 7–8, useful with a targeted expert review; above 8, more of it survives the audit — but no score removes the need for a human to own a big decision. And “no mistakes found” means none turned up in this run, not that the model never makes them.

Why these scores work this way

Some things come down to taste — how good the writing, the recommendation or the slides are. But the math doesn't. Either a model's simulation could be re-run and gave the same answer, or it couldn't. Either its cost figures added up, or they didn't. Either a screenshot actually showed the number it was quoted for, or it didn't. Those yes/no checks give us solid ground under an otherwise judgment-based score.

That's why these scores don't line up with a coding benchmark or an exam percentage. A coding test asks whether one small fix passes; here, writing code is just one tool in a much bigger job. The model has to research, build a model, decide, explain, and hand over several files that all agree with each other — then survive judges opening every one.

Alfrada also sets a floor. Its built-in checks and repeatable tools keep every model's output at a steady minimum, so the real differences show up above that floor. (We didn't test the models with Alfrada stripped away, so we can't put a number on how much the setup helps.) What the study does show is the gap between merely delivering and actually holding up: every model cleared the basic checks, yet judged quality ranged from 5.60 to 8.84.

Making things up is only part of it. Invented facts are penalised — but the harder test is spotting a break in the reasoning even when the answer looks polished: evidence that doesn't actually back the claim, math that contradicts the memo, or a check that just proves itself. That's a tougher bar than “did it run?”

06Mistakes that change the decision

The worst mistakes are the ones the recommendation actually rests on — fix one and the money, or the advice itself, changes. On that measure: Sol, Luna and K3 had none; Fable, Opus and GLM had one each; Gemini 3.5 had two; Inkling had three.

Why 5.60 is a failed job, not just last place. Inkling added up a full year's revenue every single month, overstating two-year revenue by about 12× and claiming the business broke even in its first three months. Its data file, memo and charts don't agree on the costs or the timing. On the forecast job, eight reasonable attempts to follow its stated method all failed to reproduce its published numbers. It handed in a full package — but the thinking behind it fell apart the moment anyone checked.

The mistakes in detail, model by model

By that measure, Sol, Luna and K3 had none — and left evidence the judges could actually check. Qwen 3.7 Max also got its core math right, but its evidence didn't back up its market claims, so it isn't as safe a bet. Fable, Opus and GLM each had one: Fable called month 26 “break-even” when it was really the first good month before two more bad ones; Opus asked the board to approve a partner deal whose economics its own model never included; GLM told the board break-even came in month 30–33 when its own data file said it never came at all. Gemini 3.5 had two, one in each job. Gemini 3.6 — the 22 July update — fixed its predecessor's break-even math but added a new slip: a chart built on numbers its own model doesn't produce. Inkling had three, and failed the pass mark.

A second warning: a polished answer can outlast the facts under it. Gemini 3.5's memo says break-even means profit beating costs — then quietly works it out on revenue instead, skipping the 20% it costs to deliver. By its own definition the UAE branch never breaks even within two years; the memo says month 21 and recommends going ahead partly on that number. Four judges caught it independently. On the forecast job, the same model quietly tuned its test to fit the past, then presented that near-perfect fit as if it had predicted the future — at “92.0% confidence.” Both times, the reasoning broke before the writing did.

The models agree more than the scores suggest — fourteen of the sixteen said a cautious “yes” on the business decision, and every model whose numbers checked out agreed the branch doesn't pay for itself within two years. Where they really split is the forecast:

Same question, same sources, 5.1× apart

Each model's best-guess five-year growth rate for global chip-building spend, with its own likely range where given.

Read the finding behind this chart

Qwen 3.8's 5.1% and Gemini 3.7 Flash's 25.8% answer the same prompt from the same sources. The spread comes from anchor choice and prior construction, and is invisible in the quality scores: rigorous, reproducible runs can still disagree sharply. For planning, that disagreement is the finding.

07What this means in practice

The takeaway: a good setup guarantees a solid floor, not a correct answer. Cheaper models can do real work inside strong guardrails — as long as a person still checks the reasoning. Here's how to put that to work.

What this means for how you work

Alfrada made every model deliver the full package. Fifteen gave a useful draft; one proved that ticking every box can't save broken reasoning. Cheaper models can do real work inside strong guardrails — as long as those guardrails keep enough evidence for a person to check the reasoning.

The expert's job moves up a level. Analysts, scientists and reporters spend less time building the first draft and more time deciding which assumptions matter, re-checking the calculations that count, and sorting out conflicting evidence. Managers stop asking “is the work done?” and start asking “is the recommendation actually right?”

For higher-stakes work, Alfrada has a “Beast” mode that adds a built-in critic-and-judge loop to push borderline work up into the high 7s and 8s. We're not claiming that boost here — Beast mode was switched off for these runs.

Commissioning

Match the model to the stakes

Sol when a missed mistake is costly, K3 for balanced quality, Luna for fast, high-volume work, and GLM when you need repeatable modelling and have someone to review it.

Review

Check the bones first

Start with where the sources came from, the definitions, the math, re-running the simulation, and whether the files agree. Proofreading the polished text first is checking the wrong thing.

Keeping records

Keep the working, not just the answer

For analysts, keep the underlying models. For scientists, keep the code and settings. For reporters, keep the original sources and the link from each claim to its evidence.

Accountability

AI drafts; a person owns the decision

This helps you choose a model and design the workflow. It doesn't hand your professional responsibility over to a score.

08The evidence

A benchmark you can't check is just an opinion. For every model, we publish the full record: each judge's score and reasons in full, the list of mistakes found, the charts it drew, and a downloadable pack of its actual files. It's all here — open it if you want to dig in.

Prefer a portable copy? Download the Markdown export (includes absolute URLs for every chart).

Show the full evidence — all 16 models

09Method & limitations

How the runs, scoring and cost accounting worked — and the honest limits of a pinned, single-machine study. For readers who want to audit the numbers.

Read the full method & limitations

Setup. Runs executed July 17–August 19, 2026 against a production Alfrada build on a single dev machine: one isolated bot user per run (separate workspaces, budgets and history), temperature 0, seed 42 in the case config, $15 budget cap per run. Every model ran as the orchestrator and delegated sub-agent legwork to Alfrada's fast worker pool — mostly GPT-5.6 Luna and Gemini Flash — identically across all sixteen runs. That's the configuration production users get, and what the scores measure: the model directing a job, not doing every keystroke of it. Every worker call is recorded and published in the packages. Judge sessions ran under a separate critic user with read-only tools plus code execution for replay.

Scoring. The 0–10 raw composite is the mean of the seven judge signals across both cases, kept visible for interpretability. Rank is calibrated separately: within each case, every judge is z-scored across the sixteen submissions; Sol/Luna average into one GPT-5.6 signal, Fable/Opus into one Claude signal, and K3 and GLM each contribute one. Self-judgments are excluded. The Gemini 3.6 Flash run added an eighth judge — Gemini 3.6 itself — whose verdicts on its own work and on Opus 5 pool with Gemini 3.5's into one joint Gemini signal. The Opus 5 run added a ninth, Opus 5 itself, whose two self-verdicts are published but unmapped, so every row's raw composite stays a like-for-like seven-signal mean. The DeepSeek V4 Flash, Grok 4.6, Gemini 3.7 Flash, Qwen 3.8 2.4T and GLM 5.3 runs used the same seven-signal panel; Opus 5 judged each, those verdicts likewise published but unmapped. The joint Gemini verdicts stay published but get zero calibrated weight — their 9.2–10.0 range provides almost no discrimination. This changes the fairness of the aggregation, not the winner or model order.

Usage accounting. The study processed 255,072,887 input/output tokens across 32 contestant and 250 judge sessions. Contestants cost $81.05; judging $258.91; total recorded model and tool spend $339.96. Per-model leaderboard costs show contestant spend only, so procurement comparisons stay like-for-like. Because every run delegated to the same fast worker pool, the spread between models — under a dollar to about sixteen — is orchestrator token volume: how much the model itself read, reasoned over and wrote. A model with a lower per-token rate can still cost more per job if it thinks longer, which is why the cost column doesn't track the published price list.

Evidence base. The published table is sixteen attributable entries from a seventeen-candidate cycle; the invalid Qwen swarm result is disclosed below. It sits on hundreds of internal benchmark and development runs accumulated over roughly three months. Because prompts, models and the harness evolved over that period, those runs aren't pooled into the displayed scores. They make the broad pattern less surprising; they don't turn this cross-section into a repeated-trials estimate.

Limitations, plainly. Each displayed estimate is one pinned result per model per case — scores carry roughly ±0.3 of run-to-run noise, small gaps are ties, and defect counts aren't failure rates. Provenance: four entries use July 17 outputs (Fable, Opus, Luna, Gemini 3.5), three July 18 (Sol, GLM, Qwen), two July 19 (K3, Inkling), one July 22 (Gemini 3.6), one July 25 (Opus 5), one August 1 (DeepSeek V4 Flash — the 0731 re-post-trained revision), one August 13 (Grok 4.6), two August 14 (Gemini 3.7 Flash — case 1 judged via same-day backfill after the harness idle-stream timeout — and Qwen 3.8 2.4T — case 2’s panel re-ran via same-day backfill after a dev-server reload fault killed the in-run panel; both had all artifacts and hard checks complete before backfill), and one August 19 (GLM 5.3 via OpenRouter; two isolated case bots, then merged). Provider load varies. Judges are themselves models: the panel's biases are measurable (Figure 2), one judge needed an automated correction pass after falsely reporting an empty workspace, and rationales — however well-verified — inherit their authors' blind spots. Family averaging reduces duplicate architecture weight but can't remove correlated blindness: related judges may miss the same error classes together, and z-scoring calibrates severity, not truth. There was no model-without-harness control and no Beast-mode treatment, so infer neither uplift from these scores. Two rows carry disclosures: Kimi K3's result pairs two adjacent runs (case 1 and case 2 each completed and fully judged back-to-back on July 19, after we diagnosed why earlier attempts died — a harness bug, not the model: the stream reader ignored keepalive pings during K3's legitimately long tool-history turns and killed live requests as idle; both paired cases scored 8.4 with complete panels). And a swarm-mode Qwen run was excluded after it delegated every worker to GPT-5.6 Sol against explicit instructions, making its otherwise excellent output unattributable.

Auditable, not externally reproducible. The public evidence packages let outsiders inspect each contestant's memos, models, simulation outputs and full judge record. They can't independently rerun the study — the harness is our production system — so we claim auditability of the published artifacts, not external reproducibility. Questions or corrections: community@strategize.inc.