Sol · 8.84
Best-checked work. Use it when a wrong answer would be expensive.
Sixteen models.
Two real jobs.
Every file audited.
Fifteen of sixteen models delivered useful work. One looked finished but got the decision wrong. Here’s what to use.
See the results ↓ Markdown copy Or download the full study for sharing, Notion, or LLMs
In short: a real AI model, dropped into Alfrada — our actual product — doing a real job. Same setup for every model, so the ranking reflects how they handle the work, not a quiz score. These are long-form jobs: hundreds of steps of research, modelling, writing and delivery — the kind of work that takes at least 45 minutes to complete, not a short quiz.
The model does the thinking; Alfrada gives it the tools: planning, nineteen live tools, a private workspace, the ability to run code, and to create and hand over files. Because every model used the exact same setup, the ranking shows how they handle the same real job — not how “smart” they look on a quiz.
It's not a coding test either. The job mixes research, judgment, number-crunching, writing and delivery all at once. A good setup should even the odds between models: guardrails and checks keep the work complete, on-spec and easy to inspect, so what's left to judge is the quality of the model's own thinking.
Every model ran in the exact same Alfrada setup: same tools, same instructions, a fresh, empty workspace, and no memory of any other run. No model could peek at another's answer or quietly hand the hard parts to a smarter model. Then we gave each one two long-form jobs — hundreds of steps, at least 45 minutes of work — of the kind a real strategy team might actually order.
“Should NorthLine, a $12M-a-year logistics software company, expand into the UAE in mid-2026?”
“Forecast how much the world will spend building computer chips through 2026–2030 — as a range of outcomes, not a single number.”
Automatic checks confirm the basics: every required file is there, the word and slide counts are met, the tools ran without errors, and the schedules work.
A panel of judges opens everything, re-runs the math, checks the sources, and scores the finished work from 0 to 10.
The judges are seven other AI models: Kimi K3, GLM-5.2, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Luna, Claude Fable 5 and Claude Opus 4.8. They don't just skim the final answer — they open the files, re-run the math and the simulations, recheck the costs, and even look at the slides and charts. Every model was judged by all seven.
All seven judged each of the nine original contestants. The 22 July Gemini 3.6 Flash and 25 July Opus 5 runs added Gemini 3.6 as a judge — pooled with Gemini 3.5 into one joint Gemini signal — and Opus 5 also judged itself, a verdict published below but carrying no scoring weight. The DeepSeek V4 Flash (1 Aug), Grok 4.6 (13 Aug), Gemini 3.7 Flash and Qwen 3.8 2.4T (14 Aug) and GLM 5.3 (19 Aug) runs used the same expanded panel, with Opus 5's verdicts likewise published but unmapped. Verdict and session totals are recomputed in the published data payload. Qwen, Inkling and DeepSeek Flash are contestants, not judges: Inkling's own work fell below the usability cutoff and lacked the verification discipline an evaluator needs.
In plain terms: think of 6.0 as the pass mark, and the score as a grade out of 10 — not a percentage. Fifteen models scored 7.05–8.84 and gave a usable first draft, some needing more review than others. Inkling scored 5.60: it looked finished, but the math behind its recommendation was wrong, its key files disagreed with each other, and its forecast couldn't be reproduced. Below six, treat the work as a rough starting point, not a draft to polish.
Two models finished on top: Claude Opus 5 leads the fairness-adjusted ranking, and GPT-5.6 Sol has the highest raw score (8.84). Neither was grading itself — a model's vote on its own work never counts. The gap between them is smaller than the normal wobble you'd see from one run to the next, so treat them as tied; what separates them is the judges' notes: Sol made no mistake big enough to change its recommendation, while Opus 5 made the most on the board. Grok 4.6 comes third on its first try — the single best score on the business-decision job (8.87) at a quarter of the top models' cost, though a weaker forecast pulls its overall to 8.63. Fable is fourth. GLM 5.3 lands fifth (8.55) and is the cheapest of the top five, at $2.54. Then K3 and Qwen 3.8 2.4T — Qwen's forecast could be re-run exactly by three separate judges, but slips on the first job. Luna is the bargain of the group; DeepSeek V4 Flash reaches 7.85 for a tiny fraction of the usual cost. Every model cleared the basic delivery checks — which didn't stop some of the finished work from being unusable.
Who judges whom
Rows are the judges; columns are the models being judged. Click any cell to read what that judge said, in full.
Cell shading spans scores 3–10 across both cases. The right-hand column is each judge's mean. Sol (5.85) and Luna (6.24) are materially harsher than the joint Gemini row (9.93). Corner-ticked cells contain self-judgments; they are visible but excluded from calibrated rank.
The matrix is where this study earns its keep. Raw means reward generous judges and duplicate architecture families. Calibrated rank instead z-scores each judge within each case, averages Sol/Luna into one GPT-5.6 signal and Fable/Opus into one Claude signal, and gives K3 and GLM one signal each. Self-votes are excluded. The joint Gemini row stays fully visible but gets zero ranking weight — its near-uniform marks add almost no discrimination.
How models grade their own work
Each judge's score for its own work, next to what everyone else gave it (both jobs combined). Inkling and Qwen aren't shown — they weren't judges.
Both Geminis rated their own work far above their peers' view. Luna is the opposite: it graded its own case-2 submission 4.5/10 while six peers averaged 7.6. Sol also self-graded below its peer mean. Self-votes are diagnostic only and do not affect rank.
Treat 6.0 as the pass mark, and the score as a grade out of 10. From 6–7, it's a draft that needs supervision; 7–8, useful with a targeted expert review; above 8, more of it survives the audit — but no score removes the need for a human to own a big decision. And “no mistakes found” means none turned up in this run, not that the model never makes them.
Some things come down to taste — how good the writing, the recommendation or the slides are. But the math doesn't. Either a model's simulation could be re-run and gave the same answer, or it couldn't. Either its cost figures added up, or they didn't. Either a screenshot actually showed the number it was quoted for, or it didn't. Those yes/no checks give us solid ground under an otherwise judgment-based score.
That's why these scores don't line up with a coding benchmark or an exam percentage. A coding test asks whether one small fix passes; here, writing code is just one tool in a much bigger job. The model has to research, build a model, decide, explain, and hand over several files that all agree with each other — then survive judges opening every one.
Alfrada also sets a floor. Its built-in checks and repeatable tools keep every model's output at a steady minimum, so the real differences show up above that floor. (We didn't test the models with Alfrada stripped away, so we can't put a number on how much the setup helps.) What the study does show is the gap between merely delivering and actually holding up: every model cleared the basic checks, yet judged quality ranged from 5.60 to 8.84.
Making things up is only part of it. Invented facts are penalised — but the harder test is spotting a break in the reasoning even when the answer looks polished: evidence that doesn't actually back the claim, math that contradicts the memo, or a check that just proves itself. That's a tougher bar than “did it run?”
The worst mistakes are the ones the recommendation actually rests on — fix one and the money, or the advice itself, changes. On that measure: Sol, Luna and K3 had none; Fable, Opus and GLM had one each; Gemini 3.5 had two; Inkling had three.
Why 5.60 is a failed job, not just last place. Inkling added up a full year's revenue every single month, overstating two-year revenue by about 12× and claiming the business broke even in its first three months. Its data file, memo and charts don't agree on the costs or the timing. On the forecast job, eight reasonable attempts to follow its stated method all failed to reproduce its published numbers. It handed in a full package — but the thinking behind it fell apart the moment anyone checked.
By that measure, Sol, Luna and K3 had none — and left evidence the judges could actually check. Qwen 3.7 Max also got its core math right, but its evidence didn't back up its market claims, so it isn't as safe a bet. Fable, Opus and GLM each had one: Fable called month 26 “break-even” when it was really the first good month before two more bad ones; Opus asked the board to approve a partner deal whose economics its own model never included; GLM told the board break-even came in month 30–33 when its own data file said it never came at all. Gemini 3.5 had two, one in each job. Gemini 3.6 — the 22 July update — fixed its predecessor's break-even math but added a new slip: a chart built on numbers its own model doesn't produce. Inkling had three, and failed the pass mark.
A second warning: a polished answer can outlast the facts under it. Gemini 3.5's memo says break-even means profit beating costs — then quietly works it out on revenue instead, skipping the 20% it costs to deliver. By its own definition the UAE branch never breaks even within two years; the memo says month 21 and recommends going ahead partly on that number. Four judges caught it independently. On the forecast job, the same model quietly tuned its test to fit the past, then presented that near-perfect fit as if it had predicted the future — at “92.0% confidence.” Both times, the reasoning broke before the writing did.
The models agree more than the scores suggest — fourteen of the sixteen said a cautious “yes” on the business decision, and every model whose numbers checked out agreed the branch doesn't pay for itself within two years. Where they really split is the forecast:
Same question, same sources, 5.1× apart
Each model's best-guess five-year growth rate for global chip-building spend, with its own likely range where given.
Qwen 3.8's 5.1% and Gemini 3.7 Flash's 25.8% answer the same prompt from the same sources. The spread comes from anchor choice and prior construction, and is invisible in the quality scores: rigorous, reproducible runs can still disagree sharply. For planning, that disagreement is the finding.
The takeaway: a good setup guarantees a solid floor, not a correct answer. Cheaper models can do real work inside strong guardrails — as long as a person still checks the reasoning. Here's how to put that to work.
Alfrada made every model deliver the full package. Fifteen gave a useful draft; one proved that ticking every box can't save broken reasoning. Cheaper models can do real work inside strong guardrails — as long as those guardrails keep enough evidence for a person to check the reasoning.
The expert's job moves up a level. Analysts, scientists and reporters spend less time building the first draft and more time deciding which assumptions matter, re-checking the calculations that count, and sorting out conflicting evidence. Managers stop asking “is the work done?” and start asking “is the recommendation actually right?”
For higher-stakes work, Alfrada has a “Beast” mode that adds a built-in critic-and-judge loop to push borderline work up into the high 7s and 8s. We're not claiming that boost here — Beast mode was switched off for these runs.
Sol when a missed mistake is costly, K3 for balanced quality, Luna for fast, high-volume work, and GLM when you need repeatable modelling and have someone to review it.
Start with where the sources came from, the definitions, the math, re-running the simulation, and whether the files agree. Proofreading the polished text first is checking the wrong thing.
For analysts, keep the underlying models. For scientists, keep the code and settings. For reporters, keep the original sources and the link from each claim to its evidence.
This helps you choose a model and design the workflow. It doesn't hand your professional responsibility over to a score.
A benchmark you can't check is just an opinion. For every model, we publish the full record: each judge's score and reasons in full, the list of mistakes found, the charts it drew, and a downloadable pack of its actual files. It's all here — open it if you want to dig in.
Prefer a portable copy? Download the Markdown export (includes absolute URLs for every chart).
How the runs, scoring and cost accounting worked — and the honest limits of a pinned, single-machine study. For readers who want to audit the numbers.
Setup. Runs executed July 17–August 19, 2026 against a production Alfrada build on a single dev machine: one isolated bot user per run (separate workspaces, budgets and history), temperature 0, seed 42 in the case config, $15 budget cap per run. Every model ran as the orchestrator and delegated sub-agent legwork to Alfrada's fast worker pool — mostly GPT-5.6 Luna and Gemini Flash — identically across all sixteen runs. That's the configuration production users get, and what the scores measure: the model directing a job, not doing every keystroke of it. Every worker call is recorded and published in the packages. Judge sessions ran under a separate critic user with read-only tools plus code execution for replay.
Scoring. The 0–10 raw composite is the mean of the seven judge signals across both cases, kept visible for interpretability. Rank is calibrated separately: within each case, every judge is z-scored across the sixteen submissions; Sol/Luna average into one GPT-5.6 signal, Fable/Opus into one Claude signal, and K3 and GLM each contribute one. Self-judgments are excluded. The Gemini 3.6 Flash run added an eighth judge — Gemini 3.6 itself — whose verdicts on its own work and on Opus 5 pool with Gemini 3.5's into one joint Gemini signal. The Opus 5 run added a ninth, Opus 5 itself, whose two self-verdicts are published but unmapped, so every row's raw composite stays a like-for-like seven-signal mean. The DeepSeek V4 Flash, Grok 4.6, Gemini 3.7 Flash, Qwen 3.8 2.4T and GLM 5.3 runs used the same seven-signal panel; Opus 5 judged each, those verdicts likewise published but unmapped. The joint Gemini verdicts stay published but get zero calibrated weight — their 9.2–10.0 range provides almost no discrimination. This changes the fairness of the aggregation, not the winner or model order.
Usage accounting. The study processed 255,072,887 input/output tokens across 32 contestant and 250 judge sessions. Contestants cost $81.05; judging $258.91; total recorded model and tool spend $339.96. Per-model leaderboard costs show contestant spend only, so procurement comparisons stay like-for-like. Because every run delegated to the same fast worker pool, the spread between models — under a dollar to about sixteen — is orchestrator token volume: how much the model itself read, reasoned over and wrote. A model with a lower per-token rate can still cost more per job if it thinks longer, which is why the cost column doesn't track the published price list.
Evidence base. The published table is sixteen attributable entries from a seventeen-candidate cycle; the invalid Qwen swarm result is disclosed below. It sits on hundreds of internal benchmark and development runs accumulated over roughly three months. Because prompts, models and the harness evolved over that period, those runs aren't pooled into the displayed scores. They make the broad pattern less surprising; they don't turn this cross-section into a repeated-trials estimate.
Limitations, plainly. Each displayed estimate is one pinned result per model per case — scores carry roughly ±0.3 of run-to-run noise, small gaps are ties, and defect counts aren't failure rates. Provenance: four entries use July 17 outputs (Fable, Opus, Luna, Gemini 3.5), three July 18 (Sol, GLM, Qwen), two July 19 (K3, Inkling), one July 22 (Gemini 3.6), one July 25 (Opus 5), one August 1 (DeepSeek V4 Flash — the 0731 re-post-trained revision), one August 13 (Grok 4.6), two August 14 (Gemini 3.7 Flash — case 1 judged via same-day backfill after the harness idle-stream timeout — and Qwen 3.8 2.4T — case 2’s panel re-ran via same-day backfill after a dev-server reload fault killed the in-run panel; both had all artifacts and hard checks complete before backfill), and one August 19 (GLM 5.3 via OpenRouter; two isolated case bots, then merged). Provider load varies. Judges are themselves models: the panel's biases are measurable (Figure 2), one judge needed an automated correction pass after falsely reporting an empty workspace, and rationales — however well-verified — inherit their authors' blind spots. Family averaging reduces duplicate architecture weight but can't remove correlated blindness: related judges may miss the same error classes together, and z-scoring calibrates severity, not truth. There was no model-without-harness control and no Beast-mode treatment, so infer neither uplift from these scores. Two rows carry disclosures: Kimi K3's result pairs two adjacent runs (case 1 and case 2 each completed and fully judged back-to-back on July 19, after we diagnosed why earlier attempts died — a harness bug, not the model: the stream reader ignored keepalive pings during K3's legitimately long tool-history turns and killed live requests as idle; both paired cases scored 8.4 with complete panels). And a swarm-mode Qwen run was excluded after it delegated every worker to GPT-5.6 Sol against explicit instructions, making its otherwise excellent output unattributable.
Auditable, not externally reproducible. The public evidence packages let outsiders inspect each contestant's memos, models, simulation outputs and full judge record. They can't independently rerun the study — the harness is our production system — so we claim auditability of the published artifacts, not external reproducibility. Questions or corrections: community@strategize.inc.