# Can it Run Alfrada?

_Published 20 July 2026 · Runs July 17–19, 2026_

**Interactive page:** [https://strategizelabs.com/can-it-run-alfrada/](https://strategizelabs.com/can-it-run-alfrada/)

![Can it Run Alfrada?](https://strategizelabs.com/can-it-run-alfrada/social-preview-20260720-auditable.png)

Nine frontier models ran inside the same production Alfrada harness, completed two real multidisciplinary jobs, and shipped artifacts that seven model judges audited file by file.

---

## At a glance

| | |
|---|---|
| Models tested | 9 |
| Judge verdicts | 126 |
| Tokens processed | 129.8M |
| All-in spend | $182 |
| Contestant spend | $48.11 |
| Judge spend | $133.58 |

---

## Executive summary

**This is a management-risk test, not another isolated model quiz.** Nine models each completed two long-horizon knowledge-work jobs inside the same production Alfrada harness. Seven independent model auditors then opened every deliverable, replayed calculations, and challenged the evidence chain. The raw artifacts and every verdict are published on the interactive page.

**The answer is not “all frontier models are good.”** Eight cleared a practical 6.0 usability floor. Inkling scored 5.60 and failed on core arithmetic, internal consistency, and reproducibility. Sol delivered the best audited quality; Luna the strongest cost-speed trade; K3 the best quality-integrity value balance.

- **Why it matters:** Managers buy decisions, not benchmark points.
- **Why this is different:** The unit of evaluation is the complete, inspectable work product — sources, calculations, memo, deck, charts, and schedules.
- **What to do:** Choose models by quality, decision risk, and cost together; require an evidence package; review the load-bearing chain before the prose.

---

## What we test

A model, inside a fixed production harness, doing a real job. The model supplies judgment and reasoning. Alfrada supplies planning, nineteen live tools, isolated workspaces, code execution, file creation, and delivery. Because the harness stayed fixed, the ranking shows how different models operate inside the same system — not raw model intelligence in isolation.

### Case 1 — The decision package

> “Should NorthLine, a $12M-ARR logistics SaaS, enter the UAE in Q3 2026?”

Market & competitor intelligence, a computed 24-month financial model, a ≥1,500-word decision memo, a ≥10-slide board deck, a short audio briefing, and 30/60/90-day review scaffolding.

### Case 2 — The prediction

> “Forecast global semiconductor capex 2026–2030 — as a distribution, not a number.”

Five quantified drivers with named sources, a ≥10,000-trial Monte Carlo with a printed seed and JSON artifact, fan/tornado/calibration charts, and a ≥1,500-word memo with invalidation scenarios.

**Judges:** Kimi K3, GLM-5.2, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Luna, Claude Fable 5, and Claude Opus 4.8 — 126 verdicts total.

---

## Results — leaderboard

Rank uses family-balanced judge z-scores; raw composite is the familiar 0–10 mean. Per-model costs are contestant spend only.

| Rank | Model | Composite | Case 1 | Case 2 | Cost | Time |
| ---: | --- | ---: | ---: | ---: | ---: | ---: |
| 1 | [GPT-5.6 Sol](https://strategizelabs.com/can-it-run-alfrada/#m/sol) | **8.84** | 8.77 | 8.90 | $15.41 | 51 min |
| 2 | [Claude Fable 5](https://strategizelabs.com/can-it-run-alfrada/#m/fable) | **8.54** | 8.57 | 8.51 | $11.29 | 42 min |
| 3 | [Kimi K3](https://strategizelabs.com/can-it-run-alfrada/#m/k3) | **8.40** | 8.40 | 8.40 | $3.54 | 81 min |
| 4 | [Claude Opus 4.8](https://strategizelabs.com/can-it-run-alfrada/#m/opus) | **8.19** | 8.29 | 8.09 | $5.15 | 28 min |
| 5 | [GPT-5.6 Luna](https://strategizelabs.com/can-it-run-alfrada/#m/luna) | **8.04** | 8.29 | 7.79 | $0.92 | 10 min |
| 6 | [GLM-5.2](https://strategizelabs.com/can-it-run-alfrada/#m/glm) | **7.95** | 7.90 | 8.00 | $4.12 | 56 min |
| 7 | [Gemini 3.5 Flash](https://strategizelabs.com/can-it-run-alfrada/#m/gemini) | **7.07** | 6.73 | 7.40 | $2.30 | 11 min |
| 8 | [Qwen 3.7 Max](https://strategizelabs.com/can-it-run-alfrada/#m/qwen) | **7.05** | 6.50 | 7.60 | $3.26 | 47 min |
| 9 | [Inkling](https://strategizelabs.com/can-it-run-alfrada/#m/inkling) | **5.60** | 4.79 | 6.41 | $2.13 | 17 min |

**Manager translation:** 6.0 is the practical cutoff. Eight models scored 7.05–8.84 and produced usable first passes with varying review burden. Inkling scored 5.60 — below six, commission substantial rework.

### How to choose

| Need | Pick | Why |
| --- | --- | --- |
| Best audited quality | **Sol · 8.84** | Deepest evidence package; no catalogued load-bearing defect |
| Best cost & speed | **Luna · $0.92** | 8.04 in ~10 minutes; honest framing; thinner evidence staging |
| Best all-round value | **K3 · 8.40** | $3.54, verified core reasoning, no catalogued load-bearing defect |
| Balanced middle | **GLM · 7.95** | Strong reproducibility at $4.12; review the case-1 breakeven claim |

---

## Load-bearing errors

A **load-bearing error** sits underneath the recommendation — correct it and the economics, confidence, or proposed action materially changes.

By that cut, **Sol, Luna, and K3** combine zero catalogued load-bearing errors with evidence the panel could inspect. **Fable, Opus, and GLM** each carry one. **Gemini** carries two. **Inkling** carries three and fails the cutoff.

### Same forecast question, 2.3× apart

Each model's P50 five-year CAGR for global semiconductor capex (with P10–P90 where reported):

| Model | P10 | P50 | P90 | Seed replayed |
| --- | ---: | ---: | ---: | :---: |
| Sol | — | **9.1%** | — | yes |
| Fable | 2.2% | **6.9%** | 11.9% | yes |
| K3 | 1.5% | **5.8%** | 10.1% | yes |
| Opus | 2.5% | **10.9%** | 19.2% | no |
| Luna | 3.2% | **5.9%** | 8.6% | yes |
| GLM | 5.1% | **10.8%** | 16.4% | yes |
| Gemini | 11.7% | **13.6%** | 15.7% | no |
| Inkling | -3.3% | **10.3%** | 25.7% | no |

K3's ~5.8% and Gemini's ~13.6% answer the same prompt from the same source landscape. For planning use, the disagreement *is* the finding.

---

## Evidence by model

Charts below are hosted on the live study page. Click through for full-size PNGs, judge verdicts, and downloadable work packages.

### 1. GPT-5.6 Sol — 8.84

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/sol) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/sol.zip)

- Case 1: **8.77** · $8.62 · 25 min
- Case 2: **8.90** · $6.79 · 26 min
- Combined contestant spend: **$15.41**
- Run note: clean full run

**Cosmetic / other**

- Slide 8 mislabels the cumulative-breakeven KPI as “>$48m” where it means “>M48”.
- Excel workbook outputs are hard-coded rather than formula-driven despite claiming inputs update outputs automatically.
- Audio briefing runs 62 seconds against the 40–50s spec; staged RTA/FTA screenshots don't visibly contain their cited figures.
- Case-2 driver-to-growth effect matrix, correlations, and several sigmas are judgmental overlays rather than sourced estimates — disclosed, not hidden.

> Verified, not just claimed: I opened the memo, model CSVs, workbook, slides JSON, research notes, and charts, and independently recomputed the arithmetic — TAM/SAM derivations, CAC ($37,143 = $780k S&M ÷ 21 wins), paybacks (19.0/50.0/8.1 mo), and the FX contributions all reproduce exactly.
>
> — *Claude Fable 5, case 1*

> I re-executed the staged forecast_model.py (100,000 trials, seed 20260717, PCG64) and reproduced every P10/P50/P90, mean, sigma, CAGR, sensitivity and regional figure bit-for-bit against the saved JSON.
>
> — *Kimi K3, case 2*

![Uae arr scenarios (case 1) — Sol](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_1__graph_uae_arr_scenarios.png)

*Uae arr scenarios (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_1__graph_uae_arr_scenarios.png)

![Uae cumulative cash (case 1) — Sol](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_1__graph_uae_cumulative_cash.png)

*Uae cumulative cash (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_1__graph_uae_cumulative_cash.png)

![Capex calibration (case 2) — Sol](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_2__graph_capex_calibration.png)

*Capex calibration (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_2__graph_capex_calibration.png)

![Capex fan (case 2) — Sol](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_2__graph_capex_fan.png)

*Capex fan (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_2__graph_capex_fan.png)

![Capex tornado (case 2) — Sol](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_2__graph_capex_tornado.png)

*Capex tornado (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/sol/case_2__graph_capex_tornado.png)

---

### 2. Claude Fable 5 — 8.54

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/fable) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/fable.zip)

- Case 1: **8.57** · $7.57 · 31 min
- Case 2: **8.51** · $3.72 · 11 min
- Combined contestant spend: **$11.29**
- Run note: sweep run

**Load-bearing defects**

- (case 1, judge Sol) The claimed month-26 operating breakeven is only the first positive month — months 27 and 30 dip negative again; sustained breakeven is ~month 31.

**Cosmetic / other**

- The CBUAE peg screenshot shows only a disclaimer overlay, not the cited 3.6725 rate; one Dubai South capture is a 404.
- The staged memo PDF omits the risks/reversal sections that appear in the deck and chat text — the artifact is thinner than claimed.
- The promised simulation code listing was never staged; verification was only possible because raw 20,000-trial arrays were saved.

> The npz contains 20,000 real trial paths per year and segment, and recomputing P10/P50/P90/mean/sigma from the raw samples reproduces the summary JSON to within rounding — the Monte Carlo was actually executed, not narrated.
>
> — *Kimi K3, case 2*

> The main evidence weakness is that several staged screenshots do not support the cited figures: the Dubai South capture shows AED 11,375 rather than the cited AED 12,500.
>
> — *GPT-5.6 Luna, case 1*

![Arr scenarios (case 1) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_1__graph_arr_scenarios.png)

*Arr scenarios (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_1__graph_arr_scenarios.png)

![Cash base v2 (case 1) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_1__graph_cash_base_v2.png)

*Cash base v2 (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_1__graph_cash_base_v2.png)

![Tam sam som (case 1) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_1__graph_tam_sam_som.png)

*Tam sam som (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_1__graph_tam_sam_som.png)

![Calibration (case 2) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_calibration.png)

*Calibration (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_calibration.png)

![Distribution 2030 (case 2) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_distribution_2030.png)

*Distribution 2030 (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_distribution_2030.png)

![Fanchart (case 2) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_fanchart.png)

*Fanchart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_fanchart.png)

![Segment mix (case 2) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_segment_mix.png)

*Segment mix (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_segment_mix.png)

![Tornado (case 2) — Fable](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_tornado.png)

*Tornado (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/fable/case_2__graph_tornado.png)

---

### 3. Kimi K3 — 8.40

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/k3) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/k3.zip)

- Case 1: **8.40** · $2.02 · 46 min
- Case 2: **8.40** · $1.52 · 35 min
- Combined contestant spend: **$3.54**
- Run note: paired runs (see method)

**Cosmetic / other**

- The ~300-operator market denominator rests on unsupported filtering assumptions — the most consequential soft number under an otherwise verified model.
- Several staged screenshots are 404/error pages or do not visibly contain the cited figures; citations often name organizations rather than inline primary sources.

> The memo gives a clear conditional no-go… independent arithmetic checks confirm that customer counts reproduce ARR, SAM reconciles, CAC payback is calculated correctly, and ongoing S&M/CAC is included in monthly EBITDA.
>
> — *GPT-5.6 Sol, case 1*

> The package is substantively strong: the memo is decision-oriented… the recommendation is calibrated, and the model arithmetic reconciles customer counts, ACV, TAM/SAM/SOM, payback, and the stated FX swing.
>
> — *GPT-5.6 Luna, case 1*

![Arr cash (case 1) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_arr_cash.png)

*Arr cash (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_arr_cash.png)

![Cac (case 1) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_cac.png)

*Cac (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_cac.png)

![Customers (case 1) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_customers.png)

*Customers (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_customers.png)

![Tamsamsom (case 1) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_tamsamsom.png)

*Tamsamsom (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_1__graph_tamsamsom.png)

![Calibration (case 2) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_calibration.png)

*Calibration (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_calibration.png)

![Fan chart (case 2) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_fan_chart.png)

*Fan chart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_fan_chart.png)

![Regional split (case 2) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_regional_split.png)

*Regional split (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_regional_split.png)

![Tornado (case 2) — K3](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_tornado.png)

*Tornado (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/k3/case_2__graph_tornado.png)

---

### 4. Claude Opus 4.8 — 8.19

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/opus) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/opus.zip)

- Case 1: **8.29** · $3.45 · 21 min
- Case 2: **8.09** · $1.70 · 7 min
- Combined contestant spend: **$5.15**
- Run note: sweep run

**Load-bearing defects**

- (case 1, judge Sol) The board is asked to approve a partner-led $750K pilot, but the financial files model only direct acquisition — partner commissions, the claimed 25–35% CAC reduction, and the spending cap are never modeled.

**Cosmetic / other**

- Several staged evidence screenshots (ADGM, u.ae) are homepage navigation chrome containing no figures; the ADGM “~$1,800 all-in” framing omits listed additional costs.
- Case-2 JSON records driver names but not the numeric priors the memo claims it contains; simulated contribution arrays don't match the memo's stated priors.

> The standout finding that cumulative breakeven is never reached within 24 (or 36) months is an honest, CAC-inclusive result with an explicitly stated definition, driving a calibrated 'conditional/partner-led go' rather than boilerplate.
>
> — *Claude Fable 5, case 1*

> The central decision is not actually modeled: the board is asked to approve a partner-led $750K pilot, but the financial files model only direct acquisition and omit partner commissions.
>
> — *GPT-5.6 Sol, case 1*

![Breakeven (case 1) — Opus](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_1__graph_breakeven.png)

*Breakeven (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_1__graph_breakeven.png)

![Customer ramp (case 1) — Opus](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_1__graph_customer_ramp.png)

*Customer ramp (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_1__graph_customer_ramp.png)

![Tamsamsom (case 1) — Opus](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_1__graph_tamsamsom.png)

*Tamsamsom (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_1__graph_tamsamsom.png)

![Calibration (case 2) — Opus](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_2__graph_calibration.png)

*Calibration (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_2__graph_calibration.png)

![Fan chart (case 2) — Opus](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_2__graph_fan_chart.png)

*Fan chart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_2__graph_fan_chart.png)

![Tornado (case 2) — Opus](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_2__graph_tornado.png)

*Tornado (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/opus/case_2__graph_tornado.png)

---

### 5. GPT-5.6 Luna — 8.04

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/luna) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/luna.zip)

- Case 1: **8.29** · $0.50 · 6 min
- Case 2: **7.79** · $0.42 · 4 min
- Combined contestant spend: **$0.92**
- Run note: sweep run

**Cosmetic / other**

- The FX sensitivity CSV's “usd_equivalent” column is AED-scaled — off by exactly 3.6725× (the memo's own dollar figures are correct).
- The staged TDRA evidence file is an unreadable raw binary PDF dump; audio runs 74 seconds against the 40–50s spec.
- Case-2 calibration is thin: 2025 is the model anchor (trivially zero error), leaving one real backtest year; the DOCX ships with an unpopulated table-of-contents field.

> The model honestly refuses to fabricate a SaaS TAM where none exists… breakeven is explicitly defined to include ongoing CAC and honestly reported as 'not reached in 24 months.'
>
> — *Claude Opus 4.8, case 1*

> Rebuilding the simulation from only the JSON's documented priors, anchor, seed 20260717 and 20,000 trials reproduces every reported P10/P50/P90 within MC noise — no fabricated numbers anywhere.
>
> — *Claude Fable 5, case 2*

![Northline uae arr ramp (case 1) — Luna](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_1__graph_northline_uae_arr_ramp.png)

*Northline uae arr ramp (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_1__graph_northline_uae_arr_ramp.png)

![Capex calibration (case 2) — Luna](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_2__graph_capex_calibration.png)

*Capex calibration (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_2__graph_capex_calibration.png)

![Capex fan (case 2) — Luna](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_2__graph_capex_fan.png)

*Capex fan (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_2__graph_capex_fan.png)

![Capex tornado (case 2) — Luna](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_2__graph_capex_tornado.png)

*Capex tornado (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/luna/case_2__graph_capex_tornado.png)

---

### 6. GLM-5.2 — 7.95

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/glm) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/glm.zip)

- Case 1: **7.90** · $2.02 · 36 min
- Case 2: **8.00** · $2.10 · 20 min
- Combined contestant spend: **$4.12**
- Run note: clean full run

**Load-bearing defects**

- (case 1, judge Fable) The memo asks the board to commit ~$4.3M partly on a “month 30–33 breakeven” that the model's own JSON records as null — extrapolating the actual trajectory puts CAC-inclusive breakeven beyond month 40.

**Cosmetic / other**

- Two base-case CSVs contradict each other at month 24 (58 vs 58.72 customers; $2.088M vs $2.114M ARR), undermining the single-source-of-truth claim.
- The $462M TAM rests on an uncited 1.4% software-intensity estimate and an arbitrary 25% SAM factor.
- Case-2 backtest feeds realized growth rates in as the means, making the −0.5%/−0.2% “deviations” near-tautological — disclosed on slide 10, but oversold in the framing.

> My independent code_executor replication from the JSON's priors, seed 20260717, and the $166B anchor reproduces every published P10/P50/P90, mean, sigma, and CAGR figure to within rounding.
>
> — *Claude Fable 5, case 2*

> The repeated 'breakeven ~Month 30–33' projection is unsupported and contradicts the model's own data: the JSON records breakeven=null… a material hand-waved number underpinning the board's ask.
>
> — *Claude Fable 5, case 1*

![Arr trajectory (case 1) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_arr_trajectory.png)

*Arr trajectory (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_arr_trajectory.png)

![Cash position (case 1) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_cash_position.png)

*Cash position (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_cash_position.png)

![Customers (case 1) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_customers.png)

*Customers (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_customers.png)

![Market funnel (case 1) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_market_funnel.png)

*Market funnel (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_market_funnel.png)

![Tam sam som (case 1) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_tam_sam_som.png)

*Tam sam som (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_1__graph_tam_sam_som.png)

![Calibration (case 2) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_2__graph_calibration.png)

*Calibration (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_2__graph_calibration.png)

![Fan chart (case 2) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_2__graph_fan_chart.png)

*Fan chart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_2__graph_fan_chart.png)

![Tornado (case 2) — GLM](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_2__graph_tornado.png)

*Tornado (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/glm/case_2__graph_tornado.png)

---

### 7. Gemini 3.5 Flash — 7.07

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/gemini) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/gemini.zip)

- Case 1: **6.73** · $1.19 · 6 min
- Case 2: **7.40** · $1.11 · 6 min
- Combined contestant spend: **$2.30**
- Run note: sweep run

**Load-bearing defects**

- (case 1, judge GLM) Breakeven is computed on revenue (MRR) instead of gross profit — by the memo's own definition the branch never breaks even within 24 months, yet the GO recommendation leans on the month-21 claim. Found independently by four judges.
- (case 2, judge Opus) The case-2 backtest hand-sets 2024–25 driver parameters with hindsight, then sells the sub-1% fit as “out-of-sample” with “92.0% confidence” — the calibration is partly circular and the confidence figure is unsupported.

**Cosmetic / other**

- No staged evidence pages behind any citation; competitor pricing rests solely on secondary aggregators (G2, SourceForge); the load-bearing 1,500-firm TAM count is unsourced.
- The memo overclaims its own length (“3,527 words” vs 1,637 measured) and slide 9's “±$15.4B” AI sensitivity contradicts the $77.8B range everywhere else.

> A material breakeven methodology error undermines the core GO recommendation: the memo defines breakeven as gross profit exceeding operating expenses, but the model actually computes it using revenue — by the stated definition, the branch never breaks even in 24 months.
>
> — *GLM-5.2, case 1*

> The backtest is not truly out-of-sample — the 2024/25 driver parameters were hand-set with hindsight, so the sub-1% calibration error is partly circular yet is sold as 'exceptional validation.'
>
> — *Claude Opus 4.8, case 2*

![Arr ramp (case 1) — Gemini](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_1__graph_arr_ramp.png)

*Arr ramp (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_1__graph_arr_ramp.png)

![Cumulative cash flow (case 1) — Gemini](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_1__graph_cumulative_cash_flow.png)

*Cumulative cash flow (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_1__graph_cumulative_cash_flow.png)

![Calibration plot (case 2) — Gemini](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_2__graph_calibration_plot.png)

*Calibration plot (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_2__graph_calibration_plot.png)

![Capex fan chart (case 2) — Gemini](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_2__graph_capex_fan_chart.png)

*Capex fan chart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_2__graph_capex_fan_chart.png)

![Tornado chart (case 2) — Gemini](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_2__graph_tornado_chart.png)

*Tornado chart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/gemini/case_2__graph_tornado_chart.png)

---

### 8. Qwen 3.7 Max — 7.05

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/qwen) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/qwen.zip)

- Case 1: **6.50** · $1.66 · 22 min
- Case 2: **7.60** · $1.60 · 25 min
- Combined contestant spend: **$3.26**
- Run note: clean full run

**Cosmetic / other**

- Inspected screenshots contain titles and descriptive text but no readable market sizes, pricing, or licensing figures — the evidence pack doesn't substantiate its load-bearing claims.
- The $300M market and $35M segment are constructed from secondary research and unsupported allocation assumptions; DIC/ADGM cost claims lack staged first-party evidence.
- Core arithmetic is clean — ARR, scenario-net, CAC-payback, CAGR, and FX all replay correctly. The penalty is provenance, not correctness.

> The basic ARR, scenario-net, CAC-payback, CAGR, and FX arithmetic checks out. However, the required primary-source evidence standard is largely unmet.
>
> — *GPT-5.6 Sol, case 1*

> Inspection of four screenshots found titles and descriptive text but no readable market sizes, pricing, licensing costs, or other load-bearing figures.
>
> — *GPT-5.6 Luna, case 1*

![Calibration plot (case 2) — Qwen](https://strategizelabs.com/can-it-run-alfrada/assets/qwen/case_2__graph_calibration_plot.png)

*Calibration plot (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/qwen/case_2__graph_calibration_plot.png)

![Fan chart (case 2) — Qwen](https://strategizelabs.com/can-it-run-alfrada/assets/qwen/case_2__graph_fan_chart.png)

*Fan chart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/qwen/case_2__graph_fan_chart.png)

![Tornado sensitivity (case 2) — Qwen](https://strategizelabs.com/can-it-run-alfrada/assets/qwen/case_2__graph_tornado_sensitivity.png)

*Tornado sensitivity (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/qwen/case_2__graph_tornado_sensitivity.png)

---

### 9. Inkling — 5.60

[Open on study page](https://strategizelabs.com/can-it-run-alfrada/#m/inkling) · [Download evidence package](https://strategizelabs.com/can-it-run-alfrada/packages/inkling.zip)

- Case 1: **4.79** · $1.19 · 13 min
- Case 2: **6.41** · $0.95 · 4 min
- Combined contestant spend: **$2.13**
- Run note: clean full run

**Load-bearing defects**

- (case 1, judge Fable) Reported 24-month revenue sums annualized ARR monthly — a ~12× inflation — producing headline claims of base breakeven in “month 1” and downside breakeven in “month 3.”
- (case 1, judge Luna) Three artifacts tell three stories: the JSON reports $288K base cost, the memo $539K, and the charts visually show breakeven near month 10.
- (case 2, judge GLM) The case-2 methodology collapses under reproduction: eight plausible interpretations of the stated model were tested and none reproduced the published numbers, despite a printed seed.

**Cosmetic / other**

- The purported top-five competitor list mixes software vendors, logistics operators, and vaguely identified firms; no incumbent pricing is verified.

> The financial core fails replay: reported 24-mo revenue sums annualized ARR monthly (~12x inflation), so the headline claims — base breakeven “month 1” and “even downside breaks even in month 3” — are wrong.
>
> — *Claude Fable 5, case 1*

> The core methodological claim collapses under reproduction: I tested 8 plausible interpretations of the stated model and none reproduced the published numbers.
>
> — *GLM-5.2, case 2*

![Base rev cost (case 1) — Inkling](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_1__graph_base_rev_cost.png)

*Base rev cost (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_1__graph_base_rev_cost.png)

![Downside rev cost (case 1) — Inkling](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_1__graph_downside_rev_cost.png)

*Downside rev cost (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_1__graph_downside_rev_cost.png)

![Upside rev cost (case 1) — Inkling](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_1__graph_upside_rev_cost.png)

*Upside rev cost (case 1)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_1__graph_upside_rev_cost.png)

![Calibration backtest (case 2) — Inkling](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_2__graph_calibration_backtest.png)

*Calibration backtest (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_2__graph_calibration_backtest.png)

![Fan chart (case 2) — Inkling](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_2__graph_fan_chart.png)

*Fan chart (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_2__graph_fan_chart.png)

![Tornado sensitivity (case 2) — Inkling](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_2__graph_tornado_sensitivity.png)

*Tornado sensitivity (case 2)* · [full size](https://strategizelabs.com/can-it-run-alfrada/assets/inkling/case_2__graph_tornado_sensitivity.png)

---

## What this means in practice

The practical result is a **controllable quality floor, not guaranteed correctness.** The harness made every model ship the required package. Eight produced useful drafts; one demonstrated that completeness checks cannot rescue broken reasoning.

- **Commissioning:** Route by consequence — Sol where misses are expensive, K3 for balance, Luna for fast volume, GLM where reproducible modelling matters.
- **Review:** Audit the spine first — provenance, definitions, arithmetic, simulation replay, cross-artifact consistency.
- **Accountability:** AI drafts; a person owns the decision.

---

## Method & limits (short)

Runs executed July 17–19, 2026 against a production Alfrada build, temperature 0, isolated bot user per model, $15 budget cap per run. Raw composite is the mean of all seven judges across both cases. Calibrated rank z-scores judges, averages GPT-5.6 and Claude family signals, excludes self-judgments, and gives Gemini zero ranking weight (near-uniform marks).

Each displayed estimate is one pinned result per model per case — treat small gaps as ties; defect counts are observations, not rates. Packages are auditable; the harness itself is not externally reproducible.

Questions: [community@strategize.inc](mailto:community@strategize.inc)

---

Strategize Labs · [Can it Run Alfrada?](https://strategizelabs.com/can-it-run-alfrada/) · 20 July 2026. All judge quotes verbatim from recorded rationales; all replayed numbers traceable to the packaged artifacts.
