Local AI ranking
Local AI ranking · Round 2 · 30 Sep – 03 Oct 2026

Round 2.
Same 10 tasks.
Recipes by Mia and Ash.

Same prompts, same harness (opencode 1.18.33), same sampling enforced by my proxy, one run per task, seed 1. Seven tasks run 45 minutes, tasks 2, 4 and 8 run 150. Every counted app below is clickable. Not a duel: cells differ in engine, checkpoint, context window and hardware.

9 / 9cells released
4memory classes
0cloud
How to read this page

What Tier 1, 2 and 3 mean

Every app is checked by a script that opens it and clicks through it. The checks come in three groups: must-pass (does it start and do the basics), quality (does it work properly) and top (the extra polish). The tier says how far the app got:

Tier 3 = all three groups passed. Tier 2 = must-pass and quality passed, not all top checks passed. Tier 1 = only must-pass passed: the app runs, but some quality checks failed. No counted commit = the app never passed the must-pass checks.

The text under each badge, for example "quality 8 of 9", shows how many quality checks passed. The tier belongs to the last commit before the time limit that passes must-pass. It is not a speed rating and not a ranking between setups. One run per task: a Tier 1 next to a Tier 3 can be one unlucky run.

Round 2 · 16 GB class

16 GB · One RTX 5060 Ti box, both cells

Round 2, 16 GB. Same RTX 5060 Ti, same 10 app tasks, same prompts, time limits, agent (opencode 1.18.33) and sampling. Same base model on both sides, Qwen3.8-27B in different quants. @MiaAI_lab's kit first, then my own llama.cpp setup from round 1.

Mia's EXL3 kitllama.cpp UD-Q3_K_XL
Tasks with a counted commit9 of 109 of 10
Tasks at tier 3 (top)6 of 98 of 9
Quality checks passed78 of 8281 of 82
Output tokens in total1,587,4511,242,419

Mia's EXL3 kit · Qwen3.8-27B on RTX 5060 Ti, 16 GB

@MiaAI_lab's one-click kit · recipe on GitHub

EXL3 2.5 bpw · exllamav3 1.4.4 (1.5.3 is out) · 176K context · KV 4 bit · MTP draft · images on · kit at 622c796

Model fit calculator built by Mia's EXL3 kitOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 46.9 min · first working 10.1 min
3D scene built by Mia's EXL3 kitOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 1quality 8 of 9 · counted commit at 139.5 min · first working 11.6 min
Product page built by Mia's EXL3 kitOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 46.9 min · first working 34.9 min
Platformer game built by Mia's EXL3 kitOpen app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 118.2 min · first working 20.8 min
Time tracker built by Mia's EXL3 kitOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 44.1 min · first working 21.1 min
Dashboard built by Mia's EXL3 kitOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 15.8 min · first working 15.8 min
Interactive explainer built by Mia's EXL3 kitOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 36.1 min · first working 8.6 min
Bird flock built by Mia's EXL3 kitOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 1quality 6 of 8 · counted commit at 86.6 min · first working 15 min
no counted commit
9 · Invoice generator
An invoice that is ready to print
No counted commit
No commit passed the must-pass checks
Photo editor built by Mia's EXL3 kitOpen app ↗
10 · Photo editor
A photo editor in the browser
Tier 1quality 8 of 9 · counted commit at 46.7 min · first working 46.7 min

llama.cpp UD-Q3_K_XL · Qwen3.8-27B on RTX 5060 Ti, 16 GB

my setup from round 1

llama.cpp b11151 (b11349 is out) · UD-Q3_K_XL, a 13.1 GB file, ~3.8 bpw · 98,304 context · KV q4_0 · MTP draft · no vision

Model fit calculator built by llama.cpp UD-Q3_K_XLOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 38.1 min · first working 7.3 min
3D scene built by llama.cpp UD-Q3_K_XLOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 139.8 min · first working 38.9 min
Product page built by llama.cpp UD-Q3_K_XLOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 32.5 min · first working 10.2 min
Platformer game built by llama.cpp UD-Q3_K_XLOpen app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 147.1 min · first working 12.9 min
Time tracker built by llama.cpp UD-Q3_K_XLOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 39.3 min · first working 27.4 min
Dashboard built by llama.cpp UD-Q3_K_XLOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 41.7 min · first working 22 min
Interactive explainer built by llama.cpp UD-Q3_K_XLOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 42.3 min · first working 28.4 min
Bird flock built by llama.cpp UD-Q3_K_XLOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 1quality 7 of 8 · counted commit at 138.3 min · first working 88.5 min
no counted commit
9 · Invoice generator
An invoice that is ready to print
No counted commit
No commit passed the must-pass checks
Photo editor built by llama.cpp UD-Q3_K_XLOpen app ↗
10 · Photo editor
A photo editor in the browser
Tier 3quality 9 of 9 · counted commit at 37.5 min · first working 32 min
Footnote. s09 produced no passing commit on either side. At the 45-minute mark both sides were level on tier 3; the 6|8 gap comes from s02 swinging opposite ways. On her kit, s02 had a top-tier version at 16.9 min (9/9). My harness makes every model keep polishing until the end, and only the last passing commit counts. Its last change at 139.5 min, 'realistic' dark monitor screens, killed the night glow. Its own check showed the miss, and it committed anyway. Final: still runs, 8/9, tier 1. On her kit, s08 lost one more quality check after the 45-minute mark (separation, final 6 of 8), again the model's own edits, no sign of outside interference. On her kit, s10 ended at 8 of 9 (keyboard undo), tier 1, from my runner's hard-stop commit. On her kit, s07 is a rescore of the same commit; the first scoring was killed by another agent's pkill on my shared test box. llama.cpp went the other way on s02 (tier 1 at the 45-minute mark, counted tier 3) and missed tier 3 on s08 by a hair: separation 1.29 against the 1.3 bar, first working version only at 88.5 min. One run per task, seed 1.
Fine print.

temperature 0.6 enforced by my proxy for both. Mia's one-click kit is at 622c796: EXL3 2.5 bpw on exllamav3 1.4.4 (1.5.3 is out), 176,128-token window, KV 4 bit, MTP draft, images on. 2.5 bpw is the kit's default for 16 GB (176K with images); the kit also lists 2.0 bpw, 3.0 bpw and 3.5 bpw profiles at other context sizes. It ran with --no-harness, its own harness off. My llama.cpp b11151 (b11349 is out): UD-Q3_K_XL, a 13.1 GB file, about 3.8 bpw, 98,304-token window, KV q4_0, MTP draft, no vision. Engines ran on the 16 GB box's GPU 1, the harness on System A. Both cells shared my test box with other agents; three runs on her kit got brief pkill hits in the agent phase (no counted result changed), and llama.cpp, which ran later, got none.

Mia's EXL3 kit | llama.cpp UD-Q3_K_XL Tasks with a counted commit: 9 | 9 of 10 Tasks at tier 3 (top): 6 | 8 Quality checks passed: 78 of 82 | 81 of 82 Output tokens in total: 1,587,451 | 1,242,419

Definitions. Counted commit = the last commit at the hard stop that passes all must-pass checks (the runner also commits leftover work at the stop; only a commit that passes counts). Tier 3 = must-pass, quality and top checks all pass. Tasks s02, s04 and s08 ran up to 150 minutes, the other seven up to 45.

Round 2 · 32 GB class

32 GB · One RTX 5090, both cells

Round 2, 32 GB. Same RTX 5090, same 10 app tasks, same prompts, time limits, agent and sampling. Same base model on both sides, Qwen3.8-27B in different quants. @MiaAI_lab's recipe first, then @ashxhart's engine.

Mia's NVFP4 recipeTensorFold 0.6.0
Tasks with a counted commit9 of 1010 of 10
Tasks at tier 3 (top)9 of 910
Time to first token, median of 10 task medians20.2 s0.99 s
Output tokens in total1,074,2475,236,055

Mia's NVFP4 recipe · Qwen3.8-27B on RTX 5090, 32 GB

@MiaAI_lab's recipe · recipe on GitHub

RadixArk NVFP4 · vLLM 0.27.1 + her backport of vLLM PR 40914 (vLLM 0.30.0 is out, PR not merged) · 262K · MTP 3 drafting · one sequence at a time · prefix caching off by design · recipe at a5f9bff

Model fit calculator built by Mia's NVFP4 recipeOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 44.8 min · first working 8.8 min
3D scene built by Mia's NVFP4 recipeOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 80.4 min · first working 32.9 min
Product page built by Mia's NVFP4 recipeOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 43 min · first working 10.3 min
Platformer game built by Mia's NVFP4 recipeOpen app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 145.1 min · first working 17.5 min
Time tracker built by Mia's NVFP4 recipeOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 12 min · first working 12 min
Dashboard built by Mia's NVFP4 recipeOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 35.7 min · first working 12.2 min
Interactive explainer built by Mia's NVFP4 recipeOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 40.9 min · first working 16.1 min
Bird flock built by Mia's NVFP4 recipeOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 3quality 8 of 8 · counted commit at 74.7 min · first working 40.2 min
Invoice generator built by Mia's NVFP4 recipeOpen app ↗
9 · Invoice generator
An invoice that is ready to print
Tier 3quality 9 of 9 · counted commit at 41.9 min · first working 9.7 min
no counted commit
10 · Photo editor
A photo editor in the browser
No counted commit
No commit passed the must-pass checks

TensorFold 0.6.0 · Qwen3.8-27B on RTX 5090, 32 GB

@ashxhart's engine · recipe on GitHub

Vontra MLX 4bit + a DFlash2 drafter · 156K window picked by TensorFold · recipe at c464617 (0.6.3 is out)

Model fit calculator built by TensorFold 0.6.0Open app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 33.2 min · first working 5.8 min
3D scene built by TensorFold 0.6.0Open app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 130.4 min · first working 4.5 min
Product page built by TensorFold 0.6.0Open app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 41.1 min · first working 3 min
Platformer game built by TensorFold 0.6.0Open app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 140.2 min · first working 8.9 min
Time tracker built by TensorFold 0.6.0Open app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 38.9 min · first working 9.6 min
Dashboard built by TensorFold 0.6.0Open app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 36.3 min · first working 5.8 min
Interactive explainer built by TensorFold 0.6.0Open app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 28.4 min · first working 7.3 min
Bird flock built by TensorFold 0.6.0Open app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 3quality 8 of 8 · counted commit at 131.4 min · first working 23.3 min
Invoice generator built by TensorFold 0.6.0Open app ↗
9 · Invoice generator
An invoice that is ready to print
Tier 3quality 9 of 9 · counted commit at 38.8 min · first working 6.1 min
Photo editor built by TensorFold 0.6.0Open app ↗
10 · Photo editor
A photo editor in the browser
Tier 3quality 9 of 9 · counted commit at 33.3 min · first working 8 min
Footnote. NVFP4's s10 never produced a passing commit; at the hard stop 10 of the 14 must-pass checks had passed when my checker's 8-minute cap stopped it during the crop check (the rest never ran). Its s05 is a rerun after the engine was unreachable for two minutes (cause not found). TensorFold's s08 had two checks each retried once after two cleanup commands from one other run killed its test server; the retried checks count. My automatic re-check for test-box errors came in after the NVFP4 runs were scored; it would not have retried s10's timeout. Foreign cleanup commands from other runs on my shared test box hit NVFP4's s04, s07 and s10 during the agent phase (s04 and s10 more than once); s07 may have lost its 15-minute mark to one of them, and s10's scoring check may have been hit too, likely without effect. No TensorFold 0.6.0 run was hit during the agent phase. One run per task, seed 1.
Fine print.

temperature 0.6 enforced by my proxy for both. Mia's recipe is at a5f9bff: RadixArk NVFP4 on vLLM 0.27.1 with her backport of vLLM PR 40914 (vLLM 0.30.0 is out; PR 40914 is not merged yet), 262K window, MTP 3 drafting, one sequence at a time, prefix caching off by design so every request is cold. TensorFold 0.6.0 at c464617 (0.6.3 is out): Vontra MLX 4bit plus a DFlash2 drafter, 156K window picked by TensorFold, requests mostly warm; without a cache hit its first token still took a median of 30.0 s (median of the ten task medians). That cache difference is why the first-token line sits where it does. TensorFold wrote 4.9 times the output tokens under the same time limits. More speed is not more quality.

Mia's NVFP4 recipe | TensorFold 0.6.0 Tasks with a counted commit: 9 | 10 of 10 Tasks at tier 3 (top): 9 | 10 of 10 Time to first token, median of 10 task medians: 20.2 s | 0.99 s Output tokens in total: 1074247 | 5236055

Definitions. Counted commit = the last commit at the hard stop that passes all must-pass checks (the runner also commits leftover work at the stop; only a commit that passes counts). Tier 3 = must-pass, quality and top checks all pass. Tasks s02, s04 and s08 ran up to 150 minutes, the other seven up to 45.

Round 2 · 128 GB class

128 GB · One GB10 box per cell

@MiaAI_lab ships two single-box recipes for Qwen3.8-Flash-Next. In my Round 2 each ran on its own GB10 box and got the same 10 app tasks, with the same prompts, time limits, coding agent and sampling.

Mia's TensorFold recipeMia's vLLM kit
Tasks with a counted commit10 of 1010 of 10
Tasks at tier 3 (top)1010
Quality checks passed91 of 9191 of 91

Mia's TensorFold recipe · Qwen3.8-Flash-Next on Gigabyte AI TOP ATOM (GB10), 128 GB

@MiaAI_lab's recipe · recipe on GitHub

MLX 4bit · TensorFold 0.3.6.3 from her image (0.6.3 is out) · PARALLEL=4 (her default 5 did not fit my free memory) · 262K · recipe at a3aa898 (newer commits exist)

Model fit calculator built by Mia's TensorFold recipeOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 36.6 min · first working 5.5 min
3D scene built by Mia's TensorFold recipeOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 126.5 min · first working 19.1 min
Product page built by Mia's TensorFold recipeOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 40.3 min · first working 8.2 min
Platformer game built by Mia's TensorFold recipeOpen app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 138.8 min · first working 17.7 min
Time tracker built by Mia's TensorFold recipeOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 36.3 min · first working 7 min
Dashboard built by Mia's TensorFold recipeOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 27.4 min · first working 13.4 min
Interactive explainer built by Mia's TensorFold recipeOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 18.1 min · first working 4.3 min
Bird flock built by Mia's TensorFold recipeOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 3quality 8 of 8 · counted commit at 135.4 min · first working 17 min
Invoice generator built by Mia's TensorFold recipeOpen app ↗
9 · Invoice generator
An invoice that is ready to print
Tier 3quality 9 of 9 · counted commit at 25.6 min · first working 25.6 min
Photo editor built by Mia's TensorFold recipeOpen app ↗
10 · Photo editor
A photo editor in the browser
Tier 3quality 9 of 9 · counted commit at 36.7 min · first working 33.3 min

Mia's vLLM kit · Qwen3.8-Flash-Next on Lenovo ThinkStation PGX (GB10), 128 GB

@MiaAI_lab's kit · recipe on GitHub

vLLM · NVFP4 · kit at 7d0712d on its default .env · 262K

Model fit calculator built by Mia's vLLM kitOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 36.5 min · first working 8 min
3D scene built by Mia's vLLM kitOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 137.4 min · first working 14 min
Product page built by Mia's vLLM kitOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 24.5 min · first working 8.1 min
Platformer game built by Mia's vLLM kitOpen app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 144.6 min · first working 8.5 min
Time tracker built by Mia's vLLM kitOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 36.2 min · first working 9.4 min
Dashboard built by Mia's vLLM kitOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 35.7 min · first working 8.1 min
Interactive explainer built by Mia's vLLM kitOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 42.8 min · first working 6 min
Bird flock built by Mia's vLLM kitOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 3quality 8 of 8 · counted commit at 141 min · first working 28.7 min
Invoice generator built by Mia's vLLM kitOpen app ↗
9 · Invoice generator
An invoice that is ready to print
Tier 3quality 9 of 9 · counted commit at 47 min · first working 11.8 min
Photo editor built by Mia's vLLM kitOpen app ↗
10 · Photo editor
A photo editor in the browser
Tier 3quality 9 of 9 · counted commit at 35.8 min · first working 13.6 min
Footnote. my first scoring had TensorFold s03 and vLLM s03, s05, s07 at tier 0. On my shared test machine another run's cleanup command killed the test server or browser mid-check. Same commits, same checker, scored again: tier 3. The same flaw hit TensorFold s02 and vLLM s02, s05, s09 while the agents worked.
Fine print.

one run per task, seed 1, temperature 0.6 set by my proxy for both. TensorFold recipe at a3aa898 (newer commits exist) with PARALLEL=4 (her default 5 did not fit my free memory), MLX 4bit, TensorFold 0.3.6.3 from her image (0.6.3 is out). vLLM kit at 7d0712d on its default .env, NVFP4. Both with a 262K window.

Definitions. Counted commit = the last commit at the hard stop that passes all must-pass checks (the runner also commits leftover work at the stop; only a commit that passes counts). Tier 3 = must-pass, quality and top checks all pass. Tasks s02, s04 and s08 ran up to 150 minutes, the other seven up to 45.

Round 2 · 256 GB class

256 GB · Two GB10 pairs

Round 2, 256 GB. What GLM-5.3-Flash builds on two pairs of GB10 boxes (the ASUS Ascent GX10 pair, the Gigabyte AI TOP ATOM + Lenovo ThinkStation PGX pair): the same 10 app tasks, one counted run per task, seed 1. I show @MiaAI_lab's kits first (vLLM and TensorFold), then @ashxhart's TensorFold recipe. Not a duel: the three cells used different engines, checkpoints and context windows; two of them shared the same pair.

Mia's vLLM kitMia's TensorFold kitAsh's TensorFold recipe
Tasks with a counted commit10 of 1010 of 108 of 10
Tasks at tier 3 (top)996 of 8
Quality + top checks passed90 of 9190 of 9171 of 73

Mia's vLLM kit · GLM-5.3-Flash on ASUS Ascent GX10 pair (GB10), 2×128 GB

@MiaAI_lab's kit · recipe on GitHub

vLLM · EXL3 4bpw · tp 2 · 850K window · fp8 KV pool · DFlash2 drafts · vision on · 4 parallel · kit at 674155d (newer commits exist)

Model fit calculator built by Mia's vLLM kitOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 37.2 min · first working 16.6 min
3D scene built by Mia's vLLM kitOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 146.8 min · first working 30.5 min
Product page built by Mia's vLLM kitOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 37.4 min · first working 9.8 min
Platformer game built by Mia's vLLM kitOpen app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 138 min · first working 26.8 min
Time tracker built by Mia's vLLM kitOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 36.6 min · first working 17.6 min
Dashboard built by Mia's vLLM kitOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 40.1 min · first working 12.1 min
Interactive explainer built by Mia's vLLM kitOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 38.3 min · first working 19.1 min
Bird flock built by Mia's vLLM kitOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 3quality 8 of 8 · counted commit at 149.4 min · first working 24.1 min
Invoice generator built by Mia's vLLM kitOpen app ↗
9 · Invoice generator
An invoice that is ready to print
Tier 2quality 8 of 9 · counted commit at 27.6 min · first working 27.6 min
Photo editor built by Mia's vLLM kitOpen app ↗
10 · Photo editor
A photo editor in the browser
Tier 3quality 9 of 9 · counted commit at 37.3 min · first working 19.2 min

Mia's TensorFold kit · GLM-5.3-Flash on Gigabyte AI TOP ATOM + Lenovo ThinkStation PGX pair (GB10), 2×128 GB

@MiaAI_lab's kit · recipe on GitHub

TensorFold 0.6.0 + her 53 patches · EXL3 4bpw · tp 2 · 1M window · fp8 KV · DFlash2 + copy drafts · vision on · 4 parallel · kit at 978b225 (v1.3; v1.5 is out, with a new default checkpoint)

Model fit calculator built by Mia's TensorFold kitOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 35.1 min · first working 5.5 min
3D scene built by Mia's TensorFold kitOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 139.9 min · first working 38.2 min
Product page built by Mia's TensorFold kitOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 38.1 min · first working 4.5 min
Platformer game built by Mia's TensorFold kitOpen app ↗
4 · Platformer game
A 2D jump and run with 3 levels
Tier 3quality 9 of 9 · counted commit at 140.2 min · first working 54.9 min
Time tracker built by Mia's TensorFold kitOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 37.7 min · first working 19.6 min
Dashboard built by Mia's TensorFold kitOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 3quality 9 of 9 · counted commit at 36.4 min · first working 6.2 min
Interactive explainer built by Mia's TensorFold kitOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 36.8 min · first working 24.4 min
Bird flock built by Mia's TensorFold kitOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 1quality 7 of 8 · counted commit at 141.6 min · first working 15.7 min
Invoice generator built by Mia's TensorFold kitOpen app ↗
9 · Invoice generator
An invoice that is ready to print
Tier 3quality 9 of 9 · counted commit at 36.6 min · first working 12.3 min
Photo editor built by Mia's TensorFold kitOpen app ↗
10 · Photo editor
A photo editor in the browser
Tier 3quality 9 of 9 · counted commit at 37.6 min · first working 13.8 min

Ash's TensorFold recipe · GLM-5.3-Flash on ASUS Ascent GX10 pair (GB10), 2×128 GB

@ashxhart's recipe · recipe on GitHub

TensorFold 0.6.0 (0.6.3 is out) · MLX 4bit · tp 2 · 64K window (my pick; the recipe default of 2051 was too small for the agent tasks) · recipe at c464617 (newer commits exist)

Model fit calculator built by Ash's TensorFold recipeOpen app ↗
1 · Model fit calculator
Will this AI model fit on my graphics card?
Tier 3quality 9 of 9 · counted commit at 38.5 min · first working 10.6 min
3D scene built by Ash's TensorFold recipeOpen app ↗
2 · 3D scene
The home AI lab at night, in real-time 3D
Tier 3quality 9 of 9 · counted commit at 150 min · first working 48.9 min
Product page built by Ash's TensorFold recipeOpen app ↗
3 · Product page
A landing page for an invented home AI box
Tier 3quality 10 of 10 · counted commit at 39.9 min · first working 6.9 min
no counted commit
4 · Platformer game
A 2D jump and run with 3 levels
No counted commit
No commit passed the must-pass checks
Time tracker built by Ash's TensorFold recipeOpen app ↗
5 · Time tracker
Time tracking with project billing
Tier 3quality 10 of 10 · counted commit at 42.7 min · first working 42.7 min
Dashboard built by Ash's TensorFold recipeOpen app ↗
6 · Dashboard
Local AI benchmark results as a dashboard
Tier 2quality 8 of 9 · counted commit at 38.6 min · first working 26.6 min
Interactive explainer built by Ash's TensorFold recipeOpen app ↗
7 · Interactive explainer
How a language model writes, one word at a time
Tier 3quality 9 of 9 · counted commit at 35.7 min · first working 33.4 min
Bird flock built by Ash's TensorFold recipeOpen app ↗
8 · Bird flock
A living flocking simulation you can play with
Tier 3quality 8 of 8 · counted commit at 145.2 min · first working 21.2 min
Invoice generator built by Ash's TensorFold recipeOpen app ↗
9 · Invoice generator
An invoice that is ready to print
Tier 2quality 8 of 9 · counted commit at 44.3 min · first working 44.3 min
no counted commit
10 · Photo editor
A photo editor in the browser
No counted commit
No commit passed the must-pass checks
Footnote. Ash's s04 and s10 produced no passing commit (window.APP missing both times; whether my 64K window choice played a part is unchecked); s06 and s09 landed tier 2 (layout stable / 40 random invoices). For Mia's vLLM cell, the counted s04 and s02 runs are reruns: the earlier tries on 01.10. were cut off by my own setup, not model results (s04 three times, s02 once; one was my RAM guard, likely because of my own Jarvis jobs on that box, not proven); the reruns ran in the same engine session as the other eight tasks; s09 landed tier 2 (40 random invoices). Mia's TensorFold s08 is a real narrow regression: tier 3 at 44.2 min, counted tier 1 at 141.6 min, hawk 0.503 against the 0.50 bar. Ash's s02 had processes killed 8 times by my other test agents on the shared harness box (my setup: of the three cells only the Mia cell on the Gigabyte + Lenovo pair ran as its own user), the most of any run I checked (the ten vLLM runs of 02.10. are not part of that check), and still made tier 3.
Fine print.

temperature 0.6 enforced by my proxy for all three. Mia's vLLM cell: her kit at 674155d (newer commits exist), vLLM, EXL3 4bpw, tp 2, 850K window, fp8 KV pool, DFlash2 drafts, vision on, 4 parallel, on the ASUS Ascent GX10 pair; the host side was mine: earlyoom off, oom_score_adj 800, the RAM guard armed under 2 GiB, my Jarvis timers on the first ASUS Ascent GX10 box paused from 02.10. 16:13 to cell end (3 of 10 runs started before), image rebuilt locally with one documented line change. Mia's TF kit at 978b225 (v1.3; v1.5 is out, with a new default checkpoint): TensorFold 0.6.0 plus her 53 patches, EXL3 4bpw, tp 2, 1M window, fp8 KV, DFlash2 + copy drafts, vision on, 4 parallel, on the Gigabyte AI TOP ATOM + Lenovo ThinkStation PGX pair; my harness for this cell ran as its own user (r2iso34); the host side was mine here too: earlyoom off mid-cell, the RAM guard armed under 2 GiB. The Mia TF cell ran on the Gigabyte AI TOP ATOM + Lenovo ThinkStation PGX pair, the weaker pair: a prefill check at the time gave 72-74% of Mia's README prefill (the ASUS Ascent GX10 pair: 98-99%), under Mia's 2200 MHz cap. Cause: the Lenovo box's ConnectX-7 was stuck at ~13 Gb/s after a hot-plug on my side. I fixed it after all ten runs (re-test under the same cap: 100.4-101.7%); no run was repeated. The TF cell itself ran uncapped. Ash's recipe at c464617 (newer commits exist) on TensorFold 0.6.0 (0.6.3 is out; Mia's TF cell runs 0.6.0 too): MLX 4bit, tp 2, 64K window (my pick; the recipe default of 2051 was too small for the agent tasks, its examples use larger ones), on the ASUS Ascent GX10 pair, run on 01.10. with the host as it was: earlyoom on, my Jarvis timers running. GPU clock on the ASUS Ascent GX10 pair cells was not recorded.

Definitions. Counted commit = the last commit at the hard stop that passes all must-pass checks (the runner also commits leftover work at the stop; only a commit that passes counts). Tier 3 = must-pass, quality and top checks all pass. Tasks s02, s04 and s08 ran up to 150 minutes, the other seven up to 45.

Overview

Task × cell, Round 2

Cell: tier (0–3) and quality checks passed. One run per task, seed 1.

Mia TF
128 GB
Mia vLLM
128 GB
Mia EXL3
16 GB
llama.cpp
16 GB
Mia vLLM
256 GB
Mia TF
256 GB
Ash TF
256 GB
Mia NVFP4
32 GB
TF 0.6.0
32 GB
1 · Model fit calculator3 9/93 9/93 9/93 9/93 9/93 9/93 9/93 9/93 9/9
2 · 3D scene3 9/93 9/91 8/93 9/93 9/93 9/93 9/93 9/93 9/9
3 · Product page3 10/103 10/103 10/103 10/103 10/103 10/103 10/103 10/103 10/10
4 · Platformer game3 9/93 9/93 9/93 9/93 9/93 9/9–3 9/93 9/9
5 · Time tracker3 10/103 10/103 10/103 10/103 10/103 10/103 10/103 10/103 10/10
6 · Dashboard3 9/93 9/93 9/93 9/93 9/93 9/92 8/93 9/93 9/9
7 · Interactive explainer3 9/93 9/93 9/93 9/93 9/93 9/93 9/93 9/93 9/9
8 · Bird flock3 8/83 8/81 6/81 7/83 8/81 7/83 8/83 8/83 8/8
9 · Invoice generator3 9/93 9/9––2 8/93 9/92 8/93 9/93 9/9
10 · Photo editor3 9/93 9/91 8/93 9/93 9/93 9/9––3 9/9

Tier 3 = must-pass, quality and top checks all pass · Tier 2 = must-pass and quality, not all top · Tier 1 = must-pass only · – = no counted commit

Round 1 and Round 2 side by side

Same setup, two rounds

Only one setup ran identically in both rounds: llama.cpp b11151, Unsloth UD-Q3_K_XL (Qwen3.8-27B), 98,304-token window, KV q4_0, MTP draft, on the same RTX 5060 Ti 16 GB box. The time limits differ: Round 1 gave 45 minutes for every app; Round 2 gave 45 for seven tasks and 150 for tasks 2, 4 and 8. One run per task in each round — this is not a race, it shows what more time buys.

Round 1 · 45 min · checks passedRound 2 · up to 150 min · tier
1 · Model fit calculator16/16 · first 8 mintier 3 q 9/9 · first 7.3 min
2 · 3D scene (150 min)13/13 · first 12 mintier 3 q 9/9 · first 38.9 min
3 · Product page17/17 · first 10 mintier 3 q 10/10 · first 10.2 min
4 · Platformer game (150 min)18/18 · first 15 mintier 3 q 9/9 · first 12.9 min
5 · Time tracker15/15 · first 19 mintier 3 q 10/10 · first 27.4 min
6 · Dashboard12/12 · first 8 mintier 3 q 9/9 · first 22 min
7 · Interactive explainer15/15 · first 12 mintier 3 q 9/9 · first 28.4 min
8 · Bird flock (150 min)14/14 · first 25 mintier 1 q 7/8 · first 88.5 min
9 · Invoice generator13/13 · first 18 minno counted commit
10 · Photo editor14/14 · first 20 mintier 3 q 9/9 · first 32 min
Documentation

Notes & checks

Late regressions, the cross-run interference audit (pkill), what was re-scored and re-run, and how tiers and the counted commit are defined.

Open notes & checks →