Notes & checks
Below: first a compiled summary of the definitions, then the two public audit posts kept verbatim (hostnames replaced by public device names), then a short note on what was corrected. ← Round 2
Definitions
Counted commit = the last commit at the hard stop that passes all must-pass checks (the runner also commits leftover work at the stop; only a commit that passes counts). Tasks s02, s04 and s08 ran up to 150 minutes, the other seven up to 45.
Tier 3 = must-pass, quality and top checks all pass. Tier 2 = must-pass and quality pass, not all top. Tier 1 = must-pass checks only. x/y = quality and top checks passed on the counted commit.
First working = minute of the first commit that passed the must-pass checks. Checkpoints = the same rule applied at fixed minutes (15 / 45, and 150 for the big tasks); the final checkpoint is always the counted state.
Late regressions (two runs got worse after the 45-minute mark)
In Round 2, two runs went from tier 3 at the 45-minute mark to tier 1 at the end, both on @MiaAI_lab's setups (2 of her 12 150-minute runs that were tier 3 at 45 min; 0 of 5 on the other kits). No foreign kill reached either run, and both apps still ran fine at the end. One run per task, seed 1.
the RTX 5060 Ti box, 16 GB, Mia's EXL3 kit, task s02 (150 min): at 16.9 min my checker rated it top tier, 9/9. My harness makes every model keep polishing until the end, and only the last passing commit counts. Its last commit, at 139.5 min, 'realistic' dark monitor screens, killed the night glow (0.135% against the 0.3% bar). Its own check showed the miss (0.13%). It meant to fix it, but its replies turned into one repeated harness line, and my harness's gate told it to commit the working state. It did. Counted: still runs, 8/9, tier 1.
the Gigabyte AI TOP ATOM + Lenovo ThinkStation PGX pair, 256 GB, Mia's GLM kit v1.3 (978b225, newer versions exist), task s08: tier 3 at 44.2 min (8/8), counted tier 1 at 141.6 min (7/8). The miss: hawk (the flock has to scatter from the mouse) at 0.503 against the 0.50 bar, 76 of 151 birds still near the pointer, one too many. I did not trace which commit moved it and did not re-measure the same commit, so it can be noise.
Footnote: no foreign kill reached either run (the s08 cell ran under its own user, and the s02 audit found none in its window). For s02 the miss traces to the last commit; for s08 the cause was not traced. Same task s02 went the other way on my llama.cpp setup (tier 1 at the 45-minute mark, counted tier 3). One run each, seed 1, a repeat could land differently.
Counted commit = last commit before the time limit that passes all must-pass checks.
Tier 1 = must-pass checks only. Tier 3 = must-pass, quality and top checks all pass. x/y = quality and top checks passed.
Cross-run interference: the pkill audit
Round 2 (16, 32, 128 and 256 GB classes), audited up to 02.10., 10:50. Almost every benchmark agent ran as the same Linux user. Did they disturb each other? I went through the logs in two passes: all 3,359 tool calls of the runs up to 01.10., 14:15, then the cells that were still running, up to 02.10., 10:50. Mia's GLM cell on the ASUS Ascent GX10 pair (10 runs, finished 03.10. 01:20) ran after the cutoff and is not covered.
Short answer: yes. As far as the logs show, no final counted result got worse because of it. One run (NVFP4 s07) may have lost its 15-minute mark to a foreign kill.
Broad kill commands like pkill -f "vite preview" (and a few port-based ones) fired 91 times by 01.10., 14:15 (95 kill segments, because some calls chain several kills). They hit 12 runs during agent work plus 5 scoring runs that first read tier 0 (two runs appear in both lists). All 5 rescored to tier 3 on the same commits.
Most of those commands hit the caller itself: the command text matches the pattern, so pkill killed the caller's own shell in 72 of the 91 calls, and 61 of those hung until the tool timeout. The hang only costs the caller time. These same commands also killed other runs' servers; that is how most of the hits above happened. In one case a foreign cleanup command (ps | grep | xargs kill -9, not a pkill) killed another run's bench harness mid-run. The harness resumed, and the counted result stayed the same.
Then my scoring got an automatic re-check for test-box errors: up to 2 retries when a check aborts for infrastructure reasons (active from 01.10., 13:46; not in every older runner). In the follow-up audit (42 runs from cells that were still running or missing; 37 with a counted commit) the retry fired in exactly one run: two foreign pkills from another run each killed one of its scoring rechecks, and both retries passed fully. That audit found 11 more broad kills; 10 of them hung the caller's own shell. As far as the logs show, no counted result got worse because of them. The one cell that already ran as its own user answered a foreign pkill with "Operation not permitted".
Footnote: the first audit's line "no broad kill after 11:23" held only at the time it was written; the follow-up found more. Two late regressions in these runs have their own post; no foreign kill reached either run. The audit recommends, for a later round, one Linux user per cell, its own tmp dir per run and a fixed port range per run.
Tier 3 = must-pass, quality and top checks all pass.
What was corrected
Re-score: five early scoring runs read tier 0 because another run's cleanup command killed their test server or browser mid-check. Same commits, same checker, scored again: all five are tier 3 and are shown as tier 3 on this page. Since 01.10. the scorer retries a check up to twice when it aborts for infrastructure reasons.
Re-runs: in the 256 GB class the counted s04 and s02 runs of Mia's vLLM cell are re-runs — the earlier tries on 01.10. were cut off by my own setup, not model results. The re-runs ran in the same engine session as the other eight tasks. In the 32 GB class, the counted NVFP4 s05 is a re-run after the engine was unreachable for two minutes (cause not found). Failed first tries are not model results and are not on the page.
Isolation: Mia's TensorFold cell on the Gigabyte + Lenovo pair ran as its own Linux user; a foreign pkill against it answered "Operation not permitted". For a later round the audit recommends one Linux user per cell, its own tmp dir per run and a fixed port range per run.