# SWEEP-RUN result, 8 September 2026 Work package SWEEP-RUN (Phase 2 GPU sweep) of the SGH programme (`operations/research/sgh-program-20260908/`). Three spot A100s ran the three `sweep-v1` packages to completion. **407 of 407 planned generations landed: 407 records, 407 PNGs, 0 skipped manifests, 0 failed jobs, 0 missing ids.** Everything below is engineering evidence — throughput, provenance and infrastructure. No image was assessed for quality here and nothing here is clinically validated. Nothing was committed to git. All three slots are TERMINATED; slot d (F2 training) was not touched. **Merged records: `[local]/records.jsonl`** (407 lines, 1.6 MB; every `output` rewritten to `slot/out/` and verified to resolve to a local file). | deliverable | path (under `packages/sweep-v1/`) | |---|---| | merged records | `retrieved/records.jsonl` | | merge script | `retrieved/merge_records.py` | | per-manifest coverage + duplicate log | `retrieved/coverage.json` | | per-slot evidence | `retrieved/evidence-slot{0,1,2}-steps.log`, `-manifest-summary.json` | | per-slot outputs (5.5 GB, 407 PNGs) | `retrieved/slot{0,1,2}/out/` | | token library / stats / organism library / VM cartoon manifests | `retrieved/aux/` | | VM-vs-Mac cartoon hash comparison | `retrieved/aux/cartoon-sha-comparison.json` | | early pipeline check (2 PNGs + metrics) | `early/` | | per-slot orchestrator logs | `../../fleet/logs/sweep-{a,b,c}.out`, `sweep-timeline.tsv` | Slot map: package `slot0` -> fleet slot `a` (`sgh-a100-slot-a`, us-central1-a), `slot1` -> `b` (`sgh-histo-qwen-a100-recovery-b`, us-central1-b), `slot2` -> `c` (`sgh-a100-slot-c`, us-central1-c). ## 1. Timeline All times UTC, 8 September 2026. | what | slot a | slot b | slot c | |---|---|---|---| | `instances start` issued | 14:31:13 | 14:31:46 | 14:32:14 | | ssh + GPU ready | 14:32:18 (65 s) | 14:33:02 (76 s) | 14:33:22 (68 s) | | stage (43 files, ~22 MB) | 14:32:26-14:34:41 (135 s) | 14:33:10-14:35:15 (125 s) | 14:33:31-14:35:28 (117 s) | | worker #1 start | 14:34:58 | 14:35:31 | 14:35:47 | | worker #1 end (**failed: no scipy**) | 14:36:39 ran=1 skipped=16 | 14:37:28 ran=1 skipped=16 | 14:36:03 ran=0 skipped=16 | | scipy 1.14.1 installed | 14:37-14:39 | 14:37-14:39 | 14:37-14:39 | | worker #2 start (serial render) | 14:40:05 | 14:40:05 | 14:40:04 | | parallel setup takes over | 14:42:58 / 14:48:27 / 14:49:53 | same | same | | cartoons complete (a1 60, b1 5, d1 10, e2 20) | ~15:08 | ~15:08 | ~15:08 | | token library (2016 windows) + stats | 15:14:56 / 15:15:01 | 15:14:56 / 15:15:02 | 15:14:57 / 15:15:02 | | **first manifest job (`a1-pass1`)** | **15:16:02** | **15:16:04** | **15:16:03** | | preemption | **15:38:39 (real)** | none | 15:31:08 (**false positive**, see fix 6) | | restarted + resumed | 15:39:19 up, 15:42:36 worker (`records.jsonl 45 -> 45 usable`) | - | not stopped; duplicate worker killed 15:41 | | **last manifest job** | **16:53:42** | **16:46:11** | **16:52:09** | | `WORKER_DONE` (exit 0) | 16:54:32 | 16:47:04 | 16:53:18 | | retrieve (1.8 GB each) | 16:54:32-17:06:26 (714 s) | 16:47:04-16:58:36 (692 s) | 16:53:18-17:04:52 (694 s) | | SHA256SUMS verified | 143 files OK | 131 files OK | 133 files OK | | stopped | 17:07:30 | 16:59:39 | 17:06:05 | Setup (boot to first generation) was **45 minutes**; generation was **90-98 minutes**; retrieval **~12 minutes**. Retrieval ran at ~2.6 MB/s per slot in parallel (5.5 GB total). `fleet.sh status` at the end: ``` SLOT VM ZONE STATUS a sgh-a100-slot-a us-central1-a TERMINATED b sgh-histo-qwen-a100-recovery-b us-central1-b TERMINATED c sgh-a100-slot-c us-central1-c TERMINATED d sgh-a100-slot-d us-central1-f RUNNING <- F2, untouched ``` ## 2. Coverage: expected vs present, per manifest From `retrieved/merge_records.py` (expected = the ids in each slot's own `jobs/.json`; `b1-native` is compared against the copy `worker.sh` rewrote on the VM, pulled back into `retrieved/aux/jobs/`). **Every cell is exp = rec = png. No missing ids, no skipped manifests, no failed jobs.** | manifest | a exp/rec/png | b exp/rec/png | c exp/rec/png | total | |---|---|---|---|---| | a1-pass1 | 19/19/19 | 20/20/20 | 21/21/21 | 60 | | a1-pass2 | 19/19/19 | 20/20/20 | 21/21/21 | 60 | | a1-pass2-si9 | 5/5/5 | 4/4/4 | 6/6/6 | 15 | | a1-pass2-si3 | 19/19/19 | 20/20/20 | 21/21/21 | 60 | | a1-pass3-si15 | 3/3/3 | 3/3/3 | 4/4/4 | 10 | | d1-pass1 | 8/8/8 | 8/8/8 | 9/9/9 | 25 | | d1-pass2 | 8/8/8 | 8/8/8 | 9/9/9 | 25 | | a2-pass1 | 8/8/8 | 6/6/6 | 6/6/6 | 20 | | a2-pass2 | 8/8/8 | 6/6/6 | 6/6/6 | 20 | | a3-pass1 | 6/6/6 | 7/7/7 | 7/7/7 | 20 | | a3-pass2 | 6/6/6 | 7/7/7 | 7/7/7 | 20 | | d2-pass2 | 8/8/8 | 8/8/8 | 4/4/4 | 20 | | e2-pass1 | 5/5/5 | 5/5/5 | 5/5/5 | 15 | | e2-pass2 | 5/5/5 | 5/5/5 | 5/5/5 | 15 | | b1-pass1 | 2/2/2 | 2/2/2 | 1/1/1 | 5 | | b1-pass2 | 2/2/2 | 2/2/2 | 1/1/1 | 5 | | b1-native | 12/12/12 | - (fix 2) | - | 12 | | **slot total** | **143/143/143** | **131/131/131** | **133/133/133** | **407** | `evidence/manifest-summary.json`: slot a `{"ran":17,"skipped":0,"failed":0}`, slots b and c `{"ran":16,"skipped":0,"failed":0}`. 407 vs the split's 405 planned: `b1-native` was written against a placeholder naming convention and `worker.sh` rewrites it from the real `organisms/library/*.png` listing, which holds **12** windows, not 10. All 12 ran (on slot a only — see fix 2). Provenance check on the merged file: **407 of 407 records have `token_source` equal to their `extra.expected_token_source`** — `token_reference` 285, `token_library` 50, `token_file` 40, `token_blend` 20, `reference` 12 (the native 1024 windows). The three token sources GEN-CODE added all worked on real weights; e.g. an A2 record carries `token_library_entries=2016` and 21 per-window picks with `library_file`, `library_category` and `distance`. Outputs by experiment: a1 205, d1 50, a2 40, a3 40, e2 30, d2 20, b1 10, b1-native 12. By category: intestinal_metaplasia 180, hpylori_gastritis 117, normal 55, mixed 55. 395 of the 407 canvases are 4096x2048; the 12 `b1-native` are single 1024 windows. ### The only bookkeeping wrinkle Slot c produced **5 duplicate records** (`a1_p2_hpylori_gastritis_s11_across`, `_s12_across`, `_s12_along`, `_s13_oblique`, `_s15_across`) during the 11 minutes it ran two workers (fix 6). All five duplicates are **byte-identical to the copy that was kept** (same `output_sha256`), so `merge_records.py` keeps the first and logs the drop in `coverage.json` (`duplicate_records_dropped`). 412 lines read, 407 unique written. ## 3. Throughput From the 407 merged records. `denoiser_calls` is the honest cost axis: the canvas is 21 overlapping 1024 windows on one latent, batched 4 at a time. | `start_index` | steps run | denoiser calls | n | median s | mean s | min | max | |---:|---:|---:|---:|---:|---:|---:|---:| | 15 | 5 | 30 | 10 | **25.28** | 25.34 | 25.16 | 25.66 | | 12 | 8 | 48 | 145 | **31.84** | 28.31 | 23.41 | 32.51 | | 9 | 11 | 66 | 15 | **38.71** | 37.09 | 30.16 | 39.09 | | 6 | 14 | 84 | 165 | **45.29** | 43.50 | 36.87 | 95.71 | | 3 | 17 | 102 | 60 | **52.28** | 51.43 | 43.59 | 52.85 | | (native 1024 window, si 12) | 8 | 20 | 12 | **3.02** | 3.12 | 3.01 | 4.24 | That is **0.37 s per denoiser call plus ~7 s of fixed cost** on a 4096x2048 canvas, linear across the whole range. Total generation time over all 407 jobs: **15 216 s = 4.23 GPU-hours**, mean 37.4 s per generation. This is ~30% faster than FLEET's smoke figure of 23.5 s for a *pass* because the sweep's median job is a heavier pass 2. **`start_index` 12 is bimodal** and the split is informative: 81 jobs at a median of 31.9 s are all `token_reference` (first use of that donor field in the process), 64 jobs at a median of 23.6 s are `token_library`, `token_file`, or a `token_reference` whose donor was already in `TokenCache`. The **~8.3 s difference is the UNI2-h encode of one donor field** (21 windows x 16 crops). Anything that reuses a donor across jobs in one process gets that back. The 95.71 s outlier under si6 is from slot c's double-worker window: two generators sharing one A100 each took about twice as long, which is how the fault was spotted. GPU utilisation: 4.23 GPU-hours of generation against 7.61 powered hours = **56%**. The other 44% is the 45-minute setup, the ~12-minute retrieval and the two faults below. ## 4. Fixes applied Six, in the order they were found. Every one is reproduced below with what it broke and how it was verified. ### Fix 1 — the VM venv has no scipy (fatal, and silently so) `code/tissue3d.py:44 from scipy.spatial import cKDTree` -> `ModuleNotFoundError` on all three slots. That kills `render_set.py` (all four cartoon sets) *and* `build_token_library.py` (both libraries), so `worker.sh` correctly skipped 16 of its 17 manifests (`"no job has all of its inputs: reference"`) and **exited 0 with `WORKER_DONE`**. The first run therefore looked like a success and produced 12 PNGs out of 407. A worker that skips everything and reports zero failures is the failure mode to watch for on this rig. The asset-root venv is Python 3.12.3 / numpy 1.26.4 / Pillow 11.1.0 / torch 2.9.1+cu129, with no scipy at all. Fixed with `.venv/bin/pip install --no-deps 'scipy==1.14.1'` on each slot — `--no-deps` and the pin so numpy stays at 1.26.4 (a bare `pip install scipy` can pull numpy 2.x under torch). Verified by importing all 12 package modules on slot a: 10 import cleanly; the two that do not are off the generation path (`pixcell_stain_match` needs scikit-image and is only reached when `STAIN_TARGET` is set, which it is not; `organism_detector` wants `organisms/detector.py`, which package B1 did not ship). The venv lives on the boot disk, so the fix survived the stop/start and the preemption. **Anything that boots these disks needs the same install.** ### Fix 2 — `b1-native` would have been generated twice `worker.sh` rewrites `jobs/b1-native.json` from the real `organisms/library/*.png` listing before running it, which discards the split's per-slot partition of that manifest. Slots a and b therefore each ran **all 12** organism windows, with identical ids. The accident is a free determinism check: the 12 `output_sha256` values were **identical between us-central1-a and us-central1-b**, on top of FLEET's earlier bit-identical b/d result. Slot a's copy was kept; slot b's 12 records and PNGs were removed (backed up on the VM as `evidence/records-b1native-removed.jsonl`), `jobs/b1-native.json` deleted there, and `b1-native` dropped from `packages/sweep-v1/slot1/MANIFEST-ORDER.txt` (the manifest moved to `slot1/jobs-removed/`). **Any manifest `worker.sh` rewrites from an on-VM listing must be given to exactly one slot.** ### Fix 3 — the 20-minute self-stop timer outlives the worker that armed it `worker.sh`'s EXIT trap arms `systemd-run --on-active=20m shutdown`, and it is cancelled only by the *next* worker's own trap, i.e. at the end of that worker's run. So a worker killed or replaced mid-run leaves a live timer that powers the VM off 20 minutes into its successor. (FLEET gotcha 4 fixed the re-arm; this is the other half.) Every relaunch in this package explicitly runs `systemctl stop sgh-fleet-stop.timer; systemctl reset-failed sgh-fleet-stop.service` immediately after `fleet.sh launch`, and this is now built into the driver script. Verified by `systemctl list-timers` showing no armed unit after each relaunch. ### Fix 4 — the serial setup would have idled the A100 for ~100 minutes per slot Measured on the VM (12 vCPU, single-threaded renderer): | step | serial cost per slot | |---|---| | cartoon render, 4096x2048 | 45-60 s each (15 s on this Mac); 95 cartoons ~ 75 min | | token library | 32 s per donor field (21 windows); 96 fields ~ 51 min | `worker.sh` does both **before** it runs any manifest, so ~2 hours of A100 at USD 2.12/h would have been spent with the GPU at 0%. Both jobs are embarrassingly parallel — a cartoon is a pure function of `(category, seed, cut)` and a donor field is encoded independently — so they were re-run as four concurrent chunks each, concurrently with one another: - `code/render_chunk.py` renders a contiguous slice of `render_set.set_a1`'s own job list through `render_set.build` / `render_set.write`, so the bytes are the ones the serial run would have produced; `code/manifest_dir.py` then writes the set's `MANIFEST.sha256.json`. - `build_token_library.py --include-list` on a contiguous quarter of the sorted donor listing, merged by `code/merge_library.py`. Every artefact is built in a scratch directory and moved into place only when complete, so `worker.sh`'s "already built" checks (non-empty `cartoons/`, `token-library/{tokens.pt,index.json}`, `token-stats/token-stats.json`) can never see a half-finished one. Result: the whole setup finished in **45 minutes wall clock instead of ~2 hours**, and the last 25 minutes of it had the GPU busy on the library encode (68-100% utilisation) rather than idle. ### Fix 5 — `merge_library.py` re-based `token_index` on the wrong counter First version did `e["token_index"] = base + e["token_index"]` with `base = len(grids)` — the number of *parts* merged so far (0,1,2,3), not the number of *windows*. Caught by the script's own assertion (`token_index not contiguous`), so nothing corrupt was written, but it failed after the expensive part and left the three GPUs idle from 15:09:47 to 15:15:52 (**~6 minutes x 3 slots**). Fixed to a running window offset, with the per-chunk index also asserted contiguous before use. Re-merge took 5 s per slot and produced **2016 windows = 96 fields x 21**, exactly the expected count, on all three slots; `token_stats.py` then ran in another 5 s. ### Fix 6 — `fleet.sh watch` treats a failed `describe` as a preemption, and relaunches next to a live worker At 15:31:08 slot c logged `WATCH poll 13 status= -> PREEMPTED/STOPPED, recovering`. The status was **empty** — `gcloud instances describe` had failed transiently — not `TERMINATED`. `cmd_up` then correctly reported `UP already RUNNING` and started nothing, but `cmd_watch` went on to `cmd_launch` regardless, so a **second worker started on a VM whose first worker was still running**. Two generators shared one A100 for 11 minutes: per-job elapsed went from 45 s to 93-96 s and five ids were recorded twice. Killed the newer worker's process group (`kill -TERM -`; the worker is `setsid`, so this takes its generator child with it), removed the `WORKER_DONE`/`exit_code.txt` its EXIT trap had written, cancelled the timer it armed, and patched `fleet/fleet.sh`: 1. an empty status is logged as `status-unknown (describe failed), retrying` and does **not** count as a preemption; 2. before any relaunch, `watch` probes `pgrep -f worker.sh` on the VM and refuses to launch a second worker next to a live one. The driver used for the rest of the run applies the same rule locally (launch only when the slot has no live worker), which is how slots b and c were re-attached without disturbing them. ### Determinism caveat found before the run Re-rendering cartoon set b1 on this Mac from the shipped package reproduced 14 of 15 files byte-for-byte; `hpylori_gastritis_s21_oblique_labels.png` differed in 8 bytes at the tail — same 331 856-byte length, same PNG chunk layout, same palette, **pixel arrays identical**, same zlib Adler-32; only the last deflate block and the IDAT CRC differ. PNG output is therefore not bit-reproducible even on one machine, so a hash mismatch is not by itself evidence that pixels differ. That mattered for section 6. ## 5. Preemption and resume **One real preemption, on slot a.** `watch` poll 19 at 15:38:39 saw `status=STOPPING`, `instances start` was issued at 15:38:57, the VM was back at 15:39:19, and the worker was relaunched at 15:42:29. The resume is verified in the worker's own log: ``` 2026-09-08T15:42:36Z resume: records.jsonl 45 lines -> 45 usable 2026-09-08T15:43:48Z a1-pass1 ok (45 records so far) ``` `records.jsonl` was **not** truncated by the power-off (45 of 45 lines parsed), the resumed worker re-resolved `a1-pass1`, skipped all 19 ids already present and added **zero** records for it in 72 s, then continued with `a1-pass2`. No job was regenerated. Slot a's final count (143) equals its expected count exactly. Slot c's 15:31 event was not a preemption at all (fix 6); its worker was never stopped and its `resume: records.jsonl 25 lines -> 25 usable` line comes from the spurious second worker. ## 6. Was the VM rendering byte-identical to the local cartoons? **No.** Comparing `retrieved/aux/MANIFEST.sha256..json` (rendered on slot a) against `operations/research/sgh-program-20260908/cartoons//MANIFEST.sha256.json` (rendered on this Mac by package A1): | set | files compared | byte-identical | differing (cartoon / labels / json) | |---|---:|---:|---| | a1 | 180 | 1 | 60 / 60 / 59 | | b1 | 15 | 0 | 5 / 5 / 5 | | d1 | 30 | 0 | 10 / 10 / 10 | | e2 | 60 | 0 | 20 / 20 / 20 | The difference is **real, not an encoder artefact**, and it is confined to one thing. One cartoon (`a1/intestinal_metaplasia_s11_across`) was pulled back whole and compared pixel by pixel: - the JSON sidecars differ in exactly three keys: `stroma_nuclei` (**VM 4171, Mac 4225**), `label_fractions` and `kind_fractions` (5th decimal). Every geometric key — plane, tilt, depths, tube counts, vessel profiles, subtype, seeds — is identical. - the label map differs in 11.4% of pixels, of which **99.8% are the pair 5 (stroma) <-> 6 (stroma_nucleus)**: 481 580 + 475 707 pixels. Everything else totals ~1.1 k pixels out of 8.4 M. - the RGB cartoon differs in 28.7% of channel values, mean |delta| 13/255, max 214. So the two machines agree on the tissue architecture — glands, lumens, epithelium, goblet mucin, vessels are placed identically — and disagree only on the **stochastic scatter of stromal nuclei**, where a different number of nuclei is drawn and they land in different places. The cause is library skew, not the code: **VM = Python 3.12.3 / numpy 1.26.4 / Pillow 11.1.0 / scipy 1.14.1**; **Mac = Python 3.14.7 / numpy 2.5.2 / Pillow 12.3.0 / scipy 1.18.1**. Practical consequence: the cartoons in `cartoons/` are *not* the exact inputs of this sweep. Every record carries its own `reference_sha256`, so evaluation should resolve cartoons through the record, and any re-render intended to reproduce a specific generation has to pin the numpy/scipy/Pillow versions. The three VMs are consistent with each other (same image, same venv), which is what matters for comparing arms within the sweep. `retrieved/aux/` also holds `token-library/index.json` (2016 entries), `composition-summary.md`, `build-meta.json`, `token-stats/token-stats.json`, `organism-library/index.json` (12 windows, category forced to `hpylori_gastritis`) and the VM-side `jobs/b1-native.json`. ## 7. Early pipeline check As soon as the first a1 pass-2 output existed, `a1/pass2/intestinal_metaplasia/intestinal_metaplasia_s11_across.png` and its pass-1 parent were pulled to `packages/sweep-v1/early/` and measured with `code/cartoon_metrics.py`. The metric tool was first validated against FLEET's smoke PNG and reproduced its recorded `ring_with_lumen_fraction` of 0.8321 exactly. | stage | ring_with_lumen_fraction | ring_count | lumen_tissue_ratio | stroma_fraction | |---|---:|---:|---:|---:| | cartoon (input) | 0.797 | 128 | 0.120 | 0.302 | | pass 1, `start_index` 12 | 0.545 | 132 | 0.088 | 0.319 | | pass 2, `start_index` 6 | **0.202** | 119 | 0.039 | 0.315 | This is **below the 0.7-0.95 band the brief expected**, and the run was allowed to continue. The reasoning, on the evidence: - it is not garbage: 119 detected rings, tissue fraction 1.00, nuclear density 8423/mm2, stroma fraction 0.315 — a structured H&E-like field, not a blank or a wash. - the plumbing was verified rather than assumed. The pass-2 record's `reference_sha256` (`dfe57d6f...`) equals the pass-1 record's `output_sha256` exactly, so `PASS1/` resolved to this run's own pass-1 output; the donor `token_reference` and its sha match between the two passes; `start_index` 12 -> 8 steps / 48 denoiser calls and 6 -> 14 steps / 84 calls; 21 windows, `native_mpp` 0.5, guidance 1.5. - the direction is the phenomenon the sweep exists to measure. Each pass erodes lumen structure on this cartoon (0.797 -> 0.545 -> 0.202); PLAN.md's recorded 0.917 for IM pass 2 was measured on the *older* fullset cartoons, and FLEET's smoke (0.643 -> 0.832) used that older cartoon too. On A1's new `render_set.py` cartoons the light pass is closer to the real held-out IM number (0.710) than the heavy one. That is exactly what `a1-pass2-si9`, `a1-pass2-si3` and `a1-pass3-si15` were added to quantify, and all of those cells are now generated. Stopping the fleet on this reading would have thrown away a verified 45-minute setup to re-measure something the sweep already covers. The number is reported here so Phase-4 evaluation starts from it rather than from the expectation. Note for whoever repeats this: pass-1 and pass-2 outputs share a filename and differ only by directory, so copying both into one directory silently overwrites one. The first attempt here measured the pass-1 file twice; the files in `early/` are now prefixed `a1_pass1_` / `a1_pass2_`. ## 8. Powered time and cost Reconstructed from `fleet/state//powered.log` and cross-checked against the GCE `lastStartTimestamp` / `lastStopTimestamp` of each final session (slot a: `2026-09-08T08:39:17-07:00` / `10:07:28-07:00` = 15:39:17Z / 17:07:28Z, matching the log to 2 s). Sessions before 14:31 belong to the FLEET package and are excluded. | slot | VM | zone | session | up | down | minutes | |---|---|---|---:|---|---|---:| | a | sgh-a100-slot-a | us-central1-a | 1 | 14:31:37 | ~15:39:19 (preempted) | 67.7 | | a | | | 2 | 15:39:19 | 17:07:30 | 88.2 | | a | | | **total** | | | **155.9** (USD 5.51) | | b | sgh-histo-qwen-a100-recovery-b | us-central1-b | 1 | 14:32:09 | 16:59:39 | **147.5** (USD 5.21) | | c | sgh-a100-slot-c | us-central1-c | 1 | 14:32:36 | 17:06:05 | **153.5** (USD 5.42) | | | | | | | **456.9 min = 7.61 h** | **USD 16.14** | At USD 2.12/h for a spot a2-highgpu-1g. Add roughly USD 0.66 of internet egress for the 5.5 GB retrieved. Slot a's first session ends at the moment `cmd_up` found it no longer RUNNING; the preemption itself was seen one poll earlier, so that figure is accurate to about a minute. `fleet.sh status`'s own accumulator under-reports here: it only counts `UP`/`DOWN` pairs, and a preemption writes no `DOWN`, so slot a's first 68 minutes are missing from it. The table above counts them. Where the 7.61 powered hours went, per slot: ~4 min boot + ~2 min stage, 45 min setup (of which ~5 min lost to the scipy failure and ~6 min to the merge bug), 90-98 min generating (4.23 GPU-hours of actual denoising across the fleet), ~12 min retrieving, then stop. ## 9. What this package did not do - No image was assessed for quality. The one ring-topology reading in section 7 is a pipeline sanity check on a single field, not an evaluation; the Pareto selection, copy screens and reviewer pack are Phase 4. - The `start_index` finding in section 7 is one cartoon. The 407 generations needed to make it a measurement exist and are merged, but they have not been measured. - No pathologist review, no clinical claim, and nothing here is "solved". - The VM-vs-Mac cartoon difference was characterised on one field of one set. The hash comparison covers all 285 files; the pixel-level attribution to stromal nuclei does not. - Nothing was committed to git.