SWEEP-RUN result, 8 September 2026
Work package SWEEP-RUN (Phase 2 GPU sweep) of the SGH programme
(operations/research/sgh-program-20260908/). Three spot A100s ran the three sweep-v1 packages
to completion. 407 of 407 planned generations landed: 407 records, 407 PNGs, 0 skipped
manifests, 0 failed jobs, 0 missing ids. Everything below is engineering evidence — throughput,
provenance and infrastructure. No image was assessed for quality here and nothing here is
clinically validated. Nothing was committed to git. All three slots are TERMINATED; slot d (F2
training) was not touched.
Merged records: [local]/records.jsonl (407 lines, 1.6 MB;
every output rewritten to slot<i>/out/<path> and verified to resolve to a local file).
| deliverable | path (under packages/sweep-v1/) |
|---|---|
| merged records | retrieved/records.jsonl |
| merge script | retrieved/merge_records.py |
| per-manifest coverage + duplicate log | retrieved/coverage.json |
| per-slot evidence | retrieved/evidence-slot{0,1,2}-steps.log, -manifest-summary.json |
| per-slot outputs (5.5 GB, 407 PNGs) | retrieved/slot{0,1,2}/out/ |
| token library / stats / organism library / VM cartoon manifests | retrieved/aux/ |
| VM-vs-Mac cartoon hash comparison | retrieved/aux/cartoon-sha-comparison.json |
| early pipeline check (2 PNGs + metrics) | early/ |
| per-slot orchestrator logs | ../../fleet/logs/sweep-{a,b,c}.out, sweep-timeline.tsv |
Slot map: package slot0 -> fleet slot a (sgh-a100-slot-a, us-central1-a), slot1 -> b
(sgh-histo-qwen-a100-recovery-b, us-central1-b), slot2 -> c (sgh-a100-slot-c, us-central1-c).
1. Timeline
All times UTC, 8 September 2026.
| what | slot a | slot b | slot c |
|---|---|---|---|
instances start issued |
14:31:13 | 14:31:46 | 14:32:14 |
| ssh + GPU ready | 14:32:18 (65 s) | 14:33:02 (76 s) | 14:33:22 (68 s) |
| stage (43 files, ~22 MB) | 14:32:26-14:34:41 (135 s) | 14:33:10-14:35:15 (125 s) | 14:33:31-14:35:28 (117 s) |
| worker #1 start | 14:34:58 | 14:35:31 | 14:35:47 |
| worker #1 end (failed: no scipy) | 14:36:39 ran=1 skipped=16 | 14:37:28 ran=1 skipped=16 | 14:36:03 ran=0 skipped=16 |
| scipy 1.14.1 installed | 14:37-14:39 | 14:37-14:39 | 14:37-14:39 |
| worker #2 start (serial render) | 14:40:05 | 14:40:05 | 14:40:04 |
| parallel setup takes over | 14:42:58 / 14:48:27 / 14:49:53 | same | same |
| cartoons complete (a1 60, b1 5, d1 10, e2 20) | ~15:08 | ~15:08 | ~15:08 |
| token library (2016 windows) + stats | 15:14:56 / 15:15:01 | 15:14:56 / 15:15:02 | 15:14:57 / 15:15:02 |
first manifest job (a1-pass1) |
15:16:02 | 15:16:04 | 15:16:03 |
| preemption | 15:38:39 (real) | none | 15:31:08 (false positive, see fix 6) |
| restarted + resumed | 15:39:19 up, 15:42:36 worker (records.jsonl 45 -> 45 usable) |
- | not stopped; duplicate worker killed 15:41 |
| last manifest job | 16:53:42 | 16:46:11 | 16:52:09 |
WORKER_DONE (exit 0) |
16:54:32 | 16:47:04 | 16:53:18 |
| retrieve (1.8 GB each) | 16:54:32-17:06:26 (714 s) | 16:47:04-16:58:36 (692 s) | 16:53:18-17:04:52 (694 s) |
| SHA256SUMS verified | 143 files OK | 131 files OK | 133 files OK |
| stopped | 17:07:30 | 16:59:39 | 17:06:05 |
Setup (boot to first generation) was 45 minutes; generation was 90-98 minutes; retrieval ~12 minutes. Retrieval ran at ~2.6 MB/s per slot in parallel (5.5 GB total).
fleet.sh status at the end:
SLOT VM ZONE STATUS
a sgh-a100-slot-a us-central1-a TERMINATED
b sgh-histo-qwen-a100-recovery-b us-central1-b TERMINATED
c sgh-a100-slot-c us-central1-c TERMINATED
d sgh-a100-slot-d us-central1-f RUNNING <- F2, untouched
2. Coverage: expected vs present, per manifest
From retrieved/merge_records.py (expected = the ids in each slot's own jobs/<manifest>.json;
b1-native is compared against the copy worker.sh rewrote on the VM, pulled back into
retrieved/aux/jobs/). Every cell is exp = rec = png. No missing ids, no skipped manifests, no
failed jobs.
| manifest | a exp/rec/png | b exp/rec/png | c exp/rec/png | total |
|---|---|---|---|---|
| a1-pass1 | 19/19/19 | 20/20/20 | 21/21/21 | 60 |
| a1-pass2 | 19/19/19 | 20/20/20 | 21/21/21 | 60 |
| a1-pass2-si9 | 5/5/5 | 4/4/4 | 6/6/6 | 15 |
| a1-pass2-si3 | 19/19/19 | 20/20/20 | 21/21/21 | 60 |
| a1-pass3-si15 | 3/3/3 | 3/3/3 | 4/4/4 | 10 |
| d1-pass1 | 8/8/8 | 8/8/8 | 9/9/9 | 25 |
| d1-pass2 | 8/8/8 | 8/8/8 | 9/9/9 | 25 |
| a2-pass1 | 8/8/8 | 6/6/6 | 6/6/6 | 20 |
| a2-pass2 | 8/8/8 | 6/6/6 | 6/6/6 | 20 |
| a3-pass1 | 6/6/6 | 7/7/7 | 7/7/7 | 20 |
| a3-pass2 | 6/6/6 | 7/7/7 | 7/7/7 | 20 |
| d2-pass2 | 8/8/8 | 8/8/8 | 4/4/4 | 20 |
| e2-pass1 | 5/5/5 | 5/5/5 | 5/5/5 | 15 |
| e2-pass2 | 5/5/5 | 5/5/5 | 5/5/5 | 15 |
| b1-pass1 | 2/2/2 | 2/2/2 | 1/1/1 | 5 |
| b1-pass2 | 2/2/2 | 2/2/2 | 1/1/1 | 5 |
| b1-native | 12/12/12 | - (fix 2) | - | 12 |
| slot total | 143/143/143 | 131/131/131 | 133/133/133 | 407 |
evidence/manifest-summary.json: slot a {"ran":17,"skipped":0,"failed":0}, slots b and c
{"ran":16,"skipped":0,"failed":0}.
407 vs the split's 405 planned: b1-native was written against a placeholder naming convention
and worker.sh rewrites it from the real organisms/library/*.png listing, which holds 12
windows, not 10. All 12 ran (on slot a only — see fix 2).
Provenance check on the merged file: 407 of 407 records have token_source equal to their
extra.expected_token_source — token_reference 285, token_library 50, token_file 40,
token_blend 20, reference 12 (the native 1024 windows). The three token sources GEN-CODE added
all worked on real weights; e.g. an A2 record carries token_library_entries=2016 and 21
per-window picks with library_file, library_category and distance.
Outputs by experiment: a1 205, d1 50, a2 40, a3 40, e2 30, d2 20, b1 10, b1-native 12.
By category: intestinal_metaplasia 180, hpylori_gastritis 117, normal 55, mixed 55.
395 of the 407 canvases are 4096x2048; the 12 b1-native are single 1024 windows.
The only bookkeeping wrinkle
Slot c produced 5 duplicate records (a1_p2_hpylori_gastritis_s11_across, _s12_across,
_s12_along, _s13_oblique, _s15_across) during the 11 minutes it ran two workers (fix 6).
All five duplicates are byte-identical to the copy that was kept (same output_sha256), so
merge_records.py keeps the first and logs the drop in coverage.json
(duplicate_records_dropped). 412 lines read, 407 unique written.
3. Throughput
From the 407 merged records. denoiser_calls is the honest cost axis: the canvas is 21
overlapping 1024 windows on one latent, batched 4 at a time.
start_index |
steps run | denoiser calls | n | median s | mean s | min | max |
|---|---|---|---|---|---|---|---|
| 15 | 5 | 30 | 10 | 25.28 | 25.34 | 25.16 | 25.66 |
| 12 | 8 | 48 | 145 | 31.84 | 28.31 | 23.41 | 32.51 |
| 9 | 11 | 66 | 15 | 38.71 | 37.09 | 30.16 | 39.09 |
| 6 | 14 | 84 | 165 | 45.29 | 43.50 | 36.87 | 95.71 |
| 3 | 17 | 102 | 60 | 52.28 | 51.43 | 43.59 | 52.85 |
| (native 1024 window, si 12) | 8 | 20 | 12 | 3.02 | 3.12 | 3.01 | 4.24 |
That is 0.37 s per denoiser call plus ~7 s of fixed cost on a 4096x2048 canvas, linear across the whole range. Total generation time over all 407 jobs: 15 216 s = 4.23 GPU-hours, mean 37.4 s per generation. This is ~30% faster than FLEET's smoke figure of 23.5 s for a pass because the sweep's median job is a heavier pass 2.
start_index 12 is bimodal and the split is informative: 81 jobs at a median of 31.9 s are all
token_reference (first use of that donor field in the process), 64 jobs at a median of 23.6 s are
token_library, token_file, or a token_reference whose donor was already in TokenCache. The
~8.3 s difference is the UNI2-h encode of one donor field (21 windows x 16 crops). Anything that
reuses a donor across jobs in one process gets that back.
The 95.71 s outlier under si6 is from slot c's double-worker window: two generators sharing one A100 each took about twice as long, which is how the fault was spotted.
GPU utilisation: 4.23 GPU-hours of generation against 7.61 powered hours = 56%. The other 44% is the 45-minute setup, the ~12-minute retrieval and the two faults below.
4. Fixes applied
Six, in the order they were found. Every one is reproduced below with what it broke and how it was verified.
Fix 1 — the VM venv has no scipy (fatal, and silently so)
code/tissue3d.py:44 from scipy.spatial import cKDTree -> ModuleNotFoundError on all three
slots. That kills render_set.py (all four cartoon sets) and build_token_library.py (both
libraries), so worker.sh correctly skipped 16 of its 17 manifests
("no job has all of its inputs: reference") and exited 0 with WORKER_DONE. The first run
therefore looked like a success and produced 12 PNGs out of 407. A worker that skips everything and
reports zero failures is the failure mode to watch for on this rig.
The asset-root venv is Python 3.12.3 / numpy 1.26.4 / Pillow 11.1.0 / torch 2.9.1+cu129, with no
scipy at all. Fixed with .venv/bin/pip install --no-deps 'scipy==1.14.1' on each slot — --no-deps
and the pin so numpy stays at 1.26.4 (a bare pip install scipy can pull numpy 2.x under torch).
Verified by importing all 12 package modules on slot a: 10 import cleanly; the two that do not are
off the generation path (pixcell_stain_match needs scikit-image and is only reached when
STAIN_TARGET is set, which it is not; organism_detector wants organisms/detector.py, which
package B1 did not ship). The venv lives on the boot disk, so the fix survived the stop/start and
the preemption. Anything that boots these disks needs the same install.
Fix 2 — b1-native would have been generated twice
worker.sh rewrites jobs/b1-native.json from the real organisms/library/*.png listing before
running it, which discards the split's per-slot partition of that manifest. Slots a and b therefore
each ran all 12 organism windows, with identical ids.
The accident is a free determinism check: the 12 output_sha256 values were identical between
us-central1-a and us-central1-b, on top of FLEET's earlier bit-identical b/d result. Slot a's copy
was kept; slot b's 12 records and PNGs were removed (backed up on the VM as
evidence/records-b1native-removed.jsonl), jobs/b1-native.json deleted there, and b1-native
dropped from packages/sweep-v1/slot1/MANIFEST-ORDER.txt (the manifest moved to
slot1/jobs-removed/). Any manifest worker.sh rewrites from an on-VM listing must be given to
exactly one slot.
Fix 3 — the 20-minute self-stop timer outlives the worker that armed it
worker.sh's EXIT trap arms systemd-run --on-active=20m shutdown, and it is cancelled only by the
next worker's own trap, i.e. at the end of that worker's run. So a worker killed or replaced
mid-run leaves a live timer that powers the VM off 20 minutes into its successor. (FLEET gotcha 4
fixed the re-arm; this is the other half.) Every relaunch in this package explicitly runs
systemctl stop sgh-fleet-stop.timer; systemctl reset-failed sgh-fleet-stop.service immediately
after fleet.sh launch, and this is now built into the driver script. Verified by
systemctl list-timers showing no armed unit after each relaunch.
Fix 4 — the serial setup would have idled the A100 for ~100 minutes per slot
Measured on the VM (12 vCPU, single-threaded renderer):
| step | serial cost per slot |
|---|---|
| cartoon render, 4096x2048 | 45-60 s each (15 s on this Mac); 95 cartoons ~ 75 min |
| token library | 32 s per donor field (21 windows); 96 fields ~ 51 min |
worker.sh does both before it runs any manifest, so ~2 hours of A100 at USD 2.12/h would have
been spent with the GPU at 0%. Both jobs are embarrassingly parallel — a cartoon is a pure function
of (category, seed, cut) and a donor field is encoded independently — so they were re-run as four
concurrent chunks each, concurrently with one another:
code/render_chunk.pyrenders a contiguous slice ofrender_set.set_a1's own job list throughrender_set.build/render_set.write, so the bytes are the ones the serial run would have produced;code/manifest_dir.pythen writes the set'sMANIFEST.sha256.json.build_token_library.py --include-liston a contiguous quarter of the sorted donor listing, merged bycode/merge_library.py.
Every artefact is built in a scratch directory and moved into place only when complete, so
worker.sh's "already built" checks (non-empty cartoons/<set>, token-library/{tokens.pt,index.json},
token-stats/token-stats.json) can never see a half-finished one. Result: the whole setup finished
in 45 minutes wall clock instead of ~2 hours, and the last 25 minutes of it had the GPU busy on
the library encode (68-100% utilisation) rather than idle.
Fix 5 — merge_library.py re-based token_index on the wrong counter
First version did e["token_index"] = base + e["token_index"] with base = len(grids) — the number
of parts merged so far (0,1,2,3), not the number of windows. Caught by the script's own
assertion (token_index not contiguous), so nothing corrupt was written, but it failed after the
expensive part and left the three GPUs idle from 15:09:47 to 15:15:52 (~6 minutes x 3 slots).
Fixed to a running window offset, with the per-chunk index also asserted contiguous before use.
Re-merge took 5 s per slot and produced 2016 windows = 96 fields x 21, exactly the expected
count, on all three slots; token_stats.py then ran in another 5 s.
Fix 6 — fleet.sh watch treats a failed describe as a preemption, and relaunches next to a live worker
At 15:31:08 slot c logged WATCH poll 13 status= -> PREEMPTED/STOPPED, recovering. The status was
empty — gcloud instances describe had failed transiently — not TERMINATED. cmd_up then
correctly reported UP already RUNNING and started nothing, but cmd_watch went on to
cmd_launch regardless, so a second worker started on a VM whose first worker was still
running. Two generators shared one A100 for 11 minutes: per-job elapsed went from 45 s to 93-96 s
and five ids were recorded twice.
Killed the newer worker's process group (kill -TERM -<pgid>; the worker is setsid, so this takes
its generator child with it), removed the WORKER_DONE/exit_code.txt its EXIT trap had written,
cancelled the timer it armed, and patched fleet/fleet.sh:
- an empty status is logged as
status-unknown (describe failed), retryingand does not count as a preemption; - before any relaunch,
watchprobespgrep -f worker.shon the VM and refuses to launch a second worker next to a live one.
The driver used for the rest of the run applies the same rule locally (launch only when the slot has no live worker), which is how slots b and c were re-attached without disturbing them.
Determinism caveat found before the run
Re-rendering cartoon set b1 on this Mac from the shipped package reproduced 14 of 15 files
byte-for-byte; hpylori_gastritis_s21_oblique_labels.png differed in 8 bytes at the tail — same
331 856-byte length, same PNG chunk layout, same palette, pixel arrays identical, same zlib
Adler-32; only the last deflate block and the IDAT CRC differ. PNG output is therefore not
bit-reproducible even on one machine, so a hash mismatch is not by itself evidence that pixels
differ. That mattered for section 6.
5. Preemption and resume
One real preemption, on slot a. watch poll 19 at 15:38:39 saw status=STOPPING,
instances start was issued at 15:38:57, the VM was back at 15:39:19, and the worker was relaunched
at 15:42:29. The resume is verified in the worker's own log:
2026-09-08T15:42:36Z resume: records.jsonl 45 lines -> 45 usable
2026-09-08T15:43:48Z a1-pass1 ok (45 records so far)
records.jsonl was not truncated by the power-off (45 of 45 lines parsed), the resumed worker
re-resolved a1-pass1, skipped all 19 ids already present and added zero records for it in 72 s,
then continued with a1-pass2. No job was regenerated. Slot a's final count (143) equals its
expected count exactly.
Slot c's 15:31 event was not a preemption at all (fix 6); its worker was never stopped and its
resume: records.jsonl 25 lines -> 25 usable line comes from the spurious second worker.
6. Was the VM rendering byte-identical to the local cartoons?
No. Comparing retrieved/aux/MANIFEST.sha256.<set>.json (rendered on slot a) against
operations/research/sgh-program-20260908/cartoons/<set>/MANIFEST.sha256.json (rendered on this Mac
by package A1):
| set | files compared | byte-identical | differing (cartoon / labels / json) |
|---|---|---|---|
| a1 | 180 | 1 | 60 / 60 / 59 |
| b1 | 15 | 0 | 5 / 5 / 5 |
| d1 | 30 | 0 | 10 / 10 / 10 |
| e2 | 60 | 0 | 20 / 20 / 20 |
The difference is real, not an encoder artefact, and it is confined to one thing. One cartoon
(a1/intestinal_metaplasia_s11_across) was pulled back whole and compared pixel by pixel:
- the JSON sidecars differ in exactly three keys:
stroma_nuclei(VM 4171, Mac 4225),label_fractionsandkind_fractions(5th decimal). Every geometric key — plane, tilt, depths, tube counts, vessel profiles, subtype, seeds — is identical. - the label map differs in 11.4% of pixels, of which 99.8% are the pair 5 (stroma) <-> 6 (stroma_nucleus): 481 580 + 475 707 pixels. Everything else totals ~1.1 k pixels out of 8.4 M.
- the RGB cartoon differs in 28.7% of channel values, mean |delta| 13/255, max 214.
So the two machines agree on the tissue architecture — glands, lumens, epithelium, goblet mucin, vessels are placed identically — and disagree only on the stochastic scatter of stromal nuclei, where a different number of nuclei is drawn and they land in different places. The cause is library skew, not the code: VM = Python 3.12.3 / numpy 1.26.4 / Pillow 11.1.0 / scipy 1.14.1; Mac = Python 3.14.7 / numpy 2.5.2 / Pillow 12.3.0 / scipy 1.18.1.
Practical consequence: the cartoons in cartoons/ are not the exact inputs of this sweep. Every
record carries its own reference_sha256, so evaluation should resolve cartoons through the record,
and any re-render intended to reproduce a specific generation has to pin the numpy/scipy/Pillow
versions. The three VMs are consistent with each other (same image, same venv), which is what
matters for comparing arms within the sweep.
retrieved/aux/ also holds token-library/index.json (2016 entries), composition-summary.md,
build-meta.json, token-stats/token-stats.json, organism-library/index.json (12 windows,
category forced to hpylori_gastritis) and the VM-side jobs/b1-native.json.
7. Early pipeline check
As soon as the first a1 pass-2 output existed, a1/pass2/intestinal_metaplasia/intestinal_metaplasia_s11_across.png
and its pass-1 parent were pulled to packages/sweep-v1/early/ and measured with
code/cartoon_metrics.py. The metric tool was first validated against FLEET's smoke PNG and
reproduced its recorded ring_with_lumen_fraction of 0.8321 exactly.
| stage | ring_with_lumen_fraction | ring_count | lumen_tissue_ratio | stroma_fraction |
|---|---|---|---|---|
| cartoon (input) | 0.797 | 128 | 0.120 | 0.302 |
pass 1, start_index 12 |
0.545 | 132 | 0.088 | 0.319 |
pass 2, start_index 6 |
0.202 | 119 | 0.039 | 0.315 |
This is below the 0.7-0.95 band the brief expected, and the run was allowed to continue. The reasoning, on the evidence:
- it is not garbage: 119 detected rings, tissue fraction 1.00, nuclear density 8423/mm2, stroma fraction 0.315 — a structured H&E-like field, not a blank or a wash.
- the plumbing was verified rather than assumed. The pass-2 record's
reference_sha256(dfe57d6f...) equals the pass-1 record'soutput_sha256exactly, soPASS1/resolved to this run's own pass-1 output; the donortoken_referenceand its sha match between the two passes;start_index12 -> 8 steps / 48 denoiser calls and 6 -> 14 steps / 84 calls; 21 windows,native_mpp0.5, guidance 1.5. - the direction is the phenomenon the sweep exists to measure. Each pass erodes lumen structure on
this cartoon (0.797 -> 0.545 -> 0.202); PLAN.md's recorded 0.917 for IM pass 2 was measured on the
older fullset cartoons, and FLEET's smoke (0.643 -> 0.832) used that older cartoon too. On A1's
new
render_set.pycartoons the light pass is closer to the real held-out IM number (0.710) than the heavy one. That is exactly whata1-pass2-si9,a1-pass2-si3anda1-pass3-si15were added to quantify, and all of those cells are now generated.
Stopping the fleet on this reading would have thrown away a verified 45-minute setup to re-measure something the sweep already covers. The number is reported here so Phase-4 evaluation starts from it rather than from the expectation.
Note for whoever repeats this: pass-1 and pass-2 outputs share a filename and differ only by
directory, so copying both into one directory silently overwrites one. The first attempt here
measured the pass-1 file twice; the files in early/ are now prefixed a1_pass1_ / a1_pass2_.
8. Powered time and cost
Reconstructed from fleet/state/<slot>/powered.log and cross-checked against the GCE
lastStartTimestamp / lastStopTimestamp of each final session (slot a:
2026-09-08T08:39:17-07:00 / 10:07:28-07:00 = 15:39:17Z / 17:07:28Z, matching the log to 2 s).
Sessions before 14:31 belong to the FLEET package and are excluded.
| slot | VM | zone | session | up | down | minutes |
|---|---|---|---|---|---|---|
| a | sgh-a100-slot-a | us-central1-a | 1 | 14:31:37 | ~15:39:19 (preempted) | 67.7 |
| a | 2 | 15:39:19 | 17:07:30 | 88.2 | ||
| a | total | 155.9 (USD 5.51) | ||||
| b | sgh-histo-qwen-a100-recovery-b | us-central1-b | 1 | 14:32:09 | 16:59:39 | 147.5 (USD 5.21) |
| c | sgh-a100-slot-c | us-central1-c | 1 | 14:32:36 | 17:06:05 | 153.5 (USD 5.42) |
| 456.9 min = 7.61 h | USD 16.14 |
At USD 2.12/h for a spot a2-highgpu-1g. Add roughly USD 0.66 of internet egress for the 5.5 GB
retrieved. Slot a's first session ends at the moment cmd_up found it no longer RUNNING; the
preemption itself was seen one poll earlier, so that figure is accurate to about a minute.
fleet.sh status's own accumulator under-reports here: it only counts UP/DOWN pairs, and a
preemption writes no DOWN, so slot a's first 68 minutes are missing from it. The table above
counts them.
Where the 7.61 powered hours went, per slot: ~4 min boot + ~2 min stage, 45 min setup (of which ~5 min lost to the scipy failure and ~6 min to the merge bug), 90-98 min generating (4.23 GPU-hours of actual denoising across the fleet), ~12 min retrieving, then stop.
9. What this package did not do
- No image was assessed for quality. The one ring-topology reading in section 7 is a pipeline sanity check on a single field, not an evaluation; the Pareto selection, copy screens and reviewer pack are Phase 4.
- The
start_indexfinding in section 7 is one cartoon. The 407 generations needed to make it a measurement exist and are merged, but they have not been measured. - No pathologist review, no clinical claim, and nothing here is "solved".
- The VM-vs-Mac cartoon difference was characterised on one field of one set. The hash comparison covers all 285 files; the pixel-level attribution to stromal nuclei does not.
- Nothing was committed to git.