Grail Computer · research record
Dated experiment record. Statements describe this run, not current submission readiness. Historical H. pylori category names do not establish infection; copy screens are bounded tests. Claims that a source does not exist mean none was identified in that recorded search, not proof of absence. Licensing interpretations in the notes remain unconfirmed. Local access details have been omitted. Current limitations and remaining work.

SWEEP-RUN result, 8 September 2026

Work package SWEEP-RUN (Phase 2 GPU sweep) of the SGH programme (operations/research/sgh-program-20260908/). Three spot A100s ran the three sweep-v1 packages to completion. 407 of 407 planned generations landed: 407 records, 407 PNGs, 0 skipped manifests, 0 failed jobs, 0 missing ids. Everything below is engineering evidence — throughput, provenance and infrastructure. No image was assessed for quality here and nothing here is clinically validated. Nothing was committed to git. All three slots are TERMINATED; slot d (F2 training) was not touched.

Merged records: [local]/records.jsonl (407 lines, 1.6 MB; every output rewritten to slot<i>/out/<path> and verified to resolve to a local file).

deliverable path (under packages/sweep-v1/)
merged records retrieved/records.jsonl
merge script retrieved/merge_records.py
per-manifest coverage + duplicate log retrieved/coverage.json
per-slot evidence retrieved/evidence-slot{0,1,2}-steps.log, -manifest-summary.json
per-slot outputs (5.5 GB, 407 PNGs) retrieved/slot{0,1,2}/out/
token library / stats / organism library / VM cartoon manifests retrieved/aux/
VM-vs-Mac cartoon hash comparison retrieved/aux/cartoon-sha-comparison.json
early pipeline check (2 PNGs + metrics) early/
per-slot orchestrator logs ../../fleet/logs/sweep-{a,b,c}.out, sweep-timeline.tsv

Slot map: package slot0 -> fleet slot a (sgh-a100-slot-a, us-central1-a), slot1 -> b (sgh-histo-qwen-a100-recovery-b, us-central1-b), slot2 -> c (sgh-a100-slot-c, us-central1-c).

1. Timeline

All times UTC, 8 September 2026.

what slot a slot b slot c
instances start issued 14:31:13 14:31:46 14:32:14
ssh + GPU ready 14:32:18 (65 s) 14:33:02 (76 s) 14:33:22 (68 s)
stage (43 files, ~22 MB) 14:32:26-14:34:41 (135 s) 14:33:10-14:35:15 (125 s) 14:33:31-14:35:28 (117 s)
worker #1 start 14:34:58 14:35:31 14:35:47
worker #1 end (failed: no scipy) 14:36:39 ran=1 skipped=16 14:37:28 ran=1 skipped=16 14:36:03 ran=0 skipped=16
scipy 1.14.1 installed 14:37-14:39 14:37-14:39 14:37-14:39
worker #2 start (serial render) 14:40:05 14:40:05 14:40:04
parallel setup takes over 14:42:58 / 14:48:27 / 14:49:53 same same
cartoons complete (a1 60, b1 5, d1 10, e2 20) ~15:08 ~15:08 ~15:08
token library (2016 windows) + stats 15:14:56 / 15:15:01 15:14:56 / 15:15:02 15:14:57 / 15:15:02
first manifest job (a1-pass1) 15:16:02 15:16:04 15:16:03
preemption 15:38:39 (real) none 15:31:08 (false positive, see fix 6)
restarted + resumed 15:39:19 up, 15:42:36 worker (records.jsonl 45 -> 45 usable) - not stopped; duplicate worker killed 15:41
last manifest job 16:53:42 16:46:11 16:52:09
WORKER_DONE (exit 0) 16:54:32 16:47:04 16:53:18
retrieve (1.8 GB each) 16:54:32-17:06:26 (714 s) 16:47:04-16:58:36 (692 s) 16:53:18-17:04:52 (694 s)
SHA256SUMS verified 143 files OK 131 files OK 133 files OK
stopped 17:07:30 16:59:39 17:06:05

Setup (boot to first generation) was 45 minutes; generation was 90-98 minutes; retrieval ~12 minutes. Retrieval ran at ~2.6 MB/s per slot in parallel (5.5 GB total).

fleet.sh status at the end:

SLOT  VM                               ZONE            STATUS
a     sgh-a100-slot-a                  us-central1-a   TERMINATED
b     sgh-histo-qwen-a100-recovery-b   us-central1-b   TERMINATED
c     sgh-a100-slot-c                  us-central1-c   TERMINATED
d     sgh-a100-slot-d                  us-central1-f   RUNNING      <- F2, untouched

2. Coverage: expected vs present, per manifest

From retrieved/merge_records.py (expected = the ids in each slot's own jobs/<manifest>.json; b1-native is compared against the copy worker.sh rewrote on the VM, pulled back into retrieved/aux/jobs/). Every cell is exp = rec = png. No missing ids, no skipped manifests, no failed jobs.

manifest a exp/rec/png b exp/rec/png c exp/rec/png total
a1-pass1 19/19/19 20/20/20 21/21/21 60
a1-pass2 19/19/19 20/20/20 21/21/21 60
a1-pass2-si9 5/5/5 4/4/4 6/6/6 15
a1-pass2-si3 19/19/19 20/20/20 21/21/21 60
a1-pass3-si15 3/3/3 3/3/3 4/4/4 10
d1-pass1 8/8/8 8/8/8 9/9/9 25
d1-pass2 8/8/8 8/8/8 9/9/9 25
a2-pass1 8/8/8 6/6/6 6/6/6 20
a2-pass2 8/8/8 6/6/6 6/6/6 20
a3-pass1 6/6/6 7/7/7 7/7/7 20
a3-pass2 6/6/6 7/7/7 7/7/7 20
d2-pass2 8/8/8 8/8/8 4/4/4 20
e2-pass1 5/5/5 5/5/5 5/5/5 15
e2-pass2 5/5/5 5/5/5 5/5/5 15
b1-pass1 2/2/2 2/2/2 1/1/1 5
b1-pass2 2/2/2 2/2/2 1/1/1 5
b1-native 12/12/12 - (fix 2) - 12
slot total 143/143/143 131/131/131 133/133/133 407

evidence/manifest-summary.json: slot a {"ran":17,"skipped":0,"failed":0}, slots b and c {"ran":16,"skipped":0,"failed":0}.

407 vs the split's 405 planned: b1-native was written against a placeholder naming convention and worker.sh rewrites it from the real organisms/library/*.png listing, which holds 12 windows, not 10. All 12 ran (on slot a only — see fix 2).

Provenance check on the merged file: 407 of 407 records have token_source equal to their extra.expected_token_sourcetoken_reference 285, token_library 50, token_file 40, token_blend 20, reference 12 (the native 1024 windows). The three token sources GEN-CODE added all worked on real weights; e.g. an A2 record carries token_library_entries=2016 and 21 per-window picks with library_file, library_category and distance.

Outputs by experiment: a1 205, d1 50, a2 40, a3 40, e2 30, d2 20, b1 10, b1-native 12. By category: intestinal_metaplasia 180, hpylori_gastritis 117, normal 55, mixed 55. 395 of the 407 canvases are 4096x2048; the 12 b1-native are single 1024 windows.

The only bookkeeping wrinkle

Slot c produced 5 duplicate records (a1_p2_hpylori_gastritis_s11_across, _s12_across, _s12_along, _s13_oblique, _s15_across) during the 11 minutes it ran two workers (fix 6). All five duplicates are byte-identical to the copy that was kept (same output_sha256), so merge_records.py keeps the first and logs the drop in coverage.json (duplicate_records_dropped). 412 lines read, 407 unique written.

3. Throughput

From the 407 merged records. denoiser_calls is the honest cost axis: the canvas is 21 overlapping 1024 windows on one latent, batched 4 at a time.

start_index steps run denoiser calls n median s mean s min max
15 5 30 10 25.28 25.34 25.16 25.66
12 8 48 145 31.84 28.31 23.41 32.51
9 11 66 15 38.71 37.09 30.16 39.09
6 14 84 165 45.29 43.50 36.87 95.71
3 17 102 60 52.28 51.43 43.59 52.85
(native 1024 window, si 12) 8 20 12 3.02 3.12 3.01 4.24

That is 0.37 s per denoiser call plus ~7 s of fixed cost on a 4096x2048 canvas, linear across the whole range. Total generation time over all 407 jobs: 15 216 s = 4.23 GPU-hours, mean 37.4 s per generation. This is ~30% faster than FLEET's smoke figure of 23.5 s for a pass because the sweep's median job is a heavier pass 2.

start_index 12 is bimodal and the split is informative: 81 jobs at a median of 31.9 s are all token_reference (first use of that donor field in the process), 64 jobs at a median of 23.6 s are token_library, token_file, or a token_reference whose donor was already in TokenCache. The ~8.3 s difference is the UNI2-h encode of one donor field (21 windows x 16 crops). Anything that reuses a donor across jobs in one process gets that back.

The 95.71 s outlier under si6 is from slot c's double-worker window: two generators sharing one A100 each took about twice as long, which is how the fault was spotted.

GPU utilisation: 4.23 GPU-hours of generation against 7.61 powered hours = 56%. The other 44% is the 45-minute setup, the ~12-minute retrieval and the two faults below.

4. Fixes applied

Six, in the order they were found. Every one is reproduced below with what it broke and how it was verified.

Fix 1 — the VM venv has no scipy (fatal, and silently so)

code/tissue3d.py:44 from scipy.spatial import cKDTree -> ModuleNotFoundError on all three slots. That kills render_set.py (all four cartoon sets) and build_token_library.py (both libraries), so worker.sh correctly skipped 16 of its 17 manifests ("no job has all of its inputs: reference") and exited 0 with WORKER_DONE. The first run therefore looked like a success and produced 12 PNGs out of 407. A worker that skips everything and reports zero failures is the failure mode to watch for on this rig.

The asset-root venv is Python 3.12.3 / numpy 1.26.4 / Pillow 11.1.0 / torch 2.9.1+cu129, with no scipy at all. Fixed with .venv/bin/pip install --no-deps 'scipy==1.14.1' on each slot — --no-deps and the pin so numpy stays at 1.26.4 (a bare pip install scipy can pull numpy 2.x under torch). Verified by importing all 12 package modules on slot a: 10 import cleanly; the two that do not are off the generation path (pixcell_stain_match needs scikit-image and is only reached when STAIN_TARGET is set, which it is not; organism_detector wants organisms/detector.py, which package B1 did not ship). The venv lives on the boot disk, so the fix survived the stop/start and the preemption. Anything that boots these disks needs the same install.

Fix 2 — b1-native would have been generated twice

worker.sh rewrites jobs/b1-native.json from the real organisms/library/*.png listing before running it, which discards the split's per-slot partition of that manifest. Slots a and b therefore each ran all 12 organism windows, with identical ids.

The accident is a free determinism check: the 12 output_sha256 values were identical between us-central1-a and us-central1-b, on top of FLEET's earlier bit-identical b/d result. Slot a's copy was kept; slot b's 12 records and PNGs were removed (backed up on the VM as evidence/records-b1native-removed.jsonl), jobs/b1-native.json deleted there, and b1-native dropped from packages/sweep-v1/slot1/MANIFEST-ORDER.txt (the manifest moved to slot1/jobs-removed/). Any manifest worker.sh rewrites from an on-VM listing must be given to exactly one slot.

Fix 3 — the 20-minute self-stop timer outlives the worker that armed it

worker.sh's EXIT trap arms systemd-run --on-active=20m shutdown, and it is cancelled only by the next worker's own trap, i.e. at the end of that worker's run. So a worker killed or replaced mid-run leaves a live timer that powers the VM off 20 minutes into its successor. (FLEET gotcha 4 fixed the re-arm; this is the other half.) Every relaunch in this package explicitly runs systemctl stop sgh-fleet-stop.timer; systemctl reset-failed sgh-fleet-stop.service immediately after fleet.sh launch, and this is now built into the driver script. Verified by systemctl list-timers showing no armed unit after each relaunch.

Fix 4 — the serial setup would have idled the A100 for ~100 minutes per slot

Measured on the VM (12 vCPU, single-threaded renderer):

step serial cost per slot
cartoon render, 4096x2048 45-60 s each (15 s on this Mac); 95 cartoons ~ 75 min
token library 32 s per donor field (21 windows); 96 fields ~ 51 min

worker.sh does both before it runs any manifest, so ~2 hours of A100 at USD 2.12/h would have been spent with the GPU at 0%. Both jobs are embarrassingly parallel — a cartoon is a pure function of (category, seed, cut) and a donor field is encoded independently — so they were re-run as four concurrent chunks each, concurrently with one another:

Every artefact is built in a scratch directory and moved into place only when complete, so worker.sh's "already built" checks (non-empty cartoons/<set>, token-library/{tokens.pt,index.json}, token-stats/token-stats.json) can never see a half-finished one. Result: the whole setup finished in 45 minutes wall clock instead of ~2 hours, and the last 25 minutes of it had the GPU busy on the library encode (68-100% utilisation) rather than idle.

Fix 5 — merge_library.py re-based token_index on the wrong counter

First version did e["token_index"] = base + e["token_index"] with base = len(grids) — the number of parts merged so far (0,1,2,3), not the number of windows. Caught by the script's own assertion (token_index not contiguous), so nothing corrupt was written, but it failed after the expensive part and left the three GPUs idle from 15:09:47 to 15:15:52 (~6 minutes x 3 slots). Fixed to a running window offset, with the per-chunk index also asserted contiguous before use. Re-merge took 5 s per slot and produced 2016 windows = 96 fields x 21, exactly the expected count, on all three slots; token_stats.py then ran in another 5 s.

Fix 6 — fleet.sh watch treats a failed describe as a preemption, and relaunches next to a live worker

At 15:31:08 slot c logged WATCH poll 13 status= -> PREEMPTED/STOPPED, recovering. The status was emptygcloud instances describe had failed transiently — not TERMINATED. cmd_up then correctly reported UP already RUNNING and started nothing, but cmd_watch went on to cmd_launch regardless, so a second worker started on a VM whose first worker was still running. Two generators shared one A100 for 11 minutes: per-job elapsed went from 45 s to 93-96 s and five ids were recorded twice.

Killed the newer worker's process group (kill -TERM -<pgid>; the worker is setsid, so this takes its generator child with it), removed the WORKER_DONE/exit_code.txt its EXIT trap had written, cancelled the timer it armed, and patched fleet/fleet.sh:

  1. an empty status is logged as status-unknown (describe failed), retrying and does not count as a preemption;
  2. before any relaunch, watch probes pgrep -f worker.sh on the VM and refuses to launch a second worker next to a live one.

The driver used for the rest of the run applies the same rule locally (launch only when the slot has no live worker), which is how slots b and c were re-attached without disturbing them.

Determinism caveat found before the run

Re-rendering cartoon set b1 on this Mac from the shipped package reproduced 14 of 15 files byte-for-byte; hpylori_gastritis_s21_oblique_labels.png differed in 8 bytes at the tail — same 331 856-byte length, same PNG chunk layout, same palette, pixel arrays identical, same zlib Adler-32; only the last deflate block and the IDAT CRC differ. PNG output is therefore not bit-reproducible even on one machine, so a hash mismatch is not by itself evidence that pixels differ. That mattered for section 6.

5. Preemption and resume

One real preemption, on slot a. watch poll 19 at 15:38:39 saw status=STOPPING, instances start was issued at 15:38:57, the VM was back at 15:39:19, and the worker was relaunched at 15:42:29. The resume is verified in the worker's own log:

2026-09-08T15:42:36Z resume: records.jsonl 45 lines -> 45 usable
2026-09-08T15:43:48Z a1-pass1 ok (45 records so far)

records.jsonl was not truncated by the power-off (45 of 45 lines parsed), the resumed worker re-resolved a1-pass1, skipped all 19 ids already present and added zero records for it in 72 s, then continued with a1-pass2. No job was regenerated. Slot a's final count (143) equals its expected count exactly.

Slot c's 15:31 event was not a preemption at all (fix 6); its worker was never stopped and its resume: records.jsonl 25 lines -> 25 usable line comes from the spurious second worker.

6. Was the VM rendering byte-identical to the local cartoons?

No. Comparing retrieved/aux/MANIFEST.sha256.<set>.json (rendered on slot a) against operations/research/sgh-program-20260908/cartoons/<set>/MANIFEST.sha256.json (rendered on this Mac by package A1):

set files compared byte-identical differing (cartoon / labels / json)
a1 180 1 60 / 60 / 59
b1 15 0 5 / 5 / 5
d1 30 0 10 / 10 / 10
e2 60 0 20 / 20 / 20

The difference is real, not an encoder artefact, and it is confined to one thing. One cartoon (a1/intestinal_metaplasia_s11_across) was pulled back whole and compared pixel by pixel:

So the two machines agree on the tissue architecture — glands, lumens, epithelium, goblet mucin, vessels are placed identically — and disagree only on the stochastic scatter of stromal nuclei, where a different number of nuclei is drawn and they land in different places. The cause is library skew, not the code: VM = Python 3.12.3 / numpy 1.26.4 / Pillow 11.1.0 / scipy 1.14.1; Mac = Python 3.14.7 / numpy 2.5.2 / Pillow 12.3.0 / scipy 1.18.1.

Practical consequence: the cartoons in cartoons/ are not the exact inputs of this sweep. Every record carries its own reference_sha256, so evaluation should resolve cartoons through the record, and any re-render intended to reproduce a specific generation has to pin the numpy/scipy/Pillow versions. The three VMs are consistent with each other (same image, same venv), which is what matters for comparing arms within the sweep.

retrieved/aux/ also holds token-library/index.json (2016 entries), composition-summary.md, build-meta.json, token-stats/token-stats.json, organism-library/index.json (12 windows, category forced to hpylori_gastritis) and the VM-side jobs/b1-native.json.

7. Early pipeline check

As soon as the first a1 pass-2 output existed, a1/pass2/intestinal_metaplasia/intestinal_metaplasia_s11_across.png and its pass-1 parent were pulled to packages/sweep-v1/early/ and measured with code/cartoon_metrics.py. The metric tool was first validated against FLEET's smoke PNG and reproduced its recorded ring_with_lumen_fraction of 0.8321 exactly.

stage ring_with_lumen_fraction ring_count lumen_tissue_ratio stroma_fraction
cartoon (input) 0.797 128 0.120 0.302
pass 1, start_index 12 0.545 132 0.088 0.319
pass 2, start_index 6 0.202 119 0.039 0.315

This is below the 0.7-0.95 band the brief expected, and the run was allowed to continue. The reasoning, on the evidence:

Stopping the fleet on this reading would have thrown away a verified 45-minute setup to re-measure something the sweep already covers. The number is reported here so Phase-4 evaluation starts from it rather than from the expectation.

Note for whoever repeats this: pass-1 and pass-2 outputs share a filename and differ only by directory, so copying both into one directory silently overwrites one. The first attempt here measured the pass-1 file twice; the files in early/ are now prefixed a1_pass1_ / a1_pass2_.

8. Powered time and cost

Reconstructed from fleet/state/<slot>/powered.log and cross-checked against the GCE lastStartTimestamp / lastStopTimestamp of each final session (slot a: 2026-09-08T08:39:17-07:00 / 10:07:28-07:00 = 15:39:17Z / 17:07:28Z, matching the log to 2 s). Sessions before 14:31 belong to the FLEET package and are excluded.

slot VM zone session up down minutes
a sgh-a100-slot-a us-central1-a 1 14:31:37 ~15:39:19 (preempted) 67.7
a 2 15:39:19 17:07:30 88.2
a total 155.9 (USD 5.51)
b sgh-histo-qwen-a100-recovery-b us-central1-b 1 14:32:09 16:59:39 147.5 (USD 5.21)
c sgh-a100-slot-c us-central1-c 1 14:32:36 17:06:05 153.5 (USD 5.42)
456.9 min = 7.61 h USD 16.14

At USD 2.12/h for a spot a2-highgpu-1g. Add roughly USD 0.66 of internet egress for the 5.5 GB retrieved. Slot a's first session ends at the moment cmd_up found it no longer RUNNING; the preemption itself was seen one poll earlier, so that figure is accurate to about a minute.

fleet.sh status's own accumulator under-reports here: it only counts UP/DOWN pairs, and a preemption writes no DOWN, so slot a's first 68 minutes are missing from it. The table above counts them.

Where the 7.61 powered hours went, per slot: ~4 min boot + ~2 min stage, 45 min setup (of which ~5 min lost to the scipy failure and ~6 min to the merge bug), 90-98 min generating (4.23 GPU-hours of actual denoising across the fleet), ~12 min retrieving, then stop.

9. What this package did not do

Download public Markdown export