Skip to content

Commit 9581dd6

Browse files
authored
Merge pull request #66 from EpiModel/dev/doctor-escalate-exhausted
Announce a task that exhausts its restart cap
2 parents 2568938 + 0ae29a7 commit 9581dd6

4 files changed

Lines changed: 43 additions & 4 deletions

File tree

‎DESCRIPTION‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,5 @@
11
Package: EpiModelHPC
2-
Version: 2.9.1
2+
Version: 2.9.2
33
Date: 2026-07-28
44
Title: EpiModel Extensions for High-Performance Computing
55
Description: Extension package to EpiModel to run large-scale stochastic network

‎NEWS.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,3 +1,11 @@
1+
# EpiModelHPC 2.9.2
2+
3+
## NEW FEATURES
4+
5+
- `degen_watch.sh` announces a task that exhausts `MAX_RESTARTS` instead of retiring it quietly. The cap is a terminal state: the task is confirmed pathological, the doctor stops intervening, and the task then holds its slot until walltime producing nothing. It previously said so with one lowercase `restarts exhausted; leaving it` in the middle of a verbose sweep, which is the same shape of silent ending as a TIME_LIMIT kill mailed to nobody. The line is now a `!! EXHAUSTED` token carrying the task, classification, node and restart count; the sweep summary gains an `exhausted=` counter; and setting `MAIL_TO` sends one message per newly exhausted task. Alerts are emitted once per task rather than once per sweep, deduplicated through the campaign-scoped ledger already in `STATE_FILE`, since the doctor re-probes every ten minutes and an exhausted task stays exhausted.
6+
7+
Observed on a `swfcalib` campaign that ran the pre-2.8.3 classifier: task `41767827_48` wedged in PSOCK worker startup three times, getting 3, 3 and 1 of its 8 workers through package loading before each requeue, and took about 4.5 hours to deliver one 95-minute batch. Nothing surfaced that except reading the sweep log by hand.
8+
19
# EpiModelHPC 2.9.1
210

311
## BUG FIXES

‎inst/hpc_doctor/README.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -153,6 +153,14 @@ Three properties made it hard to see, and each is now addressed:
153153

154154
Only `MAX_RESTARTS` stopped it: after three rounds the tasks reached their restart cap, the doctor logged `restarts exhausted; leaving it`, and the array was finally allowed to finish. That cap was doing work it was never meant to do.
155155

156+
### Exhausting the restart cap is a terminal state and now reads like one
157+
158+
The cap is the last thing standing between a pathological task and an unbounded requeue loop, and reaching it says something definite: the doctor has given up, and the task will hold its slot until walltime producing nothing. It used to say that with one lowercase line in the middle of a verbose sweep, which is the same shape of silent ending as a TIME_LIMIT kill mailed to nobody.
159+
160+
Seen again on a fifth campaign, this time running the pre-2.8.3 classifier: `swfcalib` task `41767827_48` wedged in PSOCK worker startup and was requeued three times. Segmenting its accumulated log by attempt (SLURM appends across requeues) shows 3 of 8 workers through package loading on the first attempt, 3 of 8 on the second, 1 of 8 on the third, then all 8 on the fourth, which ran 1.42 hours and completed. About 4.5 hours to deliver one 95-minute batch, and nothing surfaced it except reading the sweep log by hand afterwards.
161+
162+
So the cap now emits a `!! EXHAUSTED` token carrying the task, its classification, the node and the restart count; the sweep summary gains an `exhausted=` counter alongside the other two; and `MAIL_TO`, if set, receives one message per newly exhausted task. The alert fires once per task rather than once per sweep. The doctor re-probes every ten minutes and an exhausted task stays exhausted, so an undeduplicated alert would repeat until walltime and train the reader to skip it. The campaign-scoped ledger in `STATE_FILE` is the dedup key, and its `exhausted <taskid>` rows cannot collide with the node-offense rows, which are matched on a leading node name.
163+
156164
The probe now reads `SLURM_ARRAY_JOB_ID` and `SLURM_ARRAY_TASK_ID` from the same environment block and emits `taskid=`, and every `squeue`/`scontrol` call in `degen_watch.sh` addresses the task by it. Taking the id from the task's own environment rather than resolving it later is the point: PSOCK workers inherit it from their master, so every process of a task agrees, and nothing has to be inferred from a `squeue` line that may describe a sibling. Verified against the successor array while it ran: `jobid=41746875 taskid=41746874_0`, where the raw `SLURM_JOB_ID` and the array id differ by one and neither is guessable from the other. A guard before the destructive call refuses any unqualified id that `scontrol` reports as an array, so the failure mode cannot return by another route.
157165

158166
Two smaller consequences worth keeping in mind. The probe match was a bare `grep jobid=$jid`, which substring-matches any longer id sharing the prefix; it is now anchored on the field boundary, because acting on the wrong task is the one mistake this script must not make. And `Restarts=`, the age lookup, and the `ExcNodeList` read all went through the same unqualified id, so on an array they were reading whichever task `scontrol` or `squeue` happened to list first rather than the one under judgement.

‎inst/hpc_doctor/degen_watch.sh‎

Lines changed: 26 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -215,7 +215,7 @@ done
215215
echo "--- strike 1 done; re-checking suspects in ${RECHECK}s (transient dips must not trigger) ---"
216216
sleep "$RECHECK"
217217

218-
acted=0; cleared=0
218+
acted=0; cleared=0; exhausted=0
219219
for s in $suspects; do
220220
IFS=: read -r jid node tid <<< "$s"
221221
tid=${tid:-$jid}
@@ -353,7 +353,30 @@ for s in $suspects; do
353353
fi
354354

355355
[ "$REQUEUE" = "1" ] || continue
356-
if [ "$restarts" -ge "$MAX_RESTARTS" ]; then echo " restarts exhausted; leaving it"; continue; fi
356+
# Exhausting MAX_RESTARTS is a terminal state and has to read like one. The task
357+
# is confirmed pathological, the doctor has stopped intervening, and it will now
358+
# hold its slot until walltime producing nothing. Before this it announced
359+
# itself with one lowercase line in the middle of a verbose sweep, which is the
360+
# same shape of silent ending as a TIME_LIMIT kill mailed to nobody.
361+
#
362+
# Emitted once per task, not once per sweep: the doctor re-probes every 10
363+
# minutes and an exhausted task stays exhausted, so an undeduplicated alert
364+
# would repeat until walltime and train the reader to skip it. The campaign
365+
# ledger already in STATE_FILE is the natural dedup key, and is campaign-scoped
366+
# for free because the filename carries the doctor's own job id.
367+
if [ "$restarts" -ge "$MAX_RESTARTS" ]; then
368+
if ! grep -qxF "exhausted $tid" "$STATE_FILE" 2>/dev/null; then
369+
echo "exhausted $tid" >> "$STATE_FILE" 2>/dev/null || true
370+
echo " !! EXHAUSTED $tid: $kind on $node after $restarts restarts (MAX_RESTARTS=$MAX_RESTARTS); no longer intervening, it will hold its slot until walltime"
371+
exhausted=$((exhausted+1))
372+
if [ -n "${MAIL_TO:-}" ] && command -v mail >/dev/null 2>&1; then
373+
printf 'task %s exhausted %s restarts and is still %s on %s.\nThe doctor has stopped intervening; the task will hold its slot until walltime.\nDoctor job: %s\nCampaign pattern: /%s/\n' \
374+
"$tid" "$restarts" "$kind" "$node" "${SLURM_JOB_ID:-manual}" "$PATTERN" \
375+
| mail -s "deploy doctor: $tid exhausted its restarts" "$MAIL_TO" 2>/dev/null || true
376+
fi
377+
fi
378+
continue
379+
fi
357380
# Last line of defence before the destructive call. An unqualified ArrayJobId
358381
# would requeue the whole array, and the failure is silent: the sibling tasks
359382
# simply vanish from the next probe and get logged as "gone (finished/moved)".
@@ -380,5 +403,5 @@ for s in $suspects; do
380403
done
381404

382405
echo "---"
383-
echo "confirmed_starved_requeued=$acted cleared_as_transient=$cleared"
406+
echo "confirmed_starved_requeued=$acted cleared_as_transient=$cleared exhausted=$exhausted"
384407
[ "$REQUEUE" = "1" ] || echo "report-only; pass --requeue to act"

0 commit comments

Comments
 (0)