Supervise jobs, not clocks: what kept killing my self-hosted CI runners
At 04:00 on Saturday 19 September, a nightly timer on my laptop stopped Docker Desktop, then crashed before it could start it again. Three of the four self-hosted GitHub Actions runners that run the CI for ClearMyInbox went with it, and stayed offline for 8 hours 14 minutes. The watchdog I had written for exactly this failure noticed within three seconds, texted me that it was restarting Docker, and then failed in exactly the same way, 17 times. While I was digging through the logs I found a second, quieter bug that had been killing jobs at the two-hour mark for a month. This is that self-hosted runner outage, start to finish.
12 minute read. All times are BST (UTC+1), the laptop's own clock, unless a log line says otherwise.
The setup: four runner lanes on one laptop
ClearMyInbox is a bulk unsubscribe tool I build on my own, around a day job. In August I moved its CI off GitHub-hosted runners and onto my laptop, a Ryzen 9 5900HX with 8 cores, to stop paying for Actions minutes. Warm package caches on a local disk also beat pulling gigabytes through the cache service over a home connection.
The laptop runs four runner "lanes". Lanes 1, 2 and 4 are throwaway containers on Docker Desktop: each takes one job with --ephemeral, exits, and is replaced by a fresh one with a new registration. A small Bash supervisor, lane.sh, runs each lane in Git Bash, and Task Scheduler revives the supervisors every five minutes if one dies. Lane 3 is a native runner inside the WSL2 Ubuntu distro, for the two jobs that need a real Docker daemon (service containers and docker build).
Two safety nets sit around that. A heartbeat pings Better Stack every few minutes, but only while a runner container is actually up. And in the same Ubuntu distro, a watchdog under systemd --user polls Docker every 20 seconds. If Docker is down it messages me on Telegram and restarts Docker Desktop, at most once every half hour.
Timeline
| Time | What happened |
|---|---|
| Fri 11:27 | The nightly Docker Desktop restart timer is installed. It has not yet run under systemd. |
| Sat 04:00:12 | The timer fires. Its guard finds no live pipeline runs, and it asks Docker Desktop to quit. |
| 04:02:46 | Outage begins. The engine stops and all three runner containers die with unexpected EOF. Their supervisors start failing every 6 seconds. |
| 04:02:49 | The watchdog sees Docker down, texts me "Restarting Docker Desktop", and tries. FileNotFoundError. (1 of 17.) |
| 04:02:58 | The nightly restart crashes on the same line, before it can start Docker again. |
| 04:23 | Better Stack opens an incident for the missing heartbeat, by email. |
| 10:42-10:54 | A run on main and two PR runs queue. No runner picks them up. |
| about 11:45 | I notice. |
| 12:16:32 | Docker Desktop is started by hand. |
| 12:17:32 | Outage ends. All four lanes are online. |
| 12:18:13 | The queued main job starts on lane 4. |
| 12:35:38 | Main is green and the deploy runs. |
| 13:11-13:19 | The fixed supervisor is rolled out, one lane at a time as each goes idle. |
Bug 1: a restart that could only stop
The timer was less than a day old. On the Friday morning Docker Desktop had needed a restart, and its port forwarding on job networks took 6 to 12 minutes to come back afterwards, long enough to fail two CI jobs that talk to a storage emulator in a service container. A quiet restart at 04:00 looked like a cheap way to start each day fresh, and to measure whether a graceful restart recovers faster than a forced one. It even had a guard: skip the restart if my agent pipeline has anything running. It did not know CI existed, though that turned out not to matter on this particular night, because CI was idle at 04:00.
The restart lives in one Python function that does three things: stop Docker Desktop, wait until it has really gone, start it again. The lane logs (in UTC) show the stop landing:
time="2026-09-19T04:02:46+01:00" level=error msg="error waiting for container: unexpected EOF"2026-09-19T03:02:47Z container exited (job done or idle-timeout), recycling2026-09-19T03:02:53Z starting container for devin-5900hx-lane1docker: error during connect: Head "http://%2F%2F.%2Fpipe%2FdockerDesktopLinuxEngine/_ping": open //./pipe/dockerDesktopLinuxEngine: The system cannot find the file specified.2026-09-19T03:02:53Z container exited (job done or idle-timeout), recycling2026-09-19T03:02:59Z starting container for devin-5900hx-lane1And the systemd journal (trimmed) shows the function dying on the second step, twelve seconds later:
04:02:58 python3[34202]: File "notify.py", line 442, in _docker_desktop_gone04:02:58 python3[34202]: return docker_desktop_state() == "stopped" and _windows_docker_procs() == 004:02:58 python3[34202]: File "notify.py", line 420, in _windows_docker_procs04:02:58 python3[34202]: r = subprocess.run(["powershell.exe", "-NoProfile", "-Command",04:02:58 python3[34202]: FileNotFoundError: [Errno 2] No such file or directory: 'powershell.exe'04:02:58 systemd[206]: docker-nightly-restart.service: Failed with result 'exit-code'.The stop worked because the code calls Docker's Windows CLI by its full path, /mnt/c/Program Files/Docker/Docker/resources/bin/docker.exe. The "has it gone?" check calls powershell.exe by bare name, and the unit it runs in sets its own PATH. The watchdog's unit has the identical line, with a comment that explains where it came from: "the same PATH the other pipeline units use".
# ~/.config/systemd/user/docker-nightly-restart.service[Service]Type=oneshotEnvironmentFile=%h/.config/pipeline-notify/envExecStart=/usr/bin/python3 -u /home/devine/agents/scripts/docker-nightly-restart.pyEnvironment=PATH=/usr/bin:/home/devine/.local/bin:/usr/local/bin:/binThat is a perfectly good PATH for a Linux program. It has no Windows directories in it. In an interactive WSL terminal you never notice, because WSL appends the whole Windows PATH to your shell's. Take the unit's line away and you are no better off, because systemd's own user manager doesn't have them either:
$ command -v powershell.exe # in a WSL terminal/mnt/c/Windows/System32/WindowsPowerShell/v1.0/powershell.exe $ systemd-run --user --wait -P sh -c 'command -v powershell.exe || echo not found'not foundThis is the cron trap in a Windows costume: the script is correct in the environment you try it in, and wrong in the one that runs it. The absolute path on docker.exe was there for an unrelated reason (inside WSL a bare docker finds the Linux CLI, not Docker Desktop's), so the step that breaks things was robust by accident, and the step that fixes them was not. The start command further down, schtasks.exe, was a bare name too. It never got the chance to fail.
Why the watchdog didn't save it
The watchdog did its detection job brilliantly. It saw Docker go down three seconds after the engine stopped and sent my phone a Telegram message: "DOCKER DOWN ... Restarting Docker Desktop". Then it called the same restart_docker_desktop() the timer was about to crash in, under a unit with the same PATH. The exception escaped before the line that counts failed restarts, so the watchdog never gave up, and never said another word until Docker was back. It just tried again every half hour until lunchtime:
04:02:49 ERROR docker watchdog failed: [Errno 2] No such file or directory: 'powershell.exe'04:31:21 ERROR docker watchdog failed: [Errno 2] No such file or directory: 'powershell.exe' ... 14 more, one every half hour ...12:04:26 ERROR docker watchdog failed: [Errno 2] No such file or directory: 'powershell.exe'So the one alert that reached my phone told me the problem was being handled. The heartbeat stopped pinging at 04:02:48 and Better Stack opened an incident at 04:23, twenty minutes later. It sent an email, which on my plan includes a line I now find very funny: "We can also call you next time, just upgrade your account." At 04:23 on a Saturday, an email is a log entry, not a page.
GitHub didn't complain either. A job waiting for a self-hosted runner can sit in the queue for up to 24 hours before it fails, so from GitHub's side the morning's runs were simply "Queued". And the lane supervisors, which retried every six seconds and minted a fresh registration token each time, made about 4,700 failed starts per lane. After every one, my own supervisor wrote the line you can see in the log above: container exited (job done or idle-timeout), recycling. It described more than 14,000 failures as normal operation.
Bug 2: the timeout that counted idle time
Once Docker was back, I went through the other CI failures I had been writing off as flakes: jobs that died mid-step with "The runner has received a shutdown signal" or similar, nowhere near their timeouts. Three had ordinary explanations. The laptop restarted during a job on 17 September; the Wi-Fi dropped for three minutes on 13 September and the runner lost GitHub's broker; and on 15 September a cancel request came from GitHub's side with nothing on the host at all.
Then I lined up every job that died without a result against the start time of the container it ran in, and three of them had the same number:
| Date | Lane, job | Job started at container age | Ran for | Killed at container age |
|---|---|---|---|---|
| 22 Aug | 4, sonar | 1:41:59 | 18m 08s | 2:00:07 |
| 15 Sep | 2, terraform-validate | 1:58:52 | 1m 12s | 2:00:04 |
| 15 Sep | 4, sonar-web | 1:51:31 | 8m 32s | 2:00:03 |
The cause was one line in lane.sh:
# lane.sh, before: one clock around the whole containerMSYS_NO_PATHCONV=1 timeout -k 60 7200 docker run --rm --name "ci-lane-$LANE" \ --cpuset-cpus "$CPUSET_LO-$CPUSET_HI" --memory 4g --memory-swap 4g \ -e EPHEMERAL=true -e RUNNER_TOKEN="$TOKEN" ... \ ci-runner:toolchain >> "$LOG" 2>&1The two-hour timeout was there for good reasons: a hung job shouldn't hold a lane forever, and a fresh registration every couple of hours keeps the runners tidy. But it starts counting when the container starts, and an ephemeral runner spends most of its life idle, waiting for a job. The timeout had no idea whether a job was running. A container that waited 1 hour 58 minutes and 52 seconds for work gave its job 72 seconds.
The extra 3 to 7 seconds past 2:00:00 are the shutdown. At exactly 7,200 seconds, timeout sends SIGTERM to the docker run client, which passes it on to the container. The runner then takes a few seconds to cancel the job and exit, less the second or so between timeout starting and the container starting.
It hid for a month because most jobs arrive within minutes of a container starting, so the clock almost never mattered. It only bit on a quiet stretch followed by a push. And when it did, the job looked like every other runner flake, and a re-run passed.
The fix: supervise jobs, not clocks
The supervisor now watches the job. lane.sh starts the container detached and polls it every 15 seconds. The test for "mid-job" is exact: a GitHub Actions runner only has a Runner.Worker process while it is executing a job, and an idle runner is just Runner.Listener. So the two jobs the old timeout did badly are now two separate limits:
# lane.sh, after (trimmed): poll every 15s and ask the runner what it is doingbusy=$($busy_fn) # is Runner.Worker in the process tree?case "$busy" in yes) [ -z "$job_since" ] && job_since=$now if [ $((now - job_since)) -ge "$JOB_CEILING" ]; then log "job ceiling: one job has run $((now - job_since))s - stopping as hung" $stop_fn; return 0 fi ;; no) job_since="" if [ $((now - started)) -ge "$IDLE_RECYCLE" ]; then gh=$(gh_busy) # and ask GitHub, which marks busy on assignment if [ "$($busy_fn)" = "no" ] && { [ "$gh" = "false" ] || [ "$gh" = "absent" ]; }; then log "idle recycle: no job in $((now - started))s - fresh registration" $stop_fn; return 0 fi fi ;;esacThe hung-job ceiling is still two hours, but it counts from when this job started. The idle recycle only fires when there is no Runner.Worker locally and GitHub doesn't show the runner as busy either, because GitHub marks a runner busy when it assigns the job, a beat before the worker process appears. A probe that fails is never read as "idle". The supervisor also checks docker info before minting a registration token, so the next time Docker is down each lane writes one log line and waits, instead of 4,700.
The systemd units can see Windows. A drop-in adds the two Windows directories to the PATH of the watchdog, the nightly restart and one other unit in the same pipeline:
# ~/.config/systemd/user/pipeline-notify.service.d/10-windows-interop-path.conf# (and the same file for docker-nightly-restart.service)[Service]Environment=PATH=/usr/bin:/home/devine/.local/bin:/usr/local/bin:/bin:/mnt/c/Windows/System32:/mnt/c/Windows/System32/WindowsPowerShell/v1.0The nightly restart checks for CI. It now skips if any runner, container or native, is mid-job:
# docker-nightly-restart.py (simplified)def ci_job_running(): """Name of a CI runner that is mid-job, or None. Runner.Worker only exists while a GitHub Actions runner executes a job, so its presence is exact.""" names = run([DOCKER_EXE, "ps", "--filter", "name=ci-lane", "--format", "{{.Names}}"]).split() for name in names: if "Runner.Worker" in run([DOCKER_EXE, "top", name]): return name if subprocess.run(["pgrep", "-f", "Runner.Worker"]).returncode == 0: return "a native runner in the Ubuntu distro" return NoneI kept the nightly restart. It can finish what it starts now, and it is still the thing that tells me whether a graceful restart recovers faster than a forced one. The new supervisor has 14 offline tests that run against fake docker and gh commands in a few seconds, and went out lane by lane between 13:11 and 13:19, each lane swapped only once it was idle.
What went well, what didn't, where I got lucky
Went well
- Every signal existed. The watchdog noticed in 3 seconds, the heartbeat in 20 minutes, and every log had timestamps to the second, so finding the cause was reading, not guessing.
- Recovery was one command, and the ephemeral runners meant nothing was left half-done or corrupted. The queued jobs simply ran.
- Lane 3 survived, because it starts differently. Accidental diversity is still diversity.
Went badly
- The fix-it path and the break-it path were the same code in the same environment.
- The alert that reached my phone announced an action, not a result. "Restarting Docker Desktop" was the last thing the watchdog said for eight hours.
- My own log line called every failure a normal exit, 14,000 times.
- The timer's first real run was at 04:00 on a Saturday. I never ran it under systemd before I left it to run under systemd.
Where I got lucky
- It happened early on a Saturday. On a weekday it would have blocked a morning of merges.
- Production doesn't run on this laptop. CI only gates deploys, so customers saw nothing.
- Nothing was running at 04:00. The same restart mid-job would have killed it, and it would have looked like one more flake.
If you run CI on your own hardware
- Test a unit as a unit.
systemctl --user start thing.serviceand read the journal. Your shell is not the environment that will run it. - On WSL, call Windows binaries by absolute path, or set the PATH in the unit on purpose. The interop PATH is a property of your shell, not of Linux.
- Check you can start before you stop. Anything that restarts a shared dependency should prove its start path works first.
- Give the watchdog a different failure mode, and make at least one alert leave the machine in a way that wakes you.
- Report outcomes, not intentions. "Restarting..." needs a follow-up that says whether it worked, and silence after it should be an alarm of its own.
- Make timeouts measure the thing you care about. A clock around a container is not a clock around a job.
- Let your logs say what actually happened. "Exited, recycling" should have been "docker run failed after 1 second, 4,700th time".
What changed
| Change | Why | Status |
|---|---|---|
| Windows directories on the PATH of three systemd user units | Bug 1 and the watchdog | Done |
| Nightly restart skips while any CI job runs | Never stop Docker under a job | Done |
| Supervisor: job ceiling from job start, idle recycle only when idle | Bug 2 | Done |
| Docker gate before minting a registration token | 4,700 wasted starts per lane | Done |
| Closing the laptop lid on mains power does nothing | Found on the way; a closed lid would sleep all four lanes | Done |
| The watchdog reports a failed restart, and the heartbeat alert can wake me | "Restarting..." then eight hours of silence | Next |
| Turn off Windows Fast Startup | A "shut down" with it on is a hibernate, so the host never gets a clean boot | Next |
| Runners that survive a reboot to the lock screen | Docker Desktop needs a signed-in session | Known, not fixed |
How I worked it out
I did the investigation with Claude Code running on the laptop itself. It read the lane logs, the systemd journal and Docker Desktop's own logs, lined the timestamps up, found the two-hour pattern, and wrote the new supervisor and its tests, which went through the same pull request review as everything else. It also drafted this write-up from the same logs, and every number in it comes from them. The diagrams are drawn from those timestamps too. How the rest of ClearMyInbox gets built that way, and what goes wrong, is in I counted Claude Code's mistakes for six months.
If you have hit the WSL and systemd PATH trap in a different shape, I would like to hear about it: chris@clearmyinbox.ai.