Skip to content
By Chris Devine · Updated

Supervise jobs, not clocks: what kept killing my self-hosted CI runners

At 04:00 on Saturday 19 September, a nightly timer on my laptop stopped Docker Desktop, then crashed before it could start it again. Three of the four self-hosted GitHub Actions runners that run the CI for ClearMyInbox went with it, and stayed offline for 8 hours 14 minutes. The watchdog I had written for exactly this failure noticed within three seconds, texted me that it was restarting Docker, and then failed in exactly the same way, 17 times. While I was digging through the logs I found a second, quieter bug that had been killing jobs at the two-hour mark for a month. This is that self-hosted runner outage, start to finish.

12 minute read. All times are BST (UTC+1), the laptop's own clock, unless a log line says otherwise.

The setup: four runner lanes on one laptop

ClearMyInbox is a bulk unsubscribe tool I build on my own, around a day job. In August I moved its CI off GitHub-hosted runners and onto my laptop, a Ryzen 9 5900HX with 8 cores, to stop paying for Actions minutes. Warm package caches on a local disk also beat pulling gigabytes through the cache service over a home connection.

The laptop runs four runner "lanes". Lanes 1, 2 and 4 are throwaway containers on Docker Desktop: each takes one job with --ephemeral, exits, and is replaced by a fresh one with a new registration. A small Bash supervisor, lane.sh, runs each lane in Git Bash, and Task Scheduler revives the supervisors every five minutes if one dies. Lane 3 is a native runner inside the WSL2 Ubuntu distro, for the two jobs that need a real Docker daemon (service containers and docker build).

Two safety nets sit around that. A heartbeat pings Better Stack every few minutes, but only while a runner container is actually up. And in the same Ubuntu distro, a watchdog under systemd --user polls Docker every 20 seconds. If Docker is down it messages me on Telegram and restarts Docker Desktop, at most once every half hour.

Architecture diagram. Task Scheduler on a Windows laptop revives four lane.sh supervisors. Three start runner containers on Docker Desktop's engine inside the WSL2 VM; the fourth drives a native runner in the Ubuntu distro. In the same distro, a systemd user timer called docker-nightly-restart can stop Docker Desktop at 04:00.
The whole CI estate. The red arrow is the one that matters: something inside the Linux distro can stop the engine every other lane depends on.

Timeline

Timeline of 19 September. Docker Desktop's engine is stopped from 04:02 to 12:16. Lanes 1, 2 and 4 are offline over the same period while their supervisors retry every six seconds. Lane 3 stays online. The watchdog fails 17 times, every thirty minutes.
Eight hours in which every alarm was technically working. The watchdog row is 17 identical failures.
TimeWhat happened
Fri 11:27The nightly Docker Desktop restart timer is installed. It has not yet run under systemd.
Sat 04:00:12The timer fires. Its guard finds no live pipeline runs, and it asks Docker Desktop to quit.
04:02:46Outage begins. The engine stops and all three runner containers die with unexpected EOF. Their supervisors start failing every 6 seconds.
04:02:49The watchdog sees Docker down, texts me "Restarting Docker Desktop", and tries. FileNotFoundError. (1 of 17.)
04:02:58The nightly restart crashes on the same line, before it can start Docker again.
04:23Better Stack opens an incident for the missing heartbeat, by email.
10:42-10:54A run on main and two PR runs queue. No runner picks them up.
about 11:45I notice.
12:16:32Docker Desktop is started by hand.
12:17:32Outage ends. All four lanes are online.
12:18:13The queued main job starts on lane 4.
12:35:38Main is green and the deploy runs.
13:11-13:19The fixed supervisor is rolled out, one lane at a time as each goes idle.

Bug 1: a restart that could only stop

The timer was less than a day old. On the Friday morning Docker Desktop had needed a restart, and its port forwarding on job networks took 6 to 12 minutes to come back afterwards, long enough to fail two CI jobs that talk to a storage emulator in a service container. A quiet restart at 04:00 looked like a cheap way to start each day fresh, and to measure whether a graceful restart recovers faster than a forced one. It even had a guard: skip the restart if my agent pipeline has anything running. It did not know CI existed, though that turned out not to matter on this particular night, because CI was idle at 04:00.

The restart lives in one Python function that does three things: stop Docker Desktop, wait until it has really gone, start it again. The lane logs (in UTC) show the stop landing:

time="2026-09-19T04:02:46+01:00" level=error msg="error waiting for container: unexpected EOF"2026-09-19T03:02:47Z container exited (job done or idle-timeout), recycling2026-09-19T03:02:53Z starting container for devin-5900hx-lane1docker: error during connect: Head "http://%2F%2F.%2Fpipe%2FdockerDesktopLinuxEngine/_ping": open //./pipe/dockerDesktopLinuxEngine: The system cannot find the file specified.2026-09-19T03:02:53Z container exited (job done or idle-timeout), recycling2026-09-19T03:02:59Z starting container for devin-5900hx-lane1

And the systemd journal (trimmed) shows the function dying on the second step, twelve seconds later:

04:02:58 python3[34202]:   File "notify.py", line 442, in _docker_desktop_gone04:02:58 python3[34202]:     return docker_desktop_state() == "stopped" and _windows_docker_procs() == 004:02:58 python3[34202]:   File "notify.py", line 420, in _windows_docker_procs04:02:58 python3[34202]:     r = subprocess.run(["powershell.exe", "-NoProfile", "-Command",04:02:58 python3[34202]: FileNotFoundError: [Errno 2] No such file or directory: 'powershell.exe'04:02:58 systemd[206]: docker-nightly-restart.service: Failed with result 'exit-code'.

The stop worked because the code calls Docker's Windows CLI by its full path, /mnt/c/Program Files/Docker/Docker/resources/bin/docker.exe. The "has it gone?" check calls powershell.exe by bare name, and the unit it runs in sets its own PATH. The watchdog's unit has the identical line, with a comment that explains where it came from: "the same PATH the other pipeline units use".

# ~/.config/systemd/user/docker-nightly-restart.service[Service]Type=oneshotEnvironmentFile=%h/.config/pipeline-notify/envExecStart=/usr/bin/python3 -u /home/devine/agents/scripts/docker-nightly-restart.pyEnvironment=PATH=/usr/bin:/home/devine/.local/bin:/usr/local/bin:/bin

That is a perfectly good PATH for a Linux program. It has no Windows directories in it. In an interactive WSL terminal you never notice, because WSL appends the whole Windows PATH to your shell's. Take the unit's line away and you are no better off, because systemd's own user manager doesn't have them either:

$ command -v powershell.exe                    # in a WSL terminal/mnt/c/Windows/System32/WindowsPowerShell/v1.0/powershell.exe $ systemd-run --user --wait -P sh -c 'command -v powershell.exe || echo not found'not found

This is the cron trap in a Windows costume: the script is correct in the environment you try it in, and wrong in the one that runs it. The absolute path on docker.exe was there for an unrelated reason (inside WSL a bare docker finds the Linux CLI, not Docker Desktop's), so the step that breaks things was robust by accident, and the step that fixes them was not. The start command further down, schtasks.exe, was a bare name too. It never got the chance to fail.

Why the watchdog didn't save it

The watchdog did its detection job brilliantly. It saw Docker go down three seconds after the engine stopped and sent my phone a Telegram message: "DOCKER DOWN ... Restarting Docker Desktop". Then it called the same restart_docker_desktop() the timer was about to crash in, under a unit with the same PATH. The exception escaped before the line that counts failed restarts, so the watchdog never gave up, and never said another word until Docker was back. It just tried again every half hour until lunchtime:

04:02:49 ERROR docker watchdog failed: [Errno 2] No such file or directory: 'powershell.exe'04:31:21 ERROR docker watchdog failed: [Errno 2] No such file or directory: 'powershell.exe'   ... 14 more, one every half hour ...12:04:26 ERROR docker watchdog failed: [Errno 2] No such file or directory: 'powershell.exe'
Diagram of shared fate. The nightly restart timer and the Docker watchdog both call restart_docker_desktop(). Step one, docker.exe desktop stop, is called by absolute path and works. Step two calls powershell.exe by bare name and raises FileNotFoundError, because both units set a PATH with no Windows directories. Step three, starting Docker Desktop, is never reached.
The thing that broke Docker and the thing meant to fix it were the same code, in the same environment. They could only ever fail together.

So the one alert that reached my phone told me the problem was being handled. The heartbeat stopped pinging at 04:02:48 and Better Stack opened an incident at 04:23, twenty minutes later. It sent an email, which on my plan includes a line I now find very funny: "We can also call you next time, just upgrade your account." At 04:23 on a Saturday, an email is a log entry, not a page.

GitHub didn't complain either. A job waiting for a self-hosted runner can sit in the queue for up to 24 hours before it fails, so from GitHub's side the morning's runs were simply "Queued". And the lane supervisors, which retried every six seconds and minted a fresh registration token each time, made about 4,700 failed starts per lane. After every one, my own supervisor wrote the line you can see in the log above: container exited (job done or idle-timeout), recycling. It described more than 14,000 failures as normal operation.

Bug 2: the timeout that counted idle time

Once Docker was back, I went through the other CI failures I had been writing off as flakes: jobs that died mid-step with "The runner has received a shutdown signal" or similar, nowhere near their timeouts. Three had ordinary explanations. The laptop restarted during a job on 17 September; the Wi-Fi dropped for three minutes on 13 September and the runner lost GitHub's broker; and on 15 September a cancel request came from GitHub's side with nothing on the host at all.

Then I lined up every job that died without a result against the start time of the container it ran in, and three of them had the same number:

DateLane, jobJob started at container ageRan forKilled at container age
22 Aug4, sonar1:41:5918m 08s2:00:07
15 Sep2, terraform-validate1:58:521m 12s2:00:04
15 Sep4, sonar-web1:51:318m 32s2:00:03

The cause was one line in lane.sh:

# lane.sh, before: one clock around the whole containerMSYS_NO_PATHCONV=1 timeout -k 60 7200 docker run --rm --name "ci-lane-$LANE" \  --cpuset-cpus "$CPUSET_LO-$CPUSET_HI" --memory 4g --memory-swap 4g \  -e EPHEMERAL=true -e RUNNER_TOKEN="$TOKEN" ... \  ci-runner:toolchain >> "$LOG" 2>&1

The two-hour timeout was there for good reasons: a hung job shouldn't hold a lane forever, and a fresh registration every couple of hours keeps the runners tidy. But it starts counting when the container starts, and an ephemeral runner spends most of its life idle, waiting for a job. The timeout had no idea whether a job was running. A container that waited 1 hour 58 minutes and 52 seconds for work gave its job 72 seconds.

The extra 3 to 7 seconds past 2:00:00 are the shutdown. At exactly 7,200 seconds, timeout sends SIGTERM to the docker run client, which passes it on to the container. The runner then takes a few seconds to cancel the job and exit, less the second or so between timeout starting and the container starting.

Before and after chart for one container on lane 4. Before: the container idles for 1 hour 51 minutes, picks up sonar-web, and a two-hour timeout round the container kills the job 8 minutes 32 seconds in. After: the same job runs to completion, because the idle recycle only fires when no job is running and the hung-job ceiling counts from when the job starts.
The clock was measuring the container. The thing I cared about was the job.

It hid for a month because most jobs arrive within minutes of a container starting, so the clock almost never mattered. It only bit on a quiet stretch followed by a push. And when it did, the job looked like every other runner flake, and a re-run passed.

The fix: supervise jobs, not clocks

The supervisor now watches the job. lane.sh starts the container detached and polls it every 15 seconds. The test for "mid-job" is exact: a GitHub Actions runner only has a Runner.Worker process while it is executing a job, and an idle runner is just Runner.Listener. So the two jobs the old timeout did badly are now two separate limits:

# lane.sh, after (trimmed): poll every 15s and ask the runner what it is doingbusy=$($busy_fn)                  # is Runner.Worker in the process tree?case "$busy" in  yes)    [ -z "$job_since" ] && job_since=$now    if [ $((now - job_since)) -ge "$JOB_CEILING" ]; then      log "job ceiling: one job has run $((now - job_since))s - stopping as hung"      $stop_fn; return 0    fi ;;  no)    job_since=""    if [ $((now - started)) -ge "$IDLE_RECYCLE" ]; then      gh=$(gh_busy)               # and ask GitHub, which marks busy on assignment      if [ "$($busy_fn)" = "no" ] && { [ "$gh" = "false" ] || [ "$gh" = "absent" ]; }; then        log "idle recycle: no job in $((now - started))s - fresh registration"        $stop_fn; return 0      fi    fi ;;esac

The hung-job ceiling is still two hours, but it counts from when this job started. The idle recycle only fires when there is no Runner.Worker locally and GitHub doesn't show the runner as busy either, because GitHub marks a runner busy when it assigns the job, a beat before the worker process appears. A probe that fails is never read as "idle". The supervisor also checks docker info before minting a registration token, so the next time Docker is down each lane writes one log line and waits, instead of 4,700.

The systemd units can see Windows. A drop-in adds the two Windows directories to the PATH of the watchdog, the nightly restart and one other unit in the same pipeline:

# ~/.config/systemd/user/pipeline-notify.service.d/10-windows-interop-path.conf# (and the same file for docker-nightly-restart.service)[Service]Environment=PATH=/usr/bin:/home/devine/.local/bin:/usr/local/bin:/bin:/mnt/c/Windows/System32:/mnt/c/Windows/System32/WindowsPowerShell/v1.0

The nightly restart checks for CI. It now skips if any runner, container or native, is mid-job:

# docker-nightly-restart.py (simplified)def ci_job_running():    """Name of a CI runner that is mid-job, or None. Runner.Worker only exists    while a GitHub Actions runner executes a job, so its presence is exact."""    names = run([DOCKER_EXE, "ps", "--filter", "name=ci-lane", "--format", "{{.Names}}"]).split()    for name in names:        if "Runner.Worker" in run([DOCKER_EXE, "top", name]):            return name    if subprocess.run(["pgrep", "-f", "Runner.Worker"]).returncode == 0:        return "a native runner in the Ubuntu distro"    return None

I kept the nightly restart. It can finish what it starts now, and it is still the thing that tells me whether a graceful restart recovers faster than a forced one. The new supervisor has 14 offline tests that run against fake docker and gh commands in a few seconds, and went out lane by lane between 13:11 and 13:19, each lane swapped only once it was idle.

What went well, what didn't, where I got lucky

Went well

Went badly

Where I got lucky

If you run CI on your own hardware

What changed

ChangeWhyStatus
Windows directories on the PATH of three systemd user unitsBug 1 and the watchdogDone
Nightly restart skips while any CI job runsNever stop Docker under a jobDone
Supervisor: job ceiling from job start, idle recycle only when idleBug 2Done
Docker gate before minting a registration token4,700 wasted starts per laneDone
Closing the laptop lid on mains power does nothingFound on the way; a closed lid would sleep all four lanesDone
The watchdog reports a failed restart, and the heartbeat alert can wake me"Restarting..." then eight hours of silenceNext
Turn off Windows Fast StartupA "shut down" with it on is a hibernate, so the host never gets a clean bootNext
Runners that survive a reboot to the lock screenDocker Desktop needs a signed-in sessionKnown, not fixed

How I worked it out

I did the investigation with Claude Code running on the laptop itself. It read the lane logs, the systemd journal and Docker Desktop's own logs, lined the timestamps up, found the two-hour pattern, and wrote the new supervisor and its tests, which went through the same pull request review as everything else. It also drafted this write-up from the same logs, and every number in it comes from them. The diagrams are drawn from those timestamps too. How the rest of ClearMyInbox gets built that way, and what goes wrong, is in I counted Claude Code's mistakes for six months.

If you have hit the WSL and systemd PATH trap in a different shape, I would like to hear about it: chris@clearmyinbox.ai.