Files
threadx/scripts
Frédéric DesbiensandClaude Opus 5 f2d27de25e Stopped a dead apt mirror from taking the whole install down with it (#642)
install.sh reaches the network four times, and on this runner pool that is not
dependable. apt-get update stalled seven times in a single day: once for 55
minutes, once for more than two hours, and five times against the ten minute
step timeout added alongside this. The log says the same thing every time. Every
fetch from azure.archive.ubuntu.com comes back Ign, apt falls back to
archive.ubuntu.com, and then the step produces no further output at all until
something kills it.

Nothing here bounded a fetch and nothing retried one, so a mirror being down
cost a whole run instead of a few seconds.

There is a second problem in the same lines. This script has no set -e, so a
failed apt-get update did not stop the apt-get install that follows. The install
went ahead against whatever package index the image happened to have, and the
run failed later, somewhere with much less to say about why.

Bound each attempt from outside and retry it. apt's own Acquire timeouts were
tried first and are not enough: with them in place a run still sat inside a
single apt-get update for nine and a half minutes without printing a line,
having got as far as fetching noble-security InRelease. The retry loop never got
a turn, because the first attempt never returned, and the step timeout was what
eventually killed it. Whatever apt waits on there is not what
Acquire::http::Timeout covers, so the bound has to come from outside the process.
timeout does not care where the wait is. The Acquire options are kept anyway,
since they make a slow mirror give up sooner, and pip gets its own retry and
timeout flags for the same reason.

timeout goes under sudo rather than over it, so that it signals apt itself.
Signalling sudo risks the kill landing on sudo while apt carries on holding the
dpkg lock, which would leave every retry failing for a different reason than the
one being retried.

The explicit exits stop a failed fetch being carried forward into a build.

set -e is deliberately not used. rm -rf /opt/hostedtoolcache runs without sudo
against a root owned directory and its exit status is not something this script
should start depending on.

The bounds fit inside the ten minute step timeout. Two minutes per attempt,
three attempts, with 10 and 20 second backoffs, caps a command at about six and
a half minutes, and a command that exhausts its attempts exits rather than
letting the next one start.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:46:19 -04:00
..