mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
Stopped a dead apt mirror from taking the whole install down with it (#642)
install.sh reaches the network four times, and on this runner pool that is not dependable. apt-get update stalled seven times in a single day: once for 55 minutes, once for more than two hours, and five times against the ten minute step timeout added alongside this. The log says the same thing every time. Every fetch from azure.archive.ubuntu.com comes back Ign, apt falls back to archive.ubuntu.com, and then the step produces no further output at all until something kills it. Nothing here bounded a fetch and nothing retried one, so a mirror being down cost a whole run instead of a few seconds. There is a second problem in the same lines. This script has no set -e, so a failed apt-get update did not stop the apt-get install that follows. The install went ahead against whatever package index the image happened to have, and the run failed later, somewhere with much less to say about why. Bound each attempt from outside and retry it. apt's own Acquire timeouts were tried first and are not enough: with them in place a run still sat inside a single apt-get update for nine and a half minutes without printing a line, having got as far as fetching noble-security InRelease. The retry loop never got a turn, because the first attempt never returned, and the step timeout was what eventually killed it. Whatever apt waits on there is not what Acquire::http::Timeout covers, so the bound has to come from outside the process. timeout does not care where the wait is. The Acquire options are kept anyway, since they make a slow mirror give up sooner, and pip gets its own retry and timeout flags for the same reason. timeout goes under sudo rather than over it, so that it signals apt itself. Signalling sudo risks the kill landing on sudo while apt carries on holding the dpkg lock, which would leave every retry failing for a different reason than the one being retried. The explicit exits stop a failed fetch being carried forward into a build. set -e is deliberately not used. rm -rf /opt/hostedtoolcache runs without sudo against a root owned directory and its exit status is not something this script should start depending on. The bounds fit inside the ten minute step timeout. Two minutes per attempt, three attempts, with 10 and 20 second backoffs, caps a command at about six and a half minutes, and a command that exhausts its attempts exits rather than letting the next one start. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
dafb7d70cc
commit
f2d27de25e
+53
-6
@@ -16,8 +16,55 @@
|
||||
# Remove large folder to save space
|
||||
rm -rf /opt/hostedtoolcache
|
||||
|
||||
sudo apt-get update
|
||||
sudo apt-get install -y \
|
||||
# Everything below reaches the network, and on this runner pool that is not
|
||||
# dependable. apt-get update stalled seven times in a single day, once for more
|
||||
# than two hours, each time with the Azure mirror returning nothing and the
|
||||
# fallback to archive.ubuntu.com then going silent. Nothing here bounded a fetch
|
||||
# and nothing retried one, so a mirror being down cost a whole run rather than a
|
||||
# few seconds. Worse, this script has no set -e, so a failed update did not stop
|
||||
# the install that follows: it went on to install from whatever index it already
|
||||
# had, and the run failed later somewhere less obvious.
|
||||
#
|
||||
# Each command is wrapped in timeout rather than left to bound itself. apt's own
|
||||
# Acquire timeouts were tried first and did not help: a run still sat inside a
|
||||
# single apt-get update for nine and a half minutes without producing a line,
|
||||
# having got as far as fetching noble-security InRelease, so the retry loop never
|
||||
# got a turn and the step timeout was what eventually killed it. Whatever apt is
|
||||
# waiting on there, it is not something Acquire::http::Timeout covers. timeout
|
||||
# does not care where the wait is.
|
||||
#
|
||||
# The Acquire options are kept anyway, since they make a slow mirror give up
|
||||
# sooner. The loop covers a mirror that is down rather than merely slow. The
|
||||
# explicit exits stop a failed fetch from being carried forward into a build.
|
||||
APT_OPTIONS=(-o Acquire::Retries=3
|
||||
-o Acquire::http::Timeout=20
|
||||
-o Acquire::https::Timeout=20)
|
||||
|
||||
# Two minutes per attempt, killed outright if it ignores the first signal. Three
|
||||
# attempts plus backoff bounds a command at about six and a half minutes, and a
|
||||
# command that exhausts its attempts exits rather than letting the next one run.
|
||||
#
|
||||
# timeout goes under sudo, not over it, so that it signals apt itself. Signalling
|
||||
# sudo instead risks the kill landing on sudo while apt carries on holding the
|
||||
# dpkg lock, which would leave every retry failing for a different reason than
|
||||
# the one being retried.
|
||||
TIMEOUT=(timeout --kill-after=10 120)
|
||||
|
||||
retry() {
|
||||
local attempt
|
||||
for attempt in 1 2 3; do
|
||||
if "$@"; then
|
||||
return 0
|
||||
fi
|
||||
echo "install.sh: '$*' failed or timed out on attempt ${attempt}"
|
||||
sleep $((attempt * 10))
|
||||
done
|
||||
echo "install.sh: '$*' failed after 3 attempts"
|
||||
return 1
|
||||
}
|
||||
|
||||
retry sudo "${TIMEOUT[@]}" apt-get "${APT_OPTIONS[@]}" update || exit 1
|
||||
retry sudo "${TIMEOUT[@]}" apt-get "${APT_OPTIONS[@]}" install -y \
|
||||
gcc-multilib \
|
||||
git \
|
||||
g++ \
|
||||
@@ -28,10 +75,10 @@ sudo apt-get install -y \
|
||||
tofrodos \
|
||||
gawk \
|
||||
cmake \
|
||||
software-properties-common
|
||||
software-properties-common || exit 1
|
||||
|
||||
python3 -m pip install --upgrade pip
|
||||
pip3 install gcovr==4.1
|
||||
retry "${TIMEOUT[@]}" python3 -m pip install --retries 3 --timeout 30 --upgrade pip || exit 1
|
||||
retry "${TIMEOUT[@]}" pip3 install --retries 3 --timeout 30 gcovr==4.1 || exit 1
|
||||
|
||||
# Upgrade cmake to the latest version.
|
||||
pip install --upgrade cmake
|
||||
retry "${TIMEOUT[@]}" pip install --retries 3 --timeout 30 --upgrade cmake || exit 1
|
||||
|
||||
Reference in New Issue
Block a user