mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
install.sh grew a retry loop, per-command timeouts and a deliberately non-gating apt-get update after this runner pool cost several whole runs: a mirror going silent for two hours, and a Hash Sum mismatch from a third-party repository the project does not even use turning builds red. Those lessons were local to that one file. install_riscv.sh had none of them, and #717 puts it on every pull request's critical path. Under set -e its bare apt-get update was a single point of failure for the whole suite -- the precise case install.sh downgrades to a warning on purpose -- and its two wget calls, each fetching about 500 MB, had no retry and no timeout. Rather than copy the helpers and let them drift again, they move to tx_ci_common.sh and both scripts source it, following the arrangement scripts/tx_windows_common.ps1 already uses on the Windows side. install.sh keeps its behaviour exactly: same APT_OPTIONS, same 120-second TIMEOUT, same three-attempt retry, and the comments explaining each of them travel with the code they explain. Two things are new: - TIMEOUT_LONG, 180 seconds, for a single large download. Sized against the 39 seconds each tarball took on 10 Sep 2026 and deliberately not larger: the install step is capped at ten minutes, and a per-attempt timeout able to swallow that cap would leave the retry loop no turn to take, which is the failure mode the apt comment already records. - fetch(), which verifies a SHA-256 before anything is unpacked. Both digests were taken from the releases API and then checked against the bytes the CDN actually serves. This is not an independent trust root -- expected value and file come from the same host -- but it pins the bytes, so a deleted and re-pushed tag or a replaced asset stops the build instead of being picked up silently. Also verifies qemu-system-riscv32 alongside riscv64. run.sh selects one per architecture, so both are worth failing on here rather than at the first test. Verified locally: retry returns 0 on success and 1 after three attempts; fetch accepts a correct digest and, on a wrong one, fails and removes the partial file; the source line resolves from the repository root, from an absolute path and through a symlink; and both recorded digests match the bytes served for the pinned tag. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
114 lines
5.4 KiB
Bash
114 lines
5.4 KiB
Bash
##############################################################################
|
|
# Copyright (c) 2026 Eclipse ThreadX contributors
|
|
#
|
|
# This program and the accompanying materials are made available under the
|
|
# terms of the MIT License which is available at
|
|
# https://opensource.org/licenses/MIT.
|
|
#
|
|
# AI Disclosure: This file was largely AI-generated by Claude Code (Opus 5).
|
|
# The AI-generated portions may be considered public domain (CC0-1.0)
|
|
# and not subject to the project's licence. The human contributor has
|
|
# reviewed and verified that the code is correct.
|
|
#
|
|
# SPDX-License-Identifier: MIT and CC0-1.0
|
|
##############################################################################
|
|
|
|
# Network helpers shared by the CI install scripts. Sourced, not executed --
|
|
# there is no shebang, and nothing happens here beyond defining APT_OPTIONS,
|
|
# TIMEOUT, TIMEOUT_LONG, retry and fetch.
|
|
#
|
|
# They lived in install.sh until the RISC-V suite was enabled in CI, at which
|
|
# point a second install script was on every pull request's critical path with
|
|
# none of them. The lessons below were paid for once; a script that reaches the
|
|
# network in this project should not have to learn them again.
|
|
#
|
|
# The two callers differ in one way worth knowing: install_riscv.sh runs under
|
|
# set -e and install.sh does not. That is why install.sh spells out `|| exit 1`
|
|
# on the calls that must stop it -- without that, a failed fetch there would be
|
|
# carried forward into a build that then failed somewhere less obvious.
|
|
|
|
# Everything the install scripts do reaches the network, and on this runner pool
|
|
# that is not dependable. apt-get update stalled seven times in a single day,
|
|
# once for more than two hours, each time with the Azure mirror returning
|
|
# nothing and the fallback to archive.ubuntu.com then going silent. Nothing
|
|
# bounded a fetch and nothing retried one, so a mirror being down cost a whole
|
|
# run rather than a few seconds.
|
|
#
|
|
# The Acquire options make a slow mirror give up sooner. The retry loop below
|
|
# covers a mirror that is down rather than merely slow.
|
|
APT_OPTIONS=(-o Acquire::Retries=3
|
|
-o Acquire::http::Timeout=20
|
|
-o Acquire::https::Timeout=20)
|
|
|
|
# Two minutes per attempt, killed outright if it ignores the first signal. Three
|
|
# attempts plus backoff bounds a command at about six and a half minutes.
|
|
#
|
|
# Each command is wrapped in timeout rather than left to bound itself. apt's own
|
|
# Acquire timeouts were tried first and did not help: a run still sat inside a
|
|
# single apt-get update for nine and a half minutes without producing a line,
|
|
# having got as far as fetching noble-security InRelease, so the retry loop never
|
|
# got a turn and the step timeout was what eventually killed it. Whatever apt is
|
|
# waiting on there, it is not something Acquire::http::Timeout covers. timeout
|
|
# does not care where the wait is.
|
|
#
|
|
# timeout goes under sudo, not over it, so that it signals apt itself. Signalling
|
|
# sudo instead risks the kill landing on sudo while apt carries on holding the
|
|
# dpkg lock, which would leave every retry failing for a different reason than
|
|
# the one being retried.
|
|
TIMEOUT=(timeout --kill-after=10 120)
|
|
|
|
# Three minutes, for a single large download rather than a package operation.
|
|
# The two RISC-V toolchain tarballs are about 500 MB each and took 39 seconds
|
|
# apiece on 10 Sep 2026, so this is a margin of roughly four and a half.
|
|
#
|
|
# It is deliberately not larger. The install step in regression_template.yml is
|
|
# capped at ten minutes, and a per-attempt timeout long enough to swallow that
|
|
# cap would leave the retry loop below with no turn to take -- which is the
|
|
# exact failure the TIMEOUT comment above records apt producing. Three attempts
|
|
# at three minutes plus backoff still does not fit inside ten, so the step
|
|
# timeout remains the outer backstop for a server that is genuinely down; what
|
|
# the retries buy is recovery from the transient case, which is the common one,
|
|
# and a log that says which attempt failed rather than a bare cancelled step.
|
|
TIMEOUT_LONG=(timeout --kill-after=10 180)
|
|
|
|
retry() {
|
|
local attempt
|
|
for attempt in 1 2 3; do
|
|
if "$@"; then
|
|
return 0
|
|
fi
|
|
echo "tx_ci_common: '$*' failed or timed out on attempt ${attempt}"
|
|
sleep $((attempt * 10))
|
|
done
|
|
echo "tx_ci_common: '$*' failed after 3 attempts"
|
|
return 1
|
|
}
|
|
|
|
# fetch <url> <destination> <sha256>
|
|
#
|
|
# Downloads with retries and verifies the digest before the caller is allowed to
|
|
# unpack anything. --tries=1 hands retrying to the loop above rather than letting
|
|
# wget retry inside a single timeout window and burn it.
|
|
#
|
|
# The digest is not an independent trust root: it is checked against bytes from
|
|
# the same host that publishes the expected value, so it does not prove the
|
|
# release was not tampered with at source. What it does buy is that the bytes
|
|
# are pinned. A tag can be deleted and re-pushed and an asset can be replaced,
|
|
# and today either would be picked up silently; with this, the build stops and
|
|
# says which file failed and what it got.
|
|
fetch() {
|
|
local url=$1
|
|
local dest=$2
|
|
local sha=$3
|
|
|
|
retry "${TIMEOUT_LONG[@]}" wget --no-verbose --tries=1 "$url" -O "$dest" || return 1
|
|
|
|
if ! echo "${sha} ${dest}" | sha256sum --check --status; then
|
|
echo "tx_ci_common: checksum mismatch for ${url}" >&2
|
|
echo "tx_ci_common: expected ${sha}" >&2
|
|
echo "tx_ci_common: actual $(sha256sum "$dest" | cut -d' ' -f1)" >&2
|
|
rm -f "$dest"
|
|
return 1
|
|
fi
|
|
}
|