Files
threadx/scripts/tx_ci_common.sh
Frédéric Desbiens 6c84e61d19 Gave install_riscv.sh the network hardening install.sh already had, and shared it between them (#721)
install.sh grew a retry loop, per-command timeouts and a deliberately
non-gating apt-get update after this runner pool cost several whole runs: a
mirror going silent for two hours, and a Hash Sum mismatch from a third-party
repository the project does not even use turning builds red. Those lessons were
local to that one file.

install_riscv.sh had none of them, and #717 puts it on every pull request's
critical path. Under set -e its bare apt-get update was a single point of
failure for the whole suite -- the precise case install.sh downgrades to a
warning on purpose -- and its two wget calls, each fetching about 500 MB, had
no retry and no timeout.

Rather than copy the helpers and let them drift again, they move to
tx_ci_common.sh and both scripts source it, following the arrangement
scripts/tx_windows_common.ps1 already uses on the Windows side. install.sh
keeps its behaviour exactly: same APT_OPTIONS, same 120-second TIMEOUT, same
three-attempt retry, and the comments explaining each of them travel with the
code they explain.

Two things are new:

  - TIMEOUT_LONG, 180 seconds, for a single large download. Sized against the
    39 seconds each tarball took on 10 Sep 2026 and deliberately not larger:
    the install step is capped at ten minutes, and a per-attempt timeout able
    to swallow that cap would leave the retry loop no turn to take, which is
    the failure mode the apt comment already records.

  - fetch(), which verifies a SHA-256 before anything is unpacked. Both digests
    were taken from the releases API and then checked against the bytes the CDN
    actually serves. This is not an independent trust root -- expected value and
    file come from the same host -- but it pins the bytes, so a deleted and
    re-pushed tag or a replaced asset stops the build instead of being picked up
    silently.

Also verifies qemu-system-riscv32 alongside riscv64. run.sh selects one per
architecture, so both are worth failing on here rather than at the first test.

Verified locally: retry returns 0 on success and 1 after three attempts;
fetch accepts a correct digest and, on a wrong one, fails and removes the
partial file; the source line resolves from the repository root, from an
absolute path and through a symlink; and both recorded digests match the
bytes served for the pinned tag.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-10 09:28:44 -04:00

114 lines
5.4 KiB
Bash

##############################################################################
# Copyright (c) 2026 Eclipse ThreadX contributors
#
# This program and the accompanying materials are made available under the
# terms of the MIT License which is available at
# https://opensource.org/licenses/MIT.
#
# AI Disclosure: This file was largely AI-generated by Claude Code (Opus 5).
# The AI-generated portions may be considered public domain (CC0-1.0)
# and not subject to the project's licence. The human contributor has
# reviewed and verified that the code is correct.
#
# SPDX-License-Identifier: MIT and CC0-1.0
##############################################################################
# Network helpers shared by the CI install scripts. Sourced, not executed --
# there is no shebang, and nothing happens here beyond defining APT_OPTIONS,
# TIMEOUT, TIMEOUT_LONG, retry and fetch.
#
# They lived in install.sh until the RISC-V suite was enabled in CI, at which
# point a second install script was on every pull request's critical path with
# none of them. The lessons below were paid for once; a script that reaches the
# network in this project should not have to learn them again.
#
# The two callers differ in one way worth knowing: install_riscv.sh runs under
# set -e and install.sh does not. That is why install.sh spells out `|| exit 1`
# on the calls that must stop it -- without that, a failed fetch there would be
# carried forward into a build that then failed somewhere less obvious.
# Everything the install scripts do reaches the network, and on this runner pool
# that is not dependable. apt-get update stalled seven times in a single day,
# once for more than two hours, each time with the Azure mirror returning
# nothing and the fallback to archive.ubuntu.com then going silent. Nothing
# bounded a fetch and nothing retried one, so a mirror being down cost a whole
# run rather than a few seconds.
#
# The Acquire options make a slow mirror give up sooner. The retry loop below
# covers a mirror that is down rather than merely slow.
APT_OPTIONS=(-o Acquire::Retries=3
-o Acquire::http::Timeout=20
-o Acquire::https::Timeout=20)
# Two minutes per attempt, killed outright if it ignores the first signal. Three
# attempts plus backoff bounds a command at about six and a half minutes.
#
# Each command is wrapped in timeout rather than left to bound itself. apt's own
# Acquire timeouts were tried first and did not help: a run still sat inside a
# single apt-get update for nine and a half minutes without producing a line,
# having got as far as fetching noble-security InRelease, so the retry loop never
# got a turn and the step timeout was what eventually killed it. Whatever apt is
# waiting on there, it is not something Acquire::http::Timeout covers. timeout
# does not care where the wait is.
#
# timeout goes under sudo, not over it, so that it signals apt itself. Signalling
# sudo instead risks the kill landing on sudo while apt carries on holding the
# dpkg lock, which would leave every retry failing for a different reason than
# the one being retried.
TIMEOUT=(timeout --kill-after=10 120)
# Three minutes, for a single large download rather than a package operation.
# The two RISC-V toolchain tarballs are about 500 MB each and took 39 seconds
# apiece on 10 Sep 2026, so this is a margin of roughly four and a half.
#
# It is deliberately not larger. The install step in regression_template.yml is
# capped at ten minutes, and a per-attempt timeout long enough to swallow that
# cap would leave the retry loop below with no turn to take -- which is the
# exact failure the TIMEOUT comment above records apt producing. Three attempts
# at three minutes plus backoff still does not fit inside ten, so the step
# timeout remains the outer backstop for a server that is genuinely down; what
# the retries buy is recovery from the transient case, which is the common one,
# and a log that says which attempt failed rather than a bare cancelled step.
TIMEOUT_LONG=(timeout --kill-after=10 180)
retry() {
local attempt
for attempt in 1 2 3; do
if "$@"; then
return 0
fi
echo "tx_ci_common: '$*' failed or timed out on attempt ${attempt}"
sleep $((attempt * 10))
done
echo "tx_ci_common: '$*' failed after 3 attempts"
return 1
}
# fetch <url> <destination> <sha256>
#
# Downloads with retries and verifies the digest before the caller is allowed to
# unpack anything. --tries=1 hands retrying to the loop above rather than letting
# wget retry inside a single timeout window and burn it.
#
# The digest is not an independent trust root: it is checked against bytes from
# the same host that publishes the expected value, so it does not prove the
# release was not tampered with at source. What it does buy is that the bytes
# are pinned. A tag can be deleted and re-pushed and an asset can be replaced,
# and today either would be picked up silently; with this, the build stops and
# says which file failed and what it got.
fetch() {
local url=$1
local dest=$2
local sha=$3
retry "${TIMEOUT_LONG[@]}" wget --no-verbose --tries=1 "$url" -O "$dest" || return 1
if ! echo "${sha} ${dest}" | sha256sum --check --status; then
echo "tx_ci_common: checksum mismatch for ${url}" >&2
echo "tx_ci_common: expected ${sha}" >&2
echo "tx_ci_common: actual $(sha256sum "$dest" | cut -d' ' -f1)" >&2
rm -f "$dest"
return 1
fi
}