Gave install_riscv.sh the network hardening install.sh already had, and shared it between them (#721)

install.sh grew a retry loop, per-command timeouts and a deliberately
non-gating apt-get update after this runner pool cost several whole runs: a
mirror going silent for two hours, and a Hash Sum mismatch from a third-party
repository the project does not even use turning builds red. Those lessons were
local to that one file.

install_riscv.sh had none of them, and #717 puts it on every pull request's
critical path. Under set -e its bare apt-get update was a single point of
failure for the whole suite -- the precise case install.sh downgrades to a
warning on purpose -- and its two wget calls, each fetching about 500 MB, had
no retry and no timeout.

Rather than copy the helpers and let them drift again, they move to
tx_ci_common.sh and both scripts source it, following the arrangement
scripts/tx_windows_common.ps1 already uses on the Windows side. install.sh
keeps its behaviour exactly: same APT_OPTIONS, same 120-second TIMEOUT, same
three-attempt retry, and the comments explaining each of them travel with the
code they explain.

Two things are new:

  - TIMEOUT_LONG, 180 seconds, for a single large download. Sized against the
    39 seconds each tarball took on 10 Sep 2026 and deliberately not larger:
    the install step is capped at ten minutes, and a per-attempt timeout able
    to swallow that cap would leave the retry loop no turn to take, which is
    the failure mode the apt comment already records.

  - fetch(), which verifies a SHA-256 before anything is unpacked. Both digests
    were taken from the releases API and then checked against the bytes the CDN
    actually serves. This is not an independent trust root -- expected value and
    file come from the same host -- but it pins the bytes, so a deleted and
    re-pushed tag or a replaced asset stops the build instead of being picked up
    silently.

Also verifies qemu-system-riscv32 alongside riscv64. run.sh selects one per
architecture, so both are worth failing on here rather than at the first test.

Verified locally: retry returns 0 on success and 1 after three attempts;
fetch accepts a correct digest and, on a wrong one, fails and removes the
partial file; the source line resolves from the repository root, from an
absolute path and through a symlink; and both recorded digests match the
bytes served for the pinned tag.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
This commit is contained in:
Frédéric Desbiens
2026-09-10 09:28:44 -04:00
committed by GitHub
parent 18abe103af
commit 6c84e61d19
3 changed files with 158 additions and 51 deletions
+6 -47
View File
@@ -16,53 +16,12 @@
# Remove large folder to save space
rm -rf /opt/hostedtoolcache
# Everything below reaches the network, and on this runner pool that is not
# dependable. apt-get update stalled seven times in a single day, once for more
# than two hours, each time with the Azure mirror returning nothing and the
# fallback to archive.ubuntu.com then going silent. Nothing here bounded a fetch
# and nothing retried one, so a mirror being down cost a whole run rather than a
# few seconds. Worse, this script has no set -e, so a failed update did not stop
# the install that follows: it went on to install from whatever index it already
# had, and the run failed later somewhere less obvious.
#
# Each command is wrapped in timeout rather than left to bound itself. apt's own
# Acquire timeouts were tried first and did not help: a run still sat inside a
# single apt-get update for nine and a half minutes without producing a line,
# having got as far as fetching noble-security InRelease, so the retry loop never
# got a turn and the step timeout was what eventually killed it. Whatever apt is
# waiting on there, it is not something Acquire::http::Timeout covers. timeout
# does not care where the wait is.
#
# The Acquire options are kept anyway, since they make a slow mirror give up
# sooner. The loop covers a mirror that is down rather than merely slow. The
# explicit exits stop a failed fetch from being carried forward into a build,
# with one deliberate exception noted at the update below.
APT_OPTIONS=(-o Acquire::Retries=3
-o Acquire::http::Timeout=20
-o Acquire::https::Timeout=20)
# Two minutes per attempt, killed outright if it ignores the first signal. Three
# attempts plus backoff bounds a command at about six and a half minutes, and a
# command that exhausts its attempts exits rather than letting the next one run.
#
# timeout goes under sudo, not over it, so that it signals apt itself. Signalling
# sudo instead risks the kill landing on sudo while apt carries on holding the
# dpkg lock, which would leave every retry failing for a different reason than
# the one being retried.
TIMEOUT=(timeout --kill-after=10 120)
retry() {
local attempt
for attempt in 1 2 3; do
if "$@"; then
return 0
fi
echo "install.sh: '$*' failed or timed out on attempt ${attempt}"
sleep $((attempt * 10))
done
echo "install.sh: '$*' failed after 3 attempts"
return 1
}
# The network helpers -- retry, TIMEOUT, TIMEOUT_LONG and APT_OPTIONS -- live in
# tx_ci_common.sh, alongside the comments recording why each of them is shaped
# the way it is. They were defined here until the RISC-V suite was enabled in
# CI, which put a second install script on every pull request's critical path
# with none of them.
. "$(dirname "$(realpath "$0")")/tx_ci_common.sh"
# THE UPDATE IS NOT THE GATE, AND IT MUST NOT BE. apt-get update fails if ANY
# configured repository serves a bad index, including ones this project does
+39 -4
View File
@@ -13,23 +13,55 @@
# Install RISC-V bare-metal cross-compiler toolchain and QEMU for CI.
set -e
# retry, TIMEOUT, TIMEOUT_LONG, APT_OPTIONS and fetch. Shared with install.sh,
# which is where they were written and where the reasons for their shape are
# recorded.
. "$(dirname "$(realpath "$0")")/tx_ci_common.sh"
# The release tag is pinned so that what CI compiles with is answerable from the
# repository, and the digests pin the bytes behind the tag. A tag can be deleted
# and re-pushed and a release asset can be replaced; either would be picked up
# silently without these. Both digests were taken from the releases API and then
# checked against the bytes the CDN actually serves, 10 Sep 2026. Moving the tag
# means refreshing all three lines together.
RELEASE_TAG="2026.04.26"
BASE_URL="https://github.com/riscv-collab/riscv-gnu-toolchain/releases/download/${RELEASE_TAG}"
# Use ubuntu-24.04 binaries to match ubuntu-latest runners.
RV32_TARBALL="riscv32-elf-ubuntu-24.04-gcc.tar.xz"
RV32_SHA256="73c9a5adbb38b779312e5b3fc10d624484364ef1f870314571ae92c486ba9917"
RV64_TARBALL="riscv64-elf-ubuntu-24.04-gcc.tar.xz"
RV64_SHA256="4a66a329653c9cfb869b826cae6d70c603bedcc1e13d419d3e3e92b99e0816b1"
echo "=== Installing QEMU and build tools ==="
sudo apt-get update -qq
sudo apt-get install -y -qq qemu-system-misc ninja-build cmake
# The update warns rather than gates, for the reason install.sh sets out at
# length: apt-get update fails if ANY configured repository serves a bad index,
# including the third-party ones the runner image carries and this project does
# not use. The install below is the gate, and it still fails hard on a package
# it cannot get. Under set -e a bare update would have ended the job -- and with
# the RISC-V suite running on every pull request, ended it for everybody.
if ! retry sudo "${TIMEOUT[@]}" apt-get "${APT_OPTIONS[@]}" update; then
echo ""
echo "install_riscv.sh: apt-get update did not fully succeed."
echo "install_riscv.sh: continuing, because a repository this project does"
echo "install_riscv.sh: not use can fail an update. The install below is"
echo "install_riscv.sh: the real gate."
echo ""
fi
retry sudo "${TIMEOUT[@]}" apt-get "${APT_OPTIONS[@]}" install -y \
qemu-system-misc \
ninja-build \
cmake || exit 1
echo "=== Downloading RISC-V GCC toolchain (${RELEASE_TAG}) ==="
# Both tarballs extract into riscv/ with non-overlapping prefixes
# (riscv32-unknown-elf-* and riscv64-unknown-elf-*).
for tarball in "$RV32_TARBALL" "$RV64_TARBALL"; do
for spec in "${RV32_TARBALL}:${RV32_SHA256}" "${RV64_TARBALL}:${RV64_SHA256}"; do
tarball="${spec%%:*}"
sha256="${spec##*:}"
echo "Downloading ${tarball} ..."
wget --no-verbose "${BASE_URL}/${tarball}" -O "/tmp/${tarball}"
fetch "${BASE_URL}/${tarball}" "/tmp/${tarball}" "${sha256}" || exit 1
sudo tar xJf "/tmp/${tarball}" -C /opt
rm "/tmp/${tarball}"
done
@@ -40,4 +72,7 @@ echo "$TOOLCHAIN_BIN" >> "$GITHUB_PATH"
echo "=== Verifying installation ==="
"$TOOLCHAIN_BIN/riscv32-unknown-elf-gcc" --version | head -1
"$TOOLCHAIN_BIN/riscv64-unknown-elf-gcc" --version | head -1
# Both QEMU system emulators are used -- run.sh selects one per architecture --
# so both are worth failing on here rather than at the first test.
qemu-system-riscv32 --version | head -1
qemu-system-riscv64 --version | head -1
+113
View File
@@ -0,0 +1,113 @@
##############################################################################
# Copyright (c) 2026 Eclipse ThreadX contributors
#
# This program and the accompanying materials are made available under the
# terms of the MIT License which is available at
# https://opensource.org/licenses/MIT.
#
# AI Disclosure: This file was largely AI-generated by Claude Code (Opus 5).
# The AI-generated portions may be considered public domain (CC0-1.0)
# and not subject to the project's licence. The human contributor has
# reviewed and verified that the code is correct.
#
# SPDX-License-Identifier: MIT and CC0-1.0
##############################################################################
# Network helpers shared by the CI install scripts. Sourced, not executed --
# there is no shebang, and nothing happens here beyond defining APT_OPTIONS,
# TIMEOUT, TIMEOUT_LONG, retry and fetch.
#
# They lived in install.sh until the RISC-V suite was enabled in CI, at which
# point a second install script was on every pull request's critical path with
# none of them. The lessons below were paid for once; a script that reaches the
# network in this project should not have to learn them again.
#
# The two callers differ in one way worth knowing: install_riscv.sh runs under
# set -e and install.sh does not. That is why install.sh spells out `|| exit 1`
# on the calls that must stop it -- without that, a failed fetch there would be
# carried forward into a build that then failed somewhere less obvious.
# Everything the install scripts do reaches the network, and on this runner pool
# that is not dependable. apt-get update stalled seven times in a single day,
# once for more than two hours, each time with the Azure mirror returning
# nothing and the fallback to archive.ubuntu.com then going silent. Nothing
# bounded a fetch and nothing retried one, so a mirror being down cost a whole
# run rather than a few seconds.
#
# The Acquire options make a slow mirror give up sooner. The retry loop below
# covers a mirror that is down rather than merely slow.
APT_OPTIONS=(-o Acquire::Retries=3
-o Acquire::http::Timeout=20
-o Acquire::https::Timeout=20)
# Two minutes per attempt, killed outright if it ignores the first signal. Three
# attempts plus backoff bounds a command at about six and a half minutes.
#
# Each command is wrapped in timeout rather than left to bound itself. apt's own
# Acquire timeouts were tried first and did not help: a run still sat inside a
# single apt-get update for nine and a half minutes without producing a line,
# having got as far as fetching noble-security InRelease, so the retry loop never
# got a turn and the step timeout was what eventually killed it. Whatever apt is
# waiting on there, it is not something Acquire::http::Timeout covers. timeout
# does not care where the wait is.
#
# timeout goes under sudo, not over it, so that it signals apt itself. Signalling
# sudo instead risks the kill landing on sudo while apt carries on holding the
# dpkg lock, which would leave every retry failing for a different reason than
# the one being retried.
TIMEOUT=(timeout --kill-after=10 120)
# Three minutes, for a single large download rather than a package operation.
# The two RISC-V toolchain tarballs are about 500 MB each and took 39 seconds
# apiece on 10 Sep 2026, so this is a margin of roughly four and a half.
#
# It is deliberately not larger. The install step in regression_template.yml is
# capped at ten minutes, and a per-attempt timeout long enough to swallow that
# cap would leave the retry loop below with no turn to take -- which is the
# exact failure the TIMEOUT comment above records apt producing. Three attempts
# at three minutes plus backoff still does not fit inside ten, so the step
# timeout remains the outer backstop for a server that is genuinely down; what
# the retries buy is recovery from the transient case, which is the common one,
# and a log that says which attempt failed rather than a bare cancelled step.
TIMEOUT_LONG=(timeout --kill-after=10 180)
retry() {
local attempt
for attempt in 1 2 3; do
if "$@"; then
return 0
fi
echo "tx_ci_common: '$*' failed or timed out on attempt ${attempt}"
sleep $((attempt * 10))
done
echo "tx_ci_common: '$*' failed after 3 attempts"
return 1
}
# fetch <url> <destination> <sha256>
#
# Downloads with retries and verifies the digest before the caller is allowed to
# unpack anything. --tries=1 hands retrying to the loop above rather than letting
# wget retry inside a single timeout window and burn it.
#
# The digest is not an independent trust root: it is checked against bytes from
# the same host that publishes the expected value, so it does not prove the
# release was not tampered with at source. What it does buy is that the bytes
# are pinned. A tag can be deleted and re-pushed and an asset can be replaced,
# and today either would be picked up silently; with this, the build stops and
# says which file failed and what it got.
fetch() {
local url=$1
local dest=$2
local sha=$3
retry "${TIMEOUT_LONG[@]}" wget --no-verbose --tries=1 "$url" -O "$dest" || return 1
if ! echo "${sha} ${dest}" | sha256sum --check --status; then
echo "tx_ci_common: checksum mismatch for ${url}" >&2
echo "tx_ci_common: expected ${sha}" >&2
echo "tx_ci_common: actual $(sha256sum "$dest" | cut -d' ' -f1)" >&2
rm -f "$dest"
return 1
fi
}