The coverage tooling was pinned to gcovr 4.1, released in 2018, and that
version is missing the two options the coverage work needs next: --json and
--add-tracefile, which is how the five build configurations get merged into
one report. This moves the pin to 8.6, the current release. The pin stays
exact, and it stays hand-moved: it lives in a shell script, and no Dependabot
ecosystem can parse that.
Isolated deliberately, so that a movement in the coverage number caused by the
tool could not be confused with one caused by a later change. Measured on the
default_build_coverage tree of test/tx, over the same gcda with the same gcov,
varying only the gcovr version:
gcovr lines-valid branches-valid files
4.1 3827 1994 177
7.0 3827 1994 177
8.3 3827 1994 177
8.6 3827 1994 177
So the denominator does not move with the tool at all, and this bump moves no
number. The plan this came from expected 3822 to become 3827; that figure does
not reproduce, under gcc-13 or gcc-14, with or without --object-directory. The
only variant that changes the count is dropping the -f filter, which collapses
the report to nothing.
Two things found while measuring, both recorded because they matter to what
comes next.
The coverage numerator is not deterministic. On an identical tree with an
identical compiler, three consecutive runs of the full suite -- all 96 tests
passing every time -- reported 3826, 3827 and 3827 covered lines. The line that
flickers is tx_thread_system_resume.c:529, the preemption path of
_tx_thread_system_resume, and it takes its guarding branch with it. It has been
described as never executed; it is executed on some runs and not others. A
coverage floor has to be set with that in mind, and the honest fix is a test
that takes the path deliberately.
Reading gcc-13 output, the compiler the runners actually use, gcovr 8.6 runs
the existing coverage.sh unchanged: Cobertura XML and 181 HTML files, same 177
classes. --xml-pretty and --object-directory still work on 8.6 but are now
deprecated aliases for --cobertura-pretty and --gcov-object-directory, worth
knowing for whoever removes --object-directory next.
Assisted-by: Claude Opus 5 <noreply@anthropic.com>
install.sh reaches the network four times, and on this runner pool that is not
dependable. apt-get update stalled seven times in a single day: once for 55
minutes, once for more than two hours, and five times against the ten minute
step timeout added alongside this. The log says the same thing every time. Every
fetch from azure.archive.ubuntu.com comes back Ign, apt falls back to
archive.ubuntu.com, and then the step produces no further output at all until
something kills it.
Nothing here bounded a fetch and nothing retried one, so a mirror being down
cost a whole run instead of a few seconds.
There is a second problem in the same lines. This script has no set -e, so a
failed apt-get update did not stop the apt-get install that follows. The install
went ahead against whatever package index the image happened to have, and the
run failed later, somewhere with much less to say about why.
Bound each attempt from outside and retry it. apt's own Acquire timeouts were
tried first and are not enough: with them in place a run still sat inside a
single apt-get update for nine and a half minutes without printing a line,
having got as far as fetching noble-security InRelease. The retry loop never got
a turn, because the first attempt never returned, and the step timeout was what
eventually killed it. Whatever apt waits on there is not what
Acquire::http::Timeout covers, so the bound has to come from outside the process.
timeout does not care where the wait is. The Acquire options are kept anyway,
since they make a slow mirror give up sooner, and pip gets its own retry and
timeout flags for the same reason.
timeout goes under sudo rather than over it, so that it signals apt itself.
Signalling sudo risks the kill landing on sudo while apt carries on holding the
dpkg lock, which would leave every retry failing for a different reason than the
one being retried.
The explicit exits stop a failed fetch being carried forward into a build.
set -e is deliberately not used. rm -rf /opt/hostedtoolcache runs without sudo
against a root owned directory and its exit status is not something this script
should start depending on.
The bounds fit inside the ten minute step timeout. Two minutes per attempt,
three attempts, with 10 and 20 second backoffs, caps a command at about six and
a half minutes, and a command that exhausts its attempts exits rather than
letting the next one start.
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Applied the standard MIT license header to all project-owned C, header,
assembly, shell, and Python files that were missing a copyright notice.
Third-party, toolchain startup, and auto-generated files were excluded.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>