Commit Graph
67 Commits
Author SHA1 Message Date
Frédéric Desbiens 147754cc86 Enforced a coverage floor on the merged report (#667)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The coverage summary reported a percentage and could not fail. Coverage
could fall from 99.97% to anything at all and every check stayed green,
against an AGENTS.md that asks for 100% test coverage -- a stated
requirement measured with a gauge that had no failure mode.

CodeCoverageSummary already takes thresholds and fail_below_min; neither
was set. Both are now, through a new coverage_thresholds input on the
template, because the two suites do not sit at the same figure: ThreadX
99, SMP 98.

Three things were probed against the pinned action on a runner before
picking those numbers, using the real merged.xml files from the dev push
run of #666.

The floor compares the line rate and nothing else. That mattered because
branch coverage is around 78% in both suites while line coverage is
98.8-100%, so a floor aimed at the line figure would have been an
immediate red wall had it tested branches or the lower of the two. The
ThreadX report at 100.00% lines and 77.67% branches clears a floor of 99.

The thresholds are whole numbers. '99.9 100' -- the value this was meant
to be -- is rejected with 'System.ArgumentException - Threshold parameter
set incorrectly.', and the step fails whether or not fail_below_min is
set. So the choice is 99 or 100 with nothing between.

100 would fail on a race. tx_thread_system_resume.c:529 is reached by
timing rather than by construction and flaps between runs of the same
green tree, which is why #666 left it; 4502/4503 fails a floor of 100 and
clears one of 99. A coverage gate that goes red on a coin toss is how
coverage gates get switched off.

SMP is 5114/5178 lines, 98.76%, with 64 uncovered lines across 11 files
of common_smp/src -- #666 closed the equivalent gaps in common/src only.
A shared floor of 99 would have failed that job on every run while
ThreadX passed.

One limit is recorded in the file rather than fixed: an empty report
reads as 100%. gcovr writes line-rate="1.0" beside lines-valid="0" when
it finds no data, and the action prints 'Line Rate = 100% (0 / 0)' and
passes any floor. The check for that is the emptiness assertion #664 put
in each suite's coverage.sh, not this one.

Also corrected two stale filenames in the deploy job's comment: since
#665 each coverage artifact carries merged.xml, not
default_build_coverage.xml. Verified on the runner.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 09:13:35 -04:00
Frédéric Desbiens 3e85bbd431 Instrumented every build configuration and merged their coverage (#665)
Only default_build_coverage carried -fprofile-arcs, because the gate was the
build type and it is the only one of five whose name ends in _coverage. The
other four build and run all their tests and their coverage was discarded. That
is not redundancy thrown away: each configuration selects a different set of TX_
feature macros, so the code the other four compile is absent from the
denominator rather than uncovered in it.

TX_COVERAGE instruments a build regardless of its name, defaulting to OFF so a
single configuration built by hand behaves as before. coverage.sh gains a
--merge mode that unions the per-configuration JSON tracefiles, and
cmake_bootstrap.sh runs it after the test loop so a local run produces the same
merged report CI reads. The template sets TX_COVERAGE for build and test, and
coverage_name moves to the merged report.

Measured on the ThreadX suite, all 480 tests passing:

  default_build_coverage           3827 valid   3827 covered
  disable_notify_callbacks_build   3767         3766
  stack_checking_build             3857         3856
  stack_checking_rand_fill_build   3862         3861
  trace_build                      4123         4108
  merged                           4503         4487

The denominator grows by 676 lines, 17.7%, and the figure moves from 99.97% to
99.64%. The second one is honest, and the drop is the point rather than a
regression: the denominator now includes code the old report never counted. The
union also contains a file the old report did not contain at all --
tx_thread_stack_error_handler.c compiles only under TX_ENABLE_STACK_CHECKING, so
it was not listed at 0%, it was simply absent. 177 files becomes 178.

Coverage collection moved out of test() and now runs after the test loop, one
configuration at a time. gcov writes its intermediate gcov files into the
directory gcovr is rooted at, and coverage.sh roots every configuration at the
repository root so filenames come out repo-relative. Five concurrent gcovr
processes therefore share one scratch directory and delete each other's output:
the first full run of this change passed all 480 tests and produced no report
for three of the five configurations. Measured both ways -- two gcovr rooted at
the repository root fail concurrently and succeed in sequence. CI would not
have caught it, because test_tx.sh sets CTEST_PARALLEL_LEVEL=1 and takes the
serial branch.

Per-configuration output moved under coverage_report/per_configuration/ and is
excluded from the Pages artifact. The deploy job merges the ThreadX and SMP
artifacts into one tree and every configuration directory has the same name in
both, so left at the top level one suite's would overwrite the other's on the
published site.

On the SMP suite, an earlier run of this change saw trace_build fail
threadx_smp_time_slice_test and then hang, which raised the question of whether
-fprofile-arcs perturbs a timing-sensitive test. It does not. Sixteen runs
settle it, and the decisive one is that threadx_smp_time_slice_test failed
ERROR #31 -- twice in a row under --repeat until-pass:2 -- on an uninstrumented
build, in the exact shape CI runs, while three instrumented runs of that shape
passed 5 of 5. In the CI shape, CTEST_PARALLEL_LEVEL=1 run.sh test all:

  TX_COVERAGE=OFF   3 runs   2 green, one ERROR #31        310 s
  TX_COVERAGE=ON    3 runs   3 green, 5/5 each             325-329 s

So the test is a pre-existing flake on dev and instrumenting all five costs
about 5% of the suite's wall clock. Separately, and also in both instrumented
and uninstrumented builds, run.sh's parallel branch -- what a developer gets
typing run.sh test all with no CTEST_PARALLEL_LEVEL -- hangs under its own load,
four times in twelve runs. Several SMP tests create 1024 ThreadX threads by
construction and the Linux port backs each with a pthread, so five
configurations at once put on the order of 5000 threads on the machine. CI sets
CTEST_PARALLEL_LEVEL=1 and does not take that branch.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-26 08:13:33 -04:00
Frédéric Desbiens b6a00a2014 Added the Dependabot configuration the pinned actions need (#662)
The action references were pinned to commit SHAs in #660, and a SHA pin with
nothing moving it is worse than a floating tag -- it holds CI on whatever was
current the day it was written. That is exactly how actions/cache@v1 stayed in
ci_cortex_m.yml until GitHub began auto-failing every request that used it.
The drift measured before that catch-up: download-artifact four majors behind,
checkout and upload-artifact three each, cache and upload-pages-artifact two,
with nothing ever reporting it. This closes the loop, and the reference to
.github/dependabot.yml that #660 left in each workflow's pinning comment.

Weekly, github-actions only. Patch and minor are grouped into one pull request
because they are the routine traffic and a queue reviewed one item at a time is
a queue that gets ignored. Majors stay ungrouped, one each, because every
breaking change this repository has met in an action has been a major.

Two choices worth stating rather than leaving to be rediscovered.

target-branch is dev. Dependabot reads this file from the default branch, which
is master, but master is deliberately kept behind dev and pull requests belong
where the regression suites gate them. The consequence is that landing this on
dev arms it without firing it: nothing happens until a release merge carries
the file to master. Setting target-branch also opts out of Dependabot security
updates, which only run against the default branch -- a small cost for this
ecosystem, since an action advisory arrives as an ordinary bump on the weekly
run, but a real one.

The pull-request limit is raised from the default five to ten. Nine actions are
in use, and five would hold majors back with nothing saying that it had.

No other ecosystem is configured, deliberately: external dependencies are
forbidden, there are no submodules, and the one pinned tool -- gcovr in
scripts/install.sh -- lives in a shell script no ecosystem can parse, so that
pin keeps moving by hand.

No sibling eclipse-threadx repository has a Dependabot configuration, so this
sets the pattern rather than following one. The dependencies label it uses
already exists here.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 16:05:04 -04:00
Frédéric Desbiens 3d852eb451 Pinned every action to a commit SHA, and moved them off Node 20 (#660)
Node 20 is removed from the GitHub runners on 16 September 2026. Every run
in this repository currently emits the deprecation warning for it, naming
actions/checkout, actions/configure-pages, actions/upload-artifact,
LouisBrunner/checks-action and marocchino/sticky-pull-request-comment among
others. After that date those actions stop working rather than warning, so
this is a deadline and not housekeeping.

Every action is now referenced by a 40-character commit SHA with the version
in a trailing comment. A tag can be repointed at any commit; a SHA cannot, so
this is what makes "which code ran in our CI" answerable from the repository
rather than from whatever the tag meant at the time. The versions were behind
by as much as four majors -- download-artifact was on v4.3.0 against v8.0.1 --
because nothing in this repository has ever reported that an action moved.

Compatibility was checked against each new action.yml rather than assumed,
for every input this repository actually passes:

  checkout            submodules is unchanged
  cache               path and key are unchanged
  upload-artifact     name, path and retention-days are unchanged
  download-artifact   pattern, merge-multiple and path are unchanged
  configure-pages     takes no input here, and none became required
  deploy-pages        still exposes page_url, which the job reads
  upload-pages-art.   path is unchanged
  checks-action       token, name, conclusion, output and
                      output_text_description_file all survive v2 to v3
  sticky-comment      header and path survive v2 to v3, and the new
                      GITHUB_TOKEN input defaults to github.token, which is
                      what v2 used implicitly
  delete-artifact     name survives v5 to v6, and useGlob still defaults to
                      true, so the coverage_report-* glob from #655 still
                      matches
  CodeCoverageSummary already current at v1.3.0; pinned, not moved

The artifact pair moves together, as it must. The round trip was verified on
a runner before this commit: upload-artifact v7 to download-artifact v8,
through the pattern and merge-multiple selection #655 introduced, filtered 4
artifacts to 2 and produced exactly the tree the deploy expects.

Two behaviour changes worth knowing. download-artifact v8 adds a
digest-mismatch input defaulting to error, so a corrupted artifact now fails
the job instead of passing through -- the right default, but a change.
upload-artifact v6 and above require a runner of at least 2.327.1, which the
hosted runners satisfy and a self-hosted runner would need checking for.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-25 15:52:49 -04:00
Frédéric Desbiens 042049a7b2 Revived the Cortex-M build, which had compiled nothing since June (#653)
cortex_m / Cortex M0 build (push) Canceled after 0s
cortex_m / Cortex M3 build (push) Canceled after 0s
cortex_m / Cortex M4 build (push) Canceled after 0s
cortex_m / Cortex M7 build (push) Canceled after 0s
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
Two defects, and the second hid the first.

This workflow triggered on master only, for both push and pull_request,
while dev is the integration branch. So it gated no pull request that
anybody opened. That is the same defect ports_arch_check.yml carries a
comment about, where it cost eight months of ports drifting from ports_arch
unnoticed, and regression_test.yml has it too.

And it had not compiled anything since at least 2026-06-08. Every run since
then failed in six to eight seconds at "Prepare all required actions",
before checkout, because GitHub automatically fails any request that uses
actions/cache@v1. The last run of any kind was 2026-06-30. A workflow that
fails in seven seconds is normally noticed within the day; this one was not,
because of the first defect. The two together meant the project's only job
that cross-compiles a port with GCC had been reporting nothing at all.

The toolchain now follows clang_check.yml rather than third party actions:
a pinned release fetched directly from Arm, verified against the published
sha256asc, and cached with actions/cache@v4. That also moves the compiler
off 9-2019-q4, a 2019 release, onto a version matching the GCC 14 default
this project states.

The ninja install is guarded on ninja being absent rather than run
unconditionally, because scripts/install.sh already carries a long comment
about apt-get update stalling for over two hours and taking a whole
regression run with it.

fail-fast is off so that one port failing still reports the other three.

Verified before committing, with the pinned 14.3.rel1 toolchain: the
download and sha256sum -c sequence in the install step was run as written,
the archive extracts to the directory the PATH step expects, and all four
ports configure and build clean with zero warnings.

This covers four ports of the forty under ports/ that have a gnu directory.
Widening it to every Arm gnu port is separate work.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:38:24 -04:00
Frédéric Desbiens 7959aef3bc Kept the coverage report from the runs that most need one (#659)
A failing test threw away coverage that had already been collected, and the
run whose behaviour changed is exactly the run whose coverage is worth
reading. Measured on the failing run of 2026-08-18: it uploaded test_reports
for all three suites and no coverage_report artifact at all.

Two causes, and the workflow one is the smaller of them.

cmake_bootstrap.sh runs under set -e, so a failing ctest aborted test()
before ./coverage.sh was reached. The gcda files exist by that point, so
nothing was missing except the step that reads them. ctest's status is now
captured and returned at the end, and the summary grep is allowed to fail
rather than being the thing that stops the coverage behind it.

The serial branch of the test dispatch collected no status either, so under
set -e the first failing configuration stopped the remaining four from being
tested at all -- and their coverage from being collected. That was cheap
while the suites ran in parallel, because the parallel branch already
collects exit codes from its background jobs. Moving to serial execution in
#643 quietly made one failure cost the other four configurations. The serial
branch now collects status the same way the parallel branch does.

With those fixed the report exists, so the workflow steps that publish it no
longer skip on failure. They are guarded with !cancelled() rather than
always(), so a cancelled run still stops promptly, which is the idiom
deploy_code_coverage already uses. The ${{ }} wrapping is required and not
decoration: a bare ! opens a YAML tag, and the file will not parse without
it.

Verified locally by replacing one test binary with a stub that exits 1:

  before  the failing run of 2026-08-18 produced no coverage_report artifact
  after   run.sh test default_build_coverage exits 8, and produces
          coverage_report/default_build_coverage.xml with 177 files and
          3804 of 3827 lines
  after   run.sh test all exits 8, and all five configurations run rather
          than stopping at the first

The failure still fails. Only the reporting around it changed.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:36:18 -04:00
Frédéric Desbiens adc6469b91 Published the coverage report instead of everything the run produced (#655)
The download step in deploy_code_coverage asked for the artifact named
${{ steps.artifact.outputs.coverage_report }}. That output is set by the
"Coverage Report name" step of run_tests, which is a different job, and the
steps context does not cross jobs. So the expression evaluated to the empty
string and the action took its documented path for an unspecified name:

    No input name, artifact-ids or pattern filtered specified,
    downloading all artifacts
    Total of 4 artifact(s) downloaded

The four are the two coverage reports and the two test_reports bundles of
JUnit XML, each extracted into a directory named after the artifact. The
next step uploads the lot to Pages, so the published site has carried the
test reports alongside the coverage, one directory deeper than intended,
under a path containing a run timestamp that changed on every publish. Any
link to a coverage report broke the next time one was published.

Selecting by pattern with merge-multiple fixes both halves: the pattern
excludes the test_reports bundles, and merging puts the contents of the two
coverage artifacts directly into coverage_report rather than under a
directory named for each. Each artifact holds one directory named for its
suite, renamed from default_build_coverage by "Prepare Coverage GitHub
Pages", so the result is the two suite directories the deploy expects and
the timestamped artifact name no longer appears in the published path.

Verified on a runner rather than reasoned about, with an isolated workflow
that uploads artifacts shaped like the real ones and downloads them both
ways:

    OLD  coverage_report/coverage_report-<epoch>-ThreadX/ThreadX/index.html
         coverage_report/coverage_report-<epoch>-ThreadX/default_build_coverage.xml
         coverage_report/coverage_report-<epoch>-SMP/SMP/index.html
         coverage_report/coverage_report-<epoch>-SMP/default_build_coverage.xml
         coverage_report/test_reports SMP/results.xml
         coverage_report/test_reports ThreadX/results.xml

    NEW  coverage_report/ThreadX/index.html
         coverage_report/SMP/index.html
         coverage_report/default_build_coverage.xml

Both artifacts carry a default_build_coverage.xml and the merge means one
overwrites the other, which the run above also shows. That file is consumed
by CodeCoverageSummary back in run_tests and is not read here, so it is
untidy rather than wrong, and it is called out in a comment.

The delete step is fixed in the same place and for a related reason. The
artifacts are named coverage_report-<epoch>, useGlob defaults to true in
this action, and as a glob "coverage_report" matches only the literal
string. It has been deleting nothing, without failing, and retention-days: 1
on the upload is what has actually been clearing these up.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 16:31:21 -04:00
Frédéric Desbiens 9218bad4bc Kept the coverage publish on master, where the environment allows it (#654)
Running the regression suites on dev (#652) was meant to test the branch the
pull requests target. It changed what gets published as well, which was not
intended and does not work: the first push to dev after that merge failed
with

    Branch "dev" is not allowed to deploy to github-pages due to
    environment protection rules.

All three suites passed in that run -- tx, smp and freertos. The only
failure was deploy / deploy_code_coverage, rejected before it ran, because
the github-pages environment restricts deployments to master.

The guard goes here rather than in the environment settings, because the
environment rule is doing its job. Which branch the published coverage
report describes is a deliberate decision, and moving it from master to dev
is a change worth making on purpose rather than as a side effect of a
trigger fix. Doing so needs the environment setting relaxed as well as this
line removed.

The per-suite deploy_code_coverage jobs need no guard: tx, smp and freertos
all pass skip_deploy: true, and regression_template.yml already restricts
that job to push and workflow_dispatch. Only the deploy job, which is the
one that publishes, was reaching the environment.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 13:07:38 -04:00
Frédéric Desbiens 977e14e776 Ran the regression suites on dev, where the pull requests actually are (#652)
regression_test / tx (push) Canceled after 0s
regression_test / smp (push) Canceled after 0s
regression_test / freertos (push) Canceled after 0s
regression_test / deploy (push) Canceled after 0s
regression_template / run_tests (push) Canceled after 0s
regression_template / deploy_code_coverage (push) Canceled after 0s
The ThreadX, SMP and FreeRTOS-compatibility suites triggered on master only,
for both push and pull_request. dev is the integration branch, so these
suites gated no pull request that anybody opened: the last dev run of any
kind was a manual workflow_dispatch on 2026-08-18.

This is the same defect ports_arch_check.yml already carries a comment
about, where it cost eight months of ports drifting from ports_arch
unnoticed. ci_cortex_m.yml has it too and is handled separately.

That 2026-08-18 run failed, which is the reason to check before switching
this on rather than after. Two tests failed: threadx_timer_simple_test in
the ThreadX suite, with ERROR #28, and threadx_thread_priority_change in
the SMP suite, with a timeout. Both were fixed two days later -- the first
by running the suites one test at a time (#643), which is what a timer test
failing only under parallel load wants, and the second by #647 by name.

Verified before this commit rather than assumed: both suites were re-run on
this tree, and all 1030 tests pass across all ten build configurations, the
ThreadX suite in 34 to 64 seconds per configuration and the SMP suite in 61
to 63. The 2026-08-18 run took 36m19s, of which a single test that has since
been given a budget accounted for 439 seconds.

No paths filter is added deliberately. The suites build the linux port, so a
filter would have to enumerate what cannot affect them, and the failure mode
of getting that list wrong is a regression that merges because the filter
excluded the file that caused it.

The deploy job needs no guard: regression_template.yml already restricts
deploy_code_coverage to push and workflow_dispatch, and restricts the
coverage PR comment to pull requests from the repository itself, so neither
fires for a pull request from a fork.

Assisted-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 11:30:19 -04:00
Frédéric DesbiensandClaude Opus 5 dafb7d70cc Stopped a stalled install from costing a whole regression run (#641)
The Install softwares step runs apt-get update, an apt-get install and a pip
install, and it has stalled twice: 55 minutes on the SMP job of one run, and
more than two hours on the ThreadX job of the next, against a normal 27 to 152
seconds across every other run measured. In the second case the tests never
started at all.

No step in this template had a timeout, so a stall runs until the six hour job
limit. That turns a transient apt or PyPI problem into a lost run, and it hides
what happened: the job simply sits there, and the failure that eventually gets
reported says nothing about which step was stuck.

Bound the three steps that do real work. Ten minutes for the install, against a
normal worst case of 152 seconds. Fifteen for the build, which has run between 4
and 54 seconds. Sixty for the test step, which is the only one whose length
depends on the suites themselves; the longest observed is 37 minutes, and that
was with a wait in one test that has since been bounded.

A step that trips its timeout fails and names itself, which is the point. None
of these numbers is tight enough to trip on work that is merely slow.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-20 08:44:49 -04:00
Frédéric Desbiens 68eb205768 Made the Arm ports build with LLVM, and added a check that keeps them that way (#593)
The gnu ports are only ever built with GNU tooling, and GNU as accepts several
non-canonical forms that LLVM's assembler rejects. Nothing noticed, because
nothing built them with anything else. This matters beyond clang itself: Arm
Toolchain for Embedded is LLVM based and is the successor to Arm Compiler 6, so
these are the code paths ac6 users move onto.

Seven files needed changing, none of which alters the emitted code:

LDREX and STREX take no offset in A32 state; the #imm form is Thumb-2 only. GNU
as drops the redundant zero, LLVM rejects it. Removed from the Cortex-A5, A7 and
A9 SMP protect routines.

ARMv8-M Baseline has no flag-preserving MOV immediate, so GNU as already emits
MOVS. Writing MOVS in the two Cortex-M23 sources says what the assembler was
doing anyway. One of them sits in a branch only compiled for the single mode
secure configurations, which is why it had never surfaced.

The Cortex-M0 schedule routine wrote LDR r0, =#0x10000000 with a stray hash,
which its own sibling file already wrote correctly.

The Cortex-M0 system return routine selected the numbered subsection .text 32,
which makes LLVM place the constant pool beyond the range a Thumb-1 PC relative
load can reach. Plain .text fixes it and GNU accepts either form. The reason is
recorded in the file, since 32 files pair a numbered subsection with a literal
pool load and the rest only escape because Thumb-2 and A32 have far more range.

Add scripts/check_clang.sh, which assembles every Arm gnu port source and
compiles the common C sources for one core per architecture profile, and a
clang_check workflow that installs Arm Toolchain for Embedded and runs it. The
toolchain version is pinned and checksum verified, for the same reason the
runner image is pinned.

The port directory to target mapping in that script is explicit rather than
prefix matched. Prefix matching is what makes cortex_a5 also match cortex_a53
and cortex_a55, which are AArch64, and assembling those as ARM32 produces
hundreds of misleading errors; that mistake cost real time while measuring this,
so the reason is recorded next to the table.

Verified with Arm Toolchain for Embedded 22.1.0: 711 of 711 assembly sources
assemble and 185 of 185 C sources compile for all eight profiles, against 6
assembly failures before the change. arm-none-eabi-gcc still assembles all 315
ARM32 sources, so nothing regressed for GNU. The check was confirmed to fail
when any one of the fixes is reverted.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-10 11:40:54 -04:00
Frédéric Desbiens f23a1807a2 Added bash A profile update scripts and restored ports_arch as their source (#592)
The ARMv7-A and ARMv8-A ports are generated by update.ps1, which needs
PowerShell, so the cortex-a job in ports_arch_check ran on a Windows image and
nobody could reproduce it locally on Linux. Add update.sh beside each
update.ps1, with the same cores, compilers, copy sets and patches, and move the
job to the same Linux image as everything else.

The bash scripts were checked against the PowerShell ones by comparing what
each reports as drifted. They agree exactly on the 63 files the Windows job
last reported, and differ on 12 more, which turn out to be a defect in
update.ps1 rather than in the port. Its two .cproject patterns are written as
'value=`"cortex-a7`"' with backticks that survive into the pattern, so that
replacement has never matched, while the neighbouring Cortex-A7.NoFPU pattern
has no backticks and always worked. The result is that the AC6 example builds
for the A5, A8, A9, A12, A15 and A17 cores name cortex-a7 as their CPU while
their FPU string is correct. The bash scripts do what the PowerShell ones
intended, so regenerating corrects those twelve files.

Restore ports_arch as the source for the rest. The implementation of
_tx_thread_smp_time_get from #555 was applied to the twenty four generated SMP
ports and never to ports_arch, which still held MOV x0, #0 with a FIXME
comment, so regenerating would have replaced a working generic timer read with
a stub. That implementation now lives in the source. The remaining differences
are cosmetic and resolve in favour of the source: a trailing blank line in 38
copies of tx_thread_schedule.S and comment spacing in one tx_port.h.

Note that the Cortex-A VFP fix is already present in ports_arch and was never
at risk, contrary to what the description of the port consistency checks change
said before this was measured.

Extend scripts/check_ports.sh to run the A profile generators too, and make it
fail when a generator fails or is missing rather than reporting a clean tree,
which would have been a false pass.

Pin every workflow to ubuntu-24.04. ubuntu-latest already resolves to that
image, so nothing changes today, but a future migration becomes a deliberate
commit rather than something that happens underneath the -m32 builds.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 10:55:29 -04:00
Frédéric Desbiens 08b120d2bc Added port consistency checks and wired them into CI and release preparation (#591)
Three defects reached the repository through the port trees recently, and each
of them is mechanically detectable without a cross compiler. Add
scripts/check_ports.sh, which looks for exactly those three, and give CI and
the release process the same command a contributor can run locally.

The generated Cortex-M ports must be reproducible from ports_arch. Fixes were
applied to the generated copies instead of the source for eight months, and the
next run of the copy scripts would have reverted them.

Preprocessor directives must balance. A fix left the Cortex-M85 IAR tx_port.h
with one more #endif than #if, so that header could not compile.

No port header may carry a statement outside a function body. A fix left a
second, headerless copy of a function body in the Cortex-M4 AC6 tx_port.h,
which is issue 569. The check tracks brace depth while skipping preprocessor
lines, multi-line macro bodies and comments, and reports assignments,
dereferences and control statements that land at file scope. Headers under
example_build are excluded, since those trees vendor third party SDK code.

A fourth section reports, without failing the run, on port families that have
no copy script and so cannot be checked for reproducibility. It currently
observes that the Cortex-M0 ac5, ac6 and keil ports lack the barriers their gnu
and iar siblings have.

ports_arch_check now calls the script rather than inlining a copy and diff, so
CI and the command line check the same things by the same definition, and the
workflow now triggers on pull requests to dev as well as master. Triggering on
master alone is why the drift went unseen. prepare_release.sh runs the checks
before it branches or rewrites anything, and stops if they fail, with
SKIP_PORT_CHECKS=1 as the escape hatch.

Each check was verified by reintroducing the defect it exists to catch and
confirming that the script fails, then confirming it passes on a clean tree.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 10:48:29 -04:00
Frédéric Desbiens f3df5f9dde Added a regression suite for the FreeRTOS compatibility layer (#583)
* Fixed the resource leaks on the xQueueCreate error paths

xQueueCreate() allocated the queue descriptor and its backing memory, then
created two ThreadX semaphores, and returned NULL on either semaphore failure
without releasing anything. Since no handle reached the caller, vQueueDelete()
could not be used to recover, so both allocations were lost. A failure on the
second semaphore additionally abandoned the read semaphore it had already
created, leaving a live ThreadX control block inside freed memory.

Release the backing memory and the descriptor on both paths, and delete the
read semaphore before returning when the write semaphore cannot be created.
This is the teardown order vQueueDelete() already uses, and it matches the
cleanup xTaskCreate() performs on its own error paths.

Verified with a fault injection harness that intercepts the ThreadX byte pool
and semaphore entry points to force tx_semaphore_create() to fail on a chosen
call. On a read semaphore failure the layer previously performed 2 allocations
and 0 releases, and on a write semaphore failure 2 allocations, 0 releases and
0 semaphore deletions. It now performs 2 releases in both cases and deletes
the read semaphore in the second, with the byte pool restored to its prior
state.

Fixes https://github.com/eclipse-threadx/threadx/issues/570

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>

* Added a regression suite for the FreeRTOS compatibility layer

The compatibility layer had no tests in this repository, which is awkward for
its creation functions in particular. Each of them takes one or two byte pool
allocations for its bookkeeping and then creates ThreadX kernel objects, and
each returns NULL when a kernel object cannot be created. The caller is left
without a handle, so it cannot call the matching delete function, and anything
the layer failed to release is gone until the system restarts. A leaking
version and a correct version are indistinguishable from the outside, which is
how the leak in issue 570 went unnoticed.

Add a suite that counts what the layer takes and gives back. A test asks the
harness to fail a chosen kernel creation call, then checks the number of byte
pool allocations, releases, object creations and object deletions performed.
The ThreadX entry points are intercepted with the linker's --wrap so that
tx_freertos.c is compiled exactly as it ships, with no test hooks in it. Note
that tx_api.h maps the public API onto the error checking entry points, so the
_txe_ symbols are the ones wrapped. Coverage is the creation and teardown paths
of queues, tasks, semaphores, mutexes, event groups and timers, including a
regression test for the two paths fixed for issue 570.

The suite follows the layout of the existing ThreadX and SMP suites, is
registered with ctest, and runs in CI through the shared regression template.
It is built 32 bit because the Linux port defines ULONG as unsigned int on
x86_64 while the layer passes pointers through ULONG arguments, so a 64 bit
build truncates them. It is Linux only because --wrap has no MSVC equivalent,
and the CMake configuration says so rather than failing at link time.

Validated by building the suite against the layer as it stands before the
issue 570 fix, where the two expected checks fail with the leaked counts, and
against the fixed layer, where all three tests pass.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-08-09 08:03:16 -04:00
Frédéric DesbiensandCopilot 223556219+Copilot@users.noreply.github.com 7486de06c8 Refactored, consolidated, and cleaned up RV32/RV64 ports (#536)
risc-v: refactor, consolidate, and fix RV32/RV64 ports

Consolidates the RISC-V 32-bit and 64-bit GNU/Clang port sources, fixes two
pre-existing assembly bugs discovered during testing, and hardens the build
infrastructure for both the regression suite and the CORE-V MCU example.

--- Port consolidation (RV32 GNU + Clang) ---

 - Delete ports/risc-v32/clang/src/ (8 .S files had no Clang-specific
 directives; diverged from GNU only due to missing bug fixes). The Clang
 port CMakeLists.txt now compiles from ../gnu/src/.
 - Change .global -> .weak for _tx_initialize_low_level in gnu/src/ to allow
 BSP-level override without a linker conflict (adopted from Clang port).
 - Create ports/risc-v32/common/tx_port_riscv32_common.h with all definitions
 shared between GNU and Clang ports. Reduce both tx_port.h files to thin
 wrappers.
 - Add a prominent comment in risc-v64/gnu/inc/tx_port.h explaining why
 LONG/ULONG are intentionally 32-bit on RV64 (ThreadX ABI requirement,
 mirrors win64/MSVC LLP64).

--- Shared CMake helper ---

 - Add cmake/threadx_riscv_port.cmake with threadx_add_riscv_port(). All
 three port CMakeLists.txt files are reduced to ~8 lines each. Include path
 is relative to CMAKE_CURRENT_LIST_DIR so the helper works whether ports
 are built standalone or as a subdirectory of the test framework.

--- Shared example-build drivers ---

 - Create canonical driver files under ports/risc-v_common/:
  inc/csr.h                  (uintptr_t-based; portable RV32 + RV64)
  example_build/plic/        (plic.c, plic.h)
  example_build/uart/        (uart_qemu_ns16550.c/h; static inline putc_nolock)
  example_build/trap/        (trap_qemu.c; XLEN-portable mcause constants)
 - Replace per-example copies with symlinks in all qemu_virt and cva6_ariane
 example directories.
 - Fix OS_IS_INTERRUPT typo (was OS_IS_INTERUPT) in shared trap_qemu.c.
 - Gate print_hex() behind TX_RISCV_TRAP_DEBUG.

--- Bug fixes in RV32 assembly ---

tx_thread_schedule.S:

 - Solicited-return FP path: reload t0 from the mepc stack slot before
 csrw mepc, t0. After the FP restore block, t0 held the fcsr value (0 for
 new threads), which caused mepc = 0 and an immediate instruction-address
 fault on the first context switch.
 - Same path: reload t0 from the mstatus stack slot before csrw mstatus, t0
 to avoid writing the stale fcsr value into mstatus.

tx_thread_system_return.S:

 - FP callee-saved registers were saved unconditionally before the mstatus.FS
 check, causing an illegal instruction trap (mcause=0x2) when a thread with
 FS=Off (lazy FPU, thread has never used FP) voluntarily yielded.
 - Apply the same FS guard pattern used in tx_thread_context_save.S: read
 mstatus first, isolate FS[1:0], and skip fsw/fsd if FS == Off.

Both bugs were pre-existing on origin/dev and are unrelated to the
consolidation changes.

--- RV64 64-bit pointer compatibility ---

 - Add TX_TIMER_INTERNAL_EXTENSION, TX_THREAD_CREATE_TIMEOUT_SETUP, and
 TX_THREAD_TIMEOUT_POINTER_SETUP to risc-v64/gnu/inc/tx_port.h to store the
 thread timeout pointer in a VOID
  * extension field rather than truncating it
 into a 32-bit ULONG. Mirrors the win64 port pattern.
 - Define TX_TIMER_EXTENSION_PTR_DEFINED as a portable sentinel.
 - Update threadx_thread_basic_execution_test.c guard from #if defined(_WIN64)
 to #if defined(_WIN64) || defined(TX_TIMER_EXTENSION_PTR_DEFINED).
 - Disable -Wconversion for the RV64 test build: ULONG = unsigned int (32-bit)
 is intentional for ThreadX ABI but triggers spurious warnings when sizeof()
 (8 bytes on RV64) appears in arithmetic with ULONG in common/src/.

--- Regression suite cmake fixes ---

test/tx/cmake/riscv/regression/CMakeLists.txt:

 - Build testcontrol_weak_defaults.c as a separate OBJECT library and include
 it in every test executable via $<TARGET_OBJECTS:>. GNU ld does not extract
 objects from a static archive to satisfy weak symbols, so bundling it in
 test_utility was insufficient for the standalone
 threadx_initialize_kernel_setup_test.

test/tx/cmake/regression/CMakeLists.txt,
test/smp/cmake/regression/CMakeLists.txt:

 - Same fix applied to the Linux and SMP regression builds. The symbols
 abort_all_threads_suspended_on_mutex, suspend_lowest_priority, and
 abort_and_resume_byte_allocating_thread were introduced by the win64 merge
 and left the standalone test unlinkable.

--- CORE-V MCU toolchain and build fixes ---

cmake/riscv64-gcc-rv32imc.cmake:

 - Resolve riscv64-unknown-elf-gcc via PATH so the riscv-collab toolchain in
 /opt/riscv/bin is preferred when it appears first.

ports/risc-v32/gnu/example_build/core_v_mcu/bsp/clz.c (new):

 - The riscv-collab toolchain is built without rv32 multilib, so its libgcc
 does not define __clzsi2 (the helper emitted for __builtin_clz() in fll.c).
 Add a weak __clzsi2 fallback so the build is self-contained with any
 riscv64-unknown-elf toolchain. The weak attribute yields to a
 libgcc-provided strong symbol when the Ubuntu multilib package is used.

core_v_mcu/CMakeLists.txt:

 - Add bsp/clz.c to sources.
 - Reference CMAKE_TOOLCHAIN_FILE via message(STATUS) to suppress the false-
 positive "Manually-specified variables were not used by the project" CMake
 warning and to show the active toolchain at configure time.

--- Housekeeping ---

 - Rename azrtos_test_* -> threadx_test_* (eliminate Azure RTOS branding).
 - Add RV64 QEMU CI test script:
  ports/risc-v64/gnu/example_build/qemu_virt/test/
  threadx_test_tx_gnu_riscv64_qemu.py
 - Normalize entry.s -> entry.S in all 4 example directories.
 - .gitignore: exclude build_m7/ and .codex local artifacts.
 - CI: comment out the riscv regression workflow job and remove it from the
 deploy job's needs list (preserved in-place for easy re-enablement).

--- Verified ---

 - 95/95 RV32 regression tests pass (QEMU virt)
 - 95/95 RV64 regression tests pass (QEMU virt)
 - All 5 Linux build configurations build cleanly (default_build_coverage,
 disable_notify_callbacks_build, stack_checking_build,
 stack_checking_rand_fill_build, trace_build)
 - CORE-V MCU example_build links cleanly with /opt/riscv toolchain

Co-authored-by: Copilot 223556219+Copilot@users.noreply.github.com
2026-05-27 10:30:57 -04:00
Akif Ejaz c1e3678797 Add QEMU based CI regression test infra for RV32 and RV64 (#526)
Added a QEMU virt-machine BSP and CTest infrastructure to run the
ThreadX regression suite on both RISC-V 32-bit and 64-bit targets
in CI.

New components:
- BSP (entry, trap, PLIC, CLINT timer, UART, linker script) targeting
  QEMU virt machine for RV32 and RV64
- CMake build system with Ninja, supporting multiple build configs
- CI scripts: install_riscv.sh (toolchain + QEMU), build_tx_riscv.sh,
  test_tx_riscv.sh
- GitHub Actions workflow job for RISC-V regression gating

Port fixes:
- RV32 tx_thread_context_restore.S: set MPIE alongside MPP (0x1800 →
  0x1880) so mret re-enables interrupts
- RV32/RV64 tx_port.h: add TX_REGRESSION_TEST extension macros needed
  by the test harness
- RV32/RV64 example_build scripts: add compile and QEMU launch steps

Regression test portability fixes:
- Block memory tests: increase pool sizes (320 → 340) to accommodate
  larger RISC-V block-header alignment
- Byte memory test: replace hardcoded offsets with BYTE_POOL_OVERHEAD
  macro for portable pool-size computation
- Event flag timeout test: make counter tolerance unconditional,
  removing linux-only guard

Signed-off-by: Akif Ejaz <akif.ejaz@10xengineers.ai>
2026-05-19 11:03:23 -04:00
Frédéric Desbiens c3259a2160 Updated copyright headers and version number constants (#509)
* Updated version number constants

* Removed revision history from all files

* Added Eclipse ThreadX contributors' copyright header
2026-03-05 10:46:30 +01:00
Frédéric Desbiens 8616486d99 Fixed code coverage report download step in deploy_code_coverage. 2025-07-29 17:05:07 -04:00
Frédéric Desbiens 8e808e70f1 Added condition to "Coverage Report Name". Corrected formatting. 2025-07-29 16:46:45 -04:00
Frédéric Desbiens c00056bb78 Fixed code coverage artefacts upload 2025-07-29 16:23:08 -04:00
Frédéric Desbiens 754c348568 Updated all actions to their latest release.
Signed-off-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2025-07-17 19:51:12 -04:00
Frédéric Desbiens b19b468e13 Added workflow permissions.
Signed-off-by: Frédéric Desbiens <frederic.desbiens@eclipse-foundation.org>
2025-07-17 16:26:04 -04:00
Frédéric Desbiens 7ad78c40e9 Merge pull request #387 from netomi/patch-1
regression_test / tx (push) Has been skipped
regression_test / smp (push) Has been skipped
regression_test / deploy (push) Has been skipped
cortex_m / Cortex M0 build (push) Has been cancelled
cortex_m / Cortex M3 build (push) Has been cancelled
cortex_m / Cortex M4 build (push) Has been cancelled
cortex_m / Cortex M7 build (push) Has been cancelled
Add new label for all created issues, enforce the use of issue templates
2025-02-27 11:32:30 -05:00
Frédéric Desbiens 72ab369eb7 Upgrading upload-artifact version to 4.6.0. 2025-02-13 10:58:58 -05:00
dependabot[bot] 7add20186e Bump actions/download-artifact from 3 to 4.1.7 in /.github/workflows
Bumps [actions/download-artifact](https://github.com/actions/download-artifact) from 3 to 4.1.7.
- [Release notes](https://github.com/actions/download-artifact/releases)
- [Commits](https://github.com/actions/download-artifact/compare/v3...v4.1.7)

---
updated-dependencies:
- dependency-name: actions/download-artifact
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
2024-09-03 22:12:22 +00:00
Thomas Neidhart 7b2259c052 Add new label 2024-05-27 09:05:29 +02:00
Thomas Neidhart 23691895c9 Add new label 2024-05-27 09:05:07 +02:00
Thomas Neidhart f648d4b461 Add new label 2024-05-27 09:04:46 +02:00
Thomas Neidhart bfde2d5493 Add new label 2024-05-27 09:04:13 +02:00
Thomas Neidhart e7fde60363 Create config.yml 2024-05-27 09:02:53 +02:00
Stefan Wick 0093df510b Update hardware-or-architecture-support.md 2024-03-10 22:09:40 -07:00
Stefan Wick bbce1f326c Update feature_request.md 2024-03-10 22:09:22 -07:00
Stefan Wick 7a29b8b463 Update bug_report.md 2024-03-10 22:08:50 -07:00
Eric Wolz 66954968b1 Delete .github/workflows/codeql.yml 2024-01-18 11:31:07 -08:00
Eric Wolz 1047cd7261 Update CODEOWNERS 2024-01-11 13:48:24 -08:00
Wenhui Xie 23aa67c948 Add sudo to move coverage folder created by root user. 2023-12-22 01:36:38 +00:00
Ting Zhu 776ea213ce Correct condition syntax in "Prepare Coverage GitHub Pages" step.l (#329)
* Correct syntax.

* Update regression_template.yml

* Update regression_template.yml
2023-11-30 10:52:29 +08:00
Ting Zhu ebe373b1f3 Add additional condition for "Prepare Coverage GitHub Pages" step. 2023-11-29 14:17:16 +08:00
CQ XiaoandTiejunZhou a8e5d0946c Added input skip_coverage and coverage_name for customiziing. (#327)
* Update regression_template.yml

* Update regression_template.yml

* Update regression_template.yml

* Update regression_template.yml

* Update regression_template.yml

* Update regression_template.yml

Coverage name of all -> default_build_coverage

* Update regression_template.yml

* Update regression_template.yml

Check inputs.skip_coverage to do steps.

* Update regression_template.yml

* Update regression_template.yml

* Update regression_template.yml

* Update regression_template.yml

Enable coverage upload when manually triggered.

* Update regression_template.yml

* Update regression_template.yml

* Update regression_template.yml

* Update .github/workflows/regression_template.yml

Fix comments.

Co-authored-by: TiejunZhou <50469179+TiejunMS@users.noreply.github.com>

---------

Co-authored-by: TiejunZhou <50469179+TiejunMS@users.noreply.github.com>
2023-11-27 11:50:52 +08:00
TiejunZhou cad6c42ecc Fix action to run on fork repo for cortex-m builds and unify them (#326) 2023-11-24 16:23:21 +08:00
TiejunZhou e420e2fa02 Restrict deploy run condition (#325)
* Restrict deploy run condition

* Fail the job for testing purpose

* Revert "Fail the job for testing purpose"

This reverts commit 6ae18cafe2.
2023-11-24 15:41:34 +08:00
TiejunZhou 5f9c713c48 Split artifacts for multiple jobs (#324)
* Test multiple code coverage pages

* Add affix to artifacts

* Test uploading code coverage as artifact

* Deploy GitHub pages at last for multiple jobs

* Test using unified upload pages

* Disable test cases to accelerate experiment

* Fix escape character $

* Revert "Test using unified upload pages"

This reverts commit 3668d9f672.

* Set destination for downloaded artifact

* Use a different artifact name

* Fix escape value

* Revert "Disable test cases to accelerate experiment"

This reverts commit 8468f17d02.

* Override duplicated github-pages in artifact

* Revert "Override duplicated github-pages in artifact"

This reverts commit 17a83aa97d.

* Delete Duplicate Code Coverage Artifact
2023-11-24 13:59:26 +08:00
TiejunZhou 11a7db22b4 Convert ADO pipelines to GitHub actions (#321)
* Convert ADO pipelines to GitHub actions

* Remove version in uses as not valid for local workflows

* Fix cmake path and add deploy url affix

* Add SMP build job

* Fix code coverage URL

* Add affix to titles of steps

* Remove ADO pipelines

* Add affix to titles of code coverage

* separate PR results for multiple jobs

* Revert "separate PR results for multiple jobs"

This reverts commit 6da13540fd.

* separate PR results for multiple jobs
2023-11-23 13:17:52 +08:00
TiejunZhou 1ffd7c2cde Allow manual trigger for CodeQL action (#286) 2023-07-13 13:24:31 +08:00
TiejunZhou fd2bf7c19a Enable CodeQL (#285)
* Enable CodeQL

* Build cortex-m0 in CodeQL

* Trigger the CodeQL by cron only
2023-07-13 13:09:47 +08:00
TiejunZhou 08380caa77 Unify ThreadX and SMP for ARMv8-A. (#275)
* Unify ThreadX and SMP for ARMv8-A.

* Fix path in pipeline to check ports arch.

* Add ignore folders for ARM DS

* Generate ThreadX and SMP ports for ARMv8-A.

* Ignore untracked files for ports_arch check.

* Use arch instead of CPU to simplify the project management.
2023-06-21 18:23:36 +08:00
TiejunZhou 25a8fa2362 Add a pull request template (#272) 2023-06-06 14:03:35 +08:00
TiejunZhou 672c5e953e Release ARMv7-A architecture ports and add tx_user.h to GNU port assembly files (#250)
* Release ARMv7-A architecture ports

* Add tx_user.h to GNU port assembly files

* Update GitHub action to perform check for Cortex-A ports
2023-04-19 17:56:09 +08:00
TiejunZhou 23680f5e5f Release ARMv7-M and ARMv8-M architecture ports (#249)
* Release ARMv7-M and ARMv8-M architecture ports

* Add a pipeline to check ports_arch
2023-04-18 18:11:20 +08:00
TiejunZhou d64ef2ab06 Filter the path for PR trigger and add codeowners (#248)
* Filter the path for PR trigger

* Add codeowners

* Fix syntax in pipeline
2023-04-17 13:16:14 +08:00