Files
threadx/ports
Frédéric Desbiens b567428f2a Stopped the Linux port losing a mutex wake-up to a suspend signal (#754)
This is the ThreadX half of the defect fixed for ThreadX SMP in an earlier pull
request. The Linux port suspends a thread with a signal whose handler calls
sigsuspend and does not return until the thread is resumed, and it takes its
critical section with a bare pthread_mutex_lock through tx_linux_mutex_lock. A
thread can therefore be signalled while it is parked on _tx_linux_mutex.

glibc waits for a contended mutex in a loop that re-arms the futex wait after a
signal, and this handler never returns to it. The next unlock hands its wake-up to
that thread, which will not act on it, and any other thread parked on the mutex is
never woken, leaving the mutex free with waiters on it.

tx_linux_mutex_lock now calls a helper that waits with pthread_mutex_timedlock and
retries, so the wait is re-armed every TX_LINUX_MUTEX_RETRY_NSEC and a lost
wake-up costs one retry period instead of the process. The period is one
millisecond. pthread_mutex_timedlock needs _GNU_SOURCE under -std=c99, which both
of this port's build systems already define.

These are the only two ports affected: a sigsuspend-based suspend handler exists
in the Linux and SMP Linux ports and nowhere else, and those two are also the only
ports taking a critical section with pthread_mutex_lock.

Unlike the SMP port, this one has never been observed to deadlock, which is
consistent with one emulated core and far less suspend and resume traffic. The fix
is by inspection, and verified as breaking nothing: all seven configurations pass
with retries disabled, 105/105 in five and 100/100 in two.

Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>
2026-09-28 12:42:48 -04:00
..