mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
The port suspends a thread with a signal whose handler calls sigsuspend and does
not return until resumed, and takes its critical section with a bare
pthread_mutex_lock. A thread signalled while parked on _tx_linux_mutex never
returns to glibc's contended-mutex loop, so the next unlock's wake-up goes to a
thread that will not act on it and every other waiter is left parked on a mutex
that is free.
tx_linux_mutex_lock now calls a helper that waits with pthread_mutex_timedlock and
retries every TX_LINUX_MUTEX_RETRY_NSEC, one millisecond, so a lost wake-up costs a
retry period instead of the process.
The original fix was made by inspection, on a port never observed to deadlock. It
has now been observed. A stalled regression test reads the mutex free, owner zero,
with the scheduler still on its futex word; it never posts _tx_linux_isr_semaphore,
the timer thread never returns from _tx_thread_context_restore, and the simulated
clock stops. Bounding a wait cannot save it: one capture held 116 ticks of timeout
unchanged across 40 seconds.
Measured on netx_tcp_overlapping_packet_test_10, six workers on four pinned CPUs:
5 stalls in 1,266 runs before, 0 in 2,034 after, three held frozen through 60
seconds of untraced /proc sampling. netxduo/default_build_coverage is unmoved at
914/914, 121 skipped, 497 files, 11230/11256 lines and 6912/6925 branches.
(cherry picked from commit b567428f2a)
Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>