mirror of
https://github.com/eclipse-threadx/threadx.git
synced 2026-10-06 06:59:08 +08:00
The SMP Linux port suspends a thread with a signal whose handler calls sigsuspend and does not return until the thread is resumed. That signal can arrive while the thread is parked in pthread_mutex_lock on _tx_linux_mutex; the port knows it can, because _tx_linux_mutex_obtain sets tx_thread_linux_mutex_access around the lock call for exactly this case, and nothing anywhere reads that flag. glibc waits for a contended mutex in a loop that re-arms the futex wait after a signal, and this handler never returns to it. The next release hands its wake-up to that thread, which will not act on it, and any other thread parked on the mutex is never woken, leaving the mutex free with waiters on it. That deadlocks the process: the only thread that can resume the suspended one is the scheduler, and the scheduler takes this mutex on every pass. _tx_linux_mutex_obtain now waits with pthread_mutex_timedlock and retries, so the wait is re-armed every TX_LINUX_MUTEX_RETRY_NSEC and a lost wake-up costs one retry period instead of the process. The period is one millisecond, half the scheduler's own idle period. Nothing else changes. Four hung processes captured untraced, across three tests, show the same state: a thread in sigsuspend on top of pthread_mutex_lock, the scheduler blocked in pthread_mutex_lock, and the mutex reading free. Standalone, threadx_smp_random_resume_suspend_exclusion_test hung 3 times in 100 runs and 2 in 55 before the change and 0 in 400 after it. Assisted-by: Claude Code (Opus 5) <noreply@anthropic.com>