syzbot
possible deadlock in bh_worker
Status:
upstream: reported on 2026/08/18 12:48
Subsystems:
kernel
Labels:
prio:low
[Documentation on labels]
Reported-by: syzbot+1bd20115328f8254ed62@syzkaller.appspotmail.com
Fix commit:
workqueue: Annotate cb_lock nesting when draining a dead BH pool
Patched on: [ci-upstream-linux-next-kasan-gce-root ci-upstream-rust-kasan-gce], missing on: [ci-qemu-gce-upstream-auto ci-qemu-native-arm64-kvm ci-qemu-upstream ci-qemu-upstream-386 ci-qemu2-arm32 ci-qemu2-arm64 ci-qemu2-arm64-compat ci-qemu2-arm64-mte ci-qemu2-riscv64 ci-snapshot-upstream-root ci-upstream-bpf-kasan-gce ci-upstream-bpf-next-kasan-gce ci-upstream-gce-arm64 ci-upstream-gce-leak ci-upstream-kasan-badwrites-root ci-upstream-kasan-gce ci-upstream-kasan-gce-386 ci-upstream-kasan-gce-root ci-upstream-kasan-gce-selinux-root ci-upstream-kasan-gce-smack-root ci-upstream-kmsan-gce-386-root ci-upstream-kmsan-gce-root ci-upstream-net-kasan-gce ci-upstream-net-this-kasan-gce ci2-upstream-fs ci2-upstream-kcsan-gce ci2-upstream-usb]
First crash: 14d, last: 4d04h
Sample crash report:
============================================
WARNING: possible recursive locking detected
syzkaller #0 Not tainted
--------------------------------------------
kworker/u8:12/1171 is trying to acquire lock:
ffff8880b873b868 (&pool->cb_lock){+...}-{3:3}, at: spin_lock include/linux/spinlock_rt.h:45 [inline]
ffff8880b873b868 (&pool->cb_lock){+...}-{3:3}, at: worker_lock_callback kernel/workqueue.c:3200 [inline]
ffff8880b873b868 (&pool->cb_lock){+...}-{3:3}, at: bh_worker+0x86/0x890 kernel/workqueue.c:3753
but task is already holding lock:
ffff8880b863b868 (&pool->cb_lock){+...}-{3:3}, at: spin_lock include/linux/spinlock_rt.h:45 [inline]
ffff8880b863b868 (&pool->cb_lock){+...}-{3:3}, at: worker_lock_callback kernel/workqueue.c:3200 [inline]
ffff8880b863b868 (&pool->cb_lock){+...}-{3:3}, at: bh_worker+0x86/0x890 kernel/workqueue.c:3753
other info that might help us debug this:
Possible unsafe locking scenario:
CPU0
----
lock(&pool->cb_lock);
lock(&pool->cb_lock);
*** DEADLOCK ***
May be due to missing lock nesting notation
locks held by kworker/u8:12/1171: 7, last CPU#0:
#0: ffff888032868938 ((wq_completion)bat_events){+.+.}-{0:0}, at: rcu_lock_acquire include/linux/rcupdate.h:309 [inline]
#0: ffff888032868938 ((wq_completion)bat_events){+.+.}-{0:0}, at: rcu_read_lock include/linux/rcupdate.h:849 [inline]
#0: ffff888032868938 ((wq_completion)bat_events){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3352 [inline]
#0: ffff888032868938 ((wq_completion)bat_events){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630 kernel/workqueue.c:3470
#1: ffffc90006787c40 ((work_completion)(&(&bat_priv->tt.work)->work)){+.+.}-{0:0}, at: rcu_lock_acquire include/linux/rcupdate.h:309 [inline]
#1: ffffc90006787c40 ((work_completion)(&(&bat_priv->tt.work)->work)){+.+.}-{0:0}, at: rcu_read_lock include/linux/rcupdate.h:849 [inline]
#1: ffffc90006787c40 ((work_completion)(&(&bat_priv->tt.work)->work)){+.+.}-{0:0}, at: process_one_work kernel/workqueue.c:3352 [inline]
#1: ffffc90006787c40 ((work_completion)(&(&bat_priv->tt.work)->work)){+.+.}-{0:0}, at: process_scheduled_works+0x97a/0x1630 kernel/workqueue.c:3470
#2: ffffffff8e1c3800 (rcu_read_lock){....}-{1:3}, at: __local_bh_disable_ip+0x3d/0x420 kernel/softirq.c:186
#3: ffff8880b863b868 (&pool->cb_lock){+...}-{3:3}, at: spin_lock include/linux/spinlock_rt.h:45 [inline]
#3: ffff8880b863b868 (&pool->cb_lock){+...}-{3:3}, at: worker_lock_callback kernel/workqueue.c:3200 [inline]
#3: ffff8880b863b868 (&pool->cb_lock){+...}-{3:3}, at: bh_worker+0x86/0x890 kernel/workqueue.c:3753
#4: ffffffff8e1c3800 (rcu_read_lock){....}-{1:3}, at: rcu_lock_acquire include/linux/rcupdate.h:309 [inline]
#4: ffffffff8e1c3800 (rcu_read_lock){....}-{1:3}, at: rcu_read_lock include/linux/rcupdate.h:849 [inline]
#4: ffffffff8e1c3800 (rcu_read_lock){....}-{1:3}, at: __rt_spin_lock kernel/locking/spinlock_rt.c:50 [inline]
#4: ffffffff8e1c3800 (rcu_read_lock){....}-{1:3}, at: rt_spin_lock+0x1e2/0x400 kernel/locking/spinlock_rt.c:57
#5: ffff88813ff21938 ((wq_completion)events_bh_highpri){+...}-{0:0}, at: rcu_lock_acquire include/linux/rcupdate.h:309 [inline]
#5: ffff88813ff21938 ((wq_completion)events_bh_highpri){+...}-{0:0}, at: rcu_read_lock include/linux/rcupdate.h:849 [inline]
#5: ffff88813ff21938 ((wq_completion)events_bh_highpri){+...}-{0:0}, at: process_one_work kernel/workqueue.c:3352 [inline]
#5: ffff88813ff21938 ((wq_completion)events_bh_highpri){+...}-{0:0}, at: process_scheduled_works+0x97a/0x1630 kernel/workqueue.c:3470
#6: ffffc900067877a0 ((work_completion)(&dead_work.work)){+...}-{0:0}, at: rcu_lock_acquire include/linux/rcupdate.h:309 [inline]
#6: ffffc900067877a0 ((work_completion)(&dead_work.work)){+...}-{0:0}, at: rcu_read_lock include/linux/rcupdate.h:849 [inline]
#6: ffffc900067877a0 ((work_completion)(&dead_work.work)){+...}-{0:0}, at: process_one_work kernel/workqueue.c:3352 [inline]
#6: ffffc900067877a0 ((work_completion)(&dead_work.work)){+...}-{0:0}, at: process_scheduled_works+0x97a/0x1630 kernel/workqueue.c:3470
stack backtrace:
CPU: 0 UID: 0 PID: 1171 Comm: kworker/u8:12 Not tainted syzkaller #0 PREEMPT_{RT,(full)}
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 07/24/2026
Workqueue: bat_events batadv_tt_purge
Call Trace:
<TASK>
dump_stack_lvl+0xe8/0x150 lib/dump_stack.c:120
print_deadlock_bug+0x279/0x290 kernel/locking/lockdep.c:3057
check_deadlock kernel/locking/lockdep.c:3109 [inline]
validate_chain kernel/locking/lockdep.c:3911 [inline]
__lock_acquire+0x24df/0x2ce0 kernel/locking/lockdep.c:5253
lock_acquire+0x106/0x350 kernel/locking/lockdep.c:5886
rt_spin_lock+0x83/0x400 kernel/locking/spinlock_rt.c:56
spin_lock include/linux/spinlock_rt.h:45 [inline]
worker_lock_callback kernel/workqueue.c:3200 [inline]
bh_worker+0x86/0x890 kernel/workqueue.c:3753
drain_dead_softirq_workfn+0x95/0x220 kernel/workqueue.c:3828
process_one_work kernel/workqueue.c:3387 [inline]
process_scheduled_works+0xc3d/0x1630 kernel/workqueue.c:3470
bh_worker+0x44d/0x890 kernel/workqueue.c:3773
tasklet_hi_action+0xf/0x70 kernel/softirq.c:1003
handle_softirqs+0x1da/0x6d0 kernel/softirq.c:645
__do_softirq kernel/softirq.c:679 [inline]
__local_bh_enable_ip+0x16f/0x2b0 kernel/softirq.c:325
local_bh_enable include/linux/bottom_half.h:33 [inline]
spin_unlock_bh include/linux/spinlock_rt.h:122 [inline]
batadv_tt_local_purge+0x2ed/0x380 net/batman-adv/translation-table.c:1402
batadv_tt_purge+0x35/0x9f0 net/batman-adv/translation-table.c:3753
process_one_work kernel/workqueue.c:3387 [inline]
process_scheduled_works+0xc3d/0x1630 kernel/workqueue.c:3470
worker_thread+0xa47/0xfb0 kernel/workqueue.c:3551
kthread+0x38b/0x470 kernel/kthread.c:436
ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
BUG: workqueue leaked atomic, lock or RCU: kworker/u8:12[1171]
preempt=0x00000000 lock=5->0 RCU=2->2 workfn=drain_dead_softirq_workfn
INFO: lockdep is turned off.
CPU: 0 UID: 0 PID: 1171 Comm: kworker/u8:12 Not tainted syzkaller #0 PREEMPT_{RT,(full)}
Hardware name: Google Google Compute Engine/Google Compute Engine, BIOS Google 07/24/2026
Workqueue: bat_events batadv_tt_purge
Call Trace:
<TASK>
dump_stack_lvl+0xe8/0x150 lib/dump_stack.c:120
process_one_work kernel/workqueue.c:3408 [inline]
process_scheduled_works+0x12c7/0x1630 kernel/workqueue.c:3470
bh_worker+0x44d/0x890 kernel/workqueue.c:3773
tasklet_hi_action+0xf/0x70 kernel/softirq.c:1003
handle_softirqs+0x1da/0x6d0 kernel/softirq.c:645
__do_softirq kernel/softirq.c:679 [inline]
__local_bh_enable_ip+0x16f/0x2b0 kernel/softirq.c:325
local_bh_enable include/linux/bottom_half.h:33 [inline]
spin_unlock_bh include/linux/spinlock_rt.h:122 [inline]
batadv_tt_local_purge+0x2ed/0x380 net/batman-adv/translation-table.c:1402
batadv_tt_purge+0x35/0x9f0 net/batman-adv/translation-table.c:3753
process_one_work kernel/workqueue.c:3387 [inline]
process_scheduled_works+0xc3d/0x1630 kernel/workqueue.c:3470
worker_thread+0xa47/0xfb0 kernel/workqueue.c:3551
kthread+0x38b/0x470 kernel/kthread.c:436
ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
* Struck through repros no longer work on HEAD.