[syzbot] [mm?] INFO: rcu detected stall in khugepaged (3)

0 views
Skip to first unread message

syzbot

unread,
Aug 4, 2026, 8:01:51 PM (24 hours ago) Aug 4
to ak...@linux-foundation.org, han...@cmpxchg.org, jack...@google.com, linux-...@vger.kernel.org, linu...@kvack.org, mho...@suse.com, sur...@google.com, syzkall...@googlegroups.com, vba...@kernel.org, z...@nvidia.com
Hello,

syzbot found the following issue on:

HEAD commit: 3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
git tree: upstream
console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
kernel config: https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
compiler: Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8

Unfortunately, I don't have any reproducer for this issue yet.

Downloadable assets:
disk image: https://storage.googleapis.com/syzbot-assets/c0390423374e/disk-3708dd94.raw.xz
vmlinux: https://storage.googleapis.com/syzbot-assets/3130bf9c5dbf/vmlinux-3708dd94.xz
kernel image: https://storage.googleapis.com/syzbot-assets/2c8fbf8aeba4/bzImage-3708dd94.xz

IMPORTANT: if you fix the issue, please add the following tag to the commit:
Reported-by: syzbot+d2401a...@syzkaller.appspotmail.com

rcu: INFO: rcu_preempt detected stalls on CPUs/tasks:
rcu: Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l
rcu: (detected by 1, t=10503 jiffies, g=32265, q=1190 ncpus=2)
task:khugepaged state:R running task stack:26864 pid:37 tgid:37 ppid:2 task_flags:0x200040 flags:0x00080000
Call Trace:
<TASK>
context_switch kernel/sched/core.c:5510 [inline]
__schedule+0x17d9/0x56c0 kernel/sched/core.c:7234
preempt_schedule_irq+0x4d/0xa0 kernel/sched/core.c:7556
irqentry_exit_to_kernel_mode include/linux/irq-entry-common.h:539 [inline]
irqentry_exit+0x14f/0x8f0 kernel/entry/common.c:167
asm_sysvec_apic_timer_interrupt+0x1a/0x20 arch/x86/include/asm/idtentry.h:674
RIP: 0010:lock_release+0x2d7/0x3c0 kernel/locking/lockdep.c:5893
Code: 7d ca 11 00 00 00 00 eb b5 e8 55 ad 2e 0a f7 c3 00 02 00 00 74 b9 65 48 8b 05 45 38 ca 11 48 3b 44 24 28 75 44 fb 48 83 c4 30 <5b> 41 5c 41 5d 41 5e 41 5f 5d c3 cc cc cc cc cc 48 8d 3d a2 33 b8
RSP: 0018:ffffc90000ad7028 EFLAGS: 00000286
RAX: 623a32d217845600 RBX: 0000000000000206 RCX: 0000000000000046
RDX: 0000000000000000 RSI: ffffffff8e4b3351 RDI: ffffffff8c4bdd80
RBP: ffff888020ea0ba0 R08: ffffc90000ad74d0 R09: 0000000000000000
R10: ffffc90000ad7158 R11: fffff5200015ae2d R12: 0000000000000000
R13: 0000000000000000 R14: ffffffff8eb59c60 R15: ffff888020ea0000
rcu_lock_release include/linux/rcupdate.h:310 [inline]
rcu_read_unlock include/linux/rcupdate.h:871 [inline]
class_rcu_destructor include/linux/rcupdate.h:1183 [inline]
unwind_next_frame+0x1baa/0x2550 arch/x86/kernel/unwind_orc.c:709
arch_stack_walk+0x11b/0x150 arch/x86/kernel/stacktrace.c:25
stack_trace_save+0xa9/0x100 kernel/stacktrace.c:122
save_stack+0x122/0x230 mm/page_owner.c:165
__reset_page_owner+0x71/0x1f0 mm/page_owner.c:320
reset_page_owner include/linux/page_owner.h:25 [inline]
__free_pages_prepare mm/page_alloc.c:1406 [inline]
__free_frozen_pages+0xc1e/0xd10 mm/page_alloc.c:2950
__folio_put+0x4b3/0x590 mm/swap.c:112
folio_put_refs include/linux/mm.h:2144 [inline]
collapse_file mm/khugepaged.c:2643 [inline]
collapse_scan_file+0x4285/0x5210 mm/khugepaged.c:2773
collapse_single_pmd+0x2b1/0x3da0 mm/khugepaged.c:2808
collapse_scan_mm_slot mm/khugepaged.c:2913 [inline]
khugepaged_do_scan mm/khugepaged.c:2993 [inline]
khugepaged+0xa00/0x1780 mm/khugepaged.c:3048
kthread+0x388/0x470 kernel/kthread.c:436
ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>


---
This report is generated by a bot. It may contain errors.
See https://goo.gl/tpsmEJ for more information about syzbot.
syzbot engineers can be reached at syzk...@googlegroups.com.

syzbot will keep track of this issue. See:
https://goo.gl/tpsmEJ#status for how to communicate with syzbot.

If the report is already addressed, let syzbot know by replying with:
#syz fix: exact-commit-title

If you want to overwrite report's subsystems, reply with:
#syz set subsystems: new-subsystem
(See the list of subsystem names on the web dashboard)

If the report is a duplicate of another one, reply with:
#syz dup: exact-subject-of-another-report

If you want to undo deduplication, reply with:
#syz undup

Andrew Morton

unread,
3:29 PM (4 hours ago) 3:29 PM
to syzbot, han...@cmpxchg.org, jack...@google.com, linux-...@vger.kernel.org, linu...@kvack.org, mho...@suse.com, sur...@google.com, syzkall...@googlegroups.com, vba...@kernel.org, z...@nvidia.com, Paul E. McKenney
On Tue, 04 Aug 2026 17:01:48 -0700 syzbot <syzbot+d2401a...@syzkaller.appspotmail.com> wrote:

> Hello,
>
> syzbot found the following issue on:
>
> HEAD commit: 3708dd948844 Merge tag 'pm-7.2-rc6' of git://git.kernel.or..
> git tree: upstream
> console output: https://syzkaller.appspot.com/x/log.txt?x=11ac703e580000
> kernel config: https://syzkaller.appspot.com/x/.config?x=4e38b15c29e6a1d9
> dashboard link: https://syzkaller.appspot.com/bug?extid=d2401aeb74cc84adba04
> compiler: Debian clang version 22.1.8 (++20260613092233+e80beda6e255-1~exp1~20260613092250.77), Debian LLD 22.1.8
>
> Unfortunately, I don't have any reproducer for this issue yet.

Thanks.

Lazy optimists (ahem) paste this gunk into Gemini and ask "what the
heck just happened". The results are often useful, but should be
treated with skepticism. In this case I think it came usably close.

https://share.gemini.google/vq4TLhTiLBih


tl;dr: khugepaged's collapse_scan_file() is taking too long and RCU got
starved. I don't think khugepaged is doing anything wrong here,
per-se. There's a lot of work to do and we're doing it.

An appropriate fix would be to take a break, let RCU do its thing then
get back to work. But I don't think RCU offers interfaces for that?

collapse_scan_file()'s main loop has

if (need_resched()) {
xas_pause(&xas);
cond_resched_rcu();
}

but that won't help with the RCU stall detector(?).

I suggest that a suitable fix here would be to add the analogous

if (rcu_i_need_to_take_a_break()) {
rcu_read_unlock();
rcu_take_a_break())
rcu_read_lock();
}

(iirc rcu_read_unlock() does an rcu run, so rcu_take_a_break() isn't
needed here)

Paul, wdyt?

Paul E. McKenney

unread,
4:28 PM (3 hours ago) 4:28 PM
to Andrew Morton, syzbot, han...@cmpxchg.org, jack...@google.com, linux-...@vger.kernel.org, linu...@kvack.org, mho...@suse.com, sur...@google.com, syzkall...@googlegroups.com, vba...@kernel.org, z...@nvidia.com
Let's see...

The console log says "rcu_preempt detected stalls on CPUs/tasks",
which means that cond_resched() is a no-op, but it also means that
the rcu_read_unlock() in cond_resched_rcu() will directly take care of
informing RCU of the pause.

But that is clearly not happening. Why?

Well, we have this:

rcu: Tasks blocked on level-0 rcu_node (CPUs 0-1): P37/1:b..l

This means that the task whose RCU read-side critical section is blocking
the current RCU grace period isn't even running, and thus cannot invoke
cond_resched_rcu(), let alone the rcu_read_unlock() within that function.
So an RCU CPU stall warning is expected behavior. Or at least it is not
in any way ruled out.

What we need is RCU priority boosting. Except that the .config file
does not enable this. Not only is there no CONFIG_RCU_BOOST=y, there
is also no CONFIG_RCU_EXPERT=y and no CONFIG_PREEMPT_RT=y. But there
is CONFIG_RT_MUTEX=y and CONFIG_RCU_EXPERT=y.

Because we don't have RCU priority boosting, if the load on the system
is heavy enough to prevent our poor preempted RCU reader (PID 37) from
running, the grace period cannot end.

I am not sure why this task is saving its stack, but maybe that is normal
for this code path?

My bemusement aside, I recommend running this test either with
non-preemptible RCU (CONFIG_PREEMPT_LAZY=y these days) or enabling RCU
priority boosting (CONFIG_RCU_EXPERT=y and CONFIG_RCU_BOOST=y).

Maybe RCU_BOOST should no longer depend on RCU_EXPERT? I would of
course need ot remove the prompt ("Enable RCU priority boosting") to
avoid annoying Linus. Maybe as shown below.

Thoughts?

Thanx, Paul

------------------------------------------------------------------------

diff --git a/kernel/rcu/Kconfig b/kernel/rcu/Kconfig
index 1a5fb3156c062a..5141ad8d1cd029 100644
--- a/kernel/rcu/Kconfig
+++ b/kernel/rcu/Kconfig
@@ -237,17 +237,16 @@ config RCU_FANOUT_LEAF
Take the default if unsure.

config RCU_BOOST
- bool "Enable RCU priority boosting"
- depends on (RT_MUTEXES && PREEMPT_RCU && RCU_EXPERT) || PREEMPT_RT
+ bool
+ depends on (RT_MUTEXES && PREEMPT_RCU) || PREEMPT_RT
default y if PREEMPT_RT
help
This option boosts the priority of preempted RCU readers that
block the current preemptible RCU grace period for too long.
This option also prevents heavy loads from blocking RCU
- callback invocation.
+ callback invocation. It is now automatically enabled in
+ any kernel that can benefit from it and that can support it.

- Say Y here if you are working with real-time apps or heavy loads
- Say N here if you are unsure.

config RCU_BOOST_DELAY
int "Milliseconds to delay boosting after RCU grace-period start"
Reply all
Reply to author
Forward
0 new messages