You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
overlay: sandbox can never exit after a self-deadlock on renameMu when an inotify Notify under doCreateAt removes an expired watch #14680
A container running on the overlay filesystem (--overlay2, default root:self) stops responding to SIGKILL. runsc state reports stopped with pid -1 while runsc ps still lists the workload processes, runsc kill returns success and nothing happens, and the shim's Wait never returns, so containerd and the kubelet retry the stop forever. The sentry itself is healthy: it answers runsc debug --stacks, and the dump shows one goroutine that has requested renameMu for write while the same goroutine already holds it for read. Every later overlay.(*dentry).DecRef, including the ones on the task-exit path, queues behind it, so the exit can never complete.
We have 43 sandboxes in this state, the oldest blocked for over 44 hours at the time of the dump.
goroutine 257515 [sync.RWMutex.Lock, 2656 minutes]:
overlay.(*renameRWMutex).Lock
overlay.(*dentry).OnZeroWatches <- wants renameMu for write
vfs.(*Inotify).RmWatch
vfs.(*Watches).cleanupExpiredWatches
vfs.(*Watches).Notify
overlay.(*filesystem).doCreateAt <- holds renameMu for read (deferred unlock)
overlay.(*filesystem).MkdirAt
... mkdir(2)
2 x goroutine [sync.RWMutex.Lock, 2628 minutes]:
overlay.(*renameRWMutex).Lock
overlay.(*dentry).DecRef
vfs.(*FileDescription).DecRef
mm.(*MemoryManager).DecUsers
kernel.(*runExitMain).execute <- the SIGKILL exit path, blocked behind the first goroutine
goroutine 1: kernel.(*Kernel).WaitExited (never returns)
goroutine 161: kernel.(*ThreadGroup).WaitExited (the shim's Wait)
Full dump available on request.
Analysis
Three pieces of code, unchanged between release-20260622.0, release-20260831.0 and master at d94a98a:
pkg/sentry/fsimpl/overlay/filesystem.go, doCreateAt: takes fs.renameMu.RLock() with defer fs.renameMuRUnlockAndCheckDrop(ctx, &ds) and calls parent.watches.Notify(...) before returning, so the notification runs under the read lock.
pkg/sentry/vfs/inotify.go: when Notify finds watches whose inotify instance has been released, cleanupExpiredWatches calls Inotify.RmWatch, which calls w.target.OnZeroWatches(ctx) synchronously on the notifying goroutine once the target has no watches left.
Walked directories in the overlay normally sit at zero references (they are kept alive by the parent's child map under renameMu, which is why the drop-list exists), so the condition holds for the parent of a create, and the write lock is requested by the goroutine that holds the read lock. Go's RWMutex never grants that.
The DentryImpl.OnZeroWatches contract only says no inotify locks may be held by the caller; it does not say anything about filesystem locks, and the overlay calls it from a path holding its own. The gofer filesystem's OnZeroWatches has the same shape through checkCachingLocked and its create paths also notify under the read lock, so the same hazard probably exists there; we have not observed it.
Sequence that triggers it
A process adds an inotify watch with IN_ONESHOT on a leaf directory of the overlay (a directory with no cached children, so its dentry sits at zero references) and keeps the inotify fd open. IN_ONESHOT is the only path that marks a watch expired and leaves its removal to Watches.Notify; releasing an inotify instance removes its watches synchronously with no filesystem lock held, so a watcher exiting is not a trigger.
Another process produces the first event on that directory: mkdir inside it (doCreateAt, renameMu.RLock) or a rename of it (RenameAt, renameMu.Lock, through InotifyRename / IN_MOVE_SELF). Watch.Notify fires the one-shot watch and marks it expired; the same Notify call runs cleanupExpiredWatches -> RmWatch -> OnZeroWatches, which requests renameMu.Lock on the goroutine that already holds renameMu, and the sandbox is wedged.
open(O_CREAT) in the same directory does not trigger it: getChildLocked takes a parent reference before the notification, so OnZeroWatches sees refs != 0 and skips the lock.
Reproducer inside a container (alpine 3.21 rootfs with python3; the path must be on the overlay-backed rootfs, not under /tmp, which runsc backs with a tmpfs whose OnZeroWatches is a no-op):
Results with runsc built from the go branch at 1db01ca (release-20260831.0-42), platform systrap, runsc run with the default --overlay2=root:self (root:memory behaves the same):
On the hung sandbox: runsc kill -all <id> KILL never returns (blocked in Kernel.Pause, because the task inside mkdirat never reaches a stop point); runsc kill <id> KILL returns 0, runsc state then reports stopped / pid -1 while runsc ps still lists the processes, and the init task's exit path is blocked in runExitMain -> FDTable.DecRef -> overlay.(*dentry).DecRef -> renameMu.Lock. Only runsc delete -force (SIGKILL of the sandbox process) ends it. The stacks match the production dump above frame for frame; the sentry watchdog logs Sentry detected 1 stuck task(s) every 45 s and, with the default --watchdog-action=log, does nothing else.
Possible fixes
Issue the inotify notification after renameMuRUnlockAndCheckDrop, holding a reference on the parent across the call. Every Notify issued while renameMu is held is exposed, not only the one in doCreateAt.
Or make the overlay's OnZeroWatches defer the drop (queue the dentry for the next renameMuRUnlockAndCheckDrop) instead of taking the write lock inline.
runsc release-20260622.0, platform systrap, --overlay2 default; the relevant code is identical in release-20260831.0 and master d94a98a.
containerd 2.2.5 with containerd-shim-runsc-v1, Kubernetes 1.35, Linux 6.12 x86_64.
Workload: Racket builds and tests. Racket's runtime adds its inotify watches with IN_ONESHOT (racket/src/rktio/rktio_fs_change.c) on the directories it watches (racket/src/expander/eval/collection.rkt), and the build creates directories under them.
I think this quirk is very specific to overlayfs, because in every other filesystem, when a child is created via doCreateAt(), the child takes a ref on the parent. But looking at overlayfs create path, it passes along the create call to upper layer dentry and doesn't construct an overlayfs child dentry which takes a ref on parent. So by the time parent.watches.Notify() is called, the parent is back to zero refs.
mkdir inside it (doCreateAt, renameMu.RLock) or a rename of it (RenameAt, renameMu.Lock, through InotifyRename / IN_MOVE_SELF).
Urgh, yeah it seems like it also impacts rename, because the old dir can then have no refs left on it? I suppose gofer is not impacted because it does not take the rename lock on OnZeroWatches()? It just takes some low level caching mutexes and we don't really destroy the dentry because it has an LRU cache. And overlayfs doesn't have an LRU cache so it tries to immediately destroy.
Description
A container running on the overlay filesystem (
--overlay2, defaultroot:self) stops responding to SIGKILL.runsc statereportsstoppedwith pid -1 whilerunsc psstill lists the workload processes,runsc killreturns success and nothing happens, and the shim's Wait never returns, so containerd and the kubelet retry the stop forever. The sentry itself is healthy: it answersrunsc debug --stacks, and the dump shows one goroutine that has requestedrenameMufor write while the same goroutine already holds it for read. Every lateroverlay.(*dentry).DecRef, including the ones on the task-exit path, queues behind it, so the exit can never complete.We have 43 sandboxes in this state, the oldest blocked for over 44 hours at the time of the dump.
Stack (condensed, innermost first,
runsc debug --stacks)Full dump available on request.
Analysis
Three pieces of code, unchanged between release-20260622.0, release-20260831.0 and master at d94a98a:
pkg/sentry/fsimpl/overlay/filesystem.go,doCreateAt: takesfs.renameMu.RLock()withdefer fs.renameMuRUnlockAndCheckDrop(ctx, &ds)and callsparent.watches.Notify(...)before returning, so the notification runs under the read lock.pkg/sentry/vfs/inotify.go: whenNotifyfinds watches whose inotify instance has been released,cleanupExpiredWatchescallsInotify.RmWatch, which callsw.target.OnZeroWatches(ctx)synchronously on the notifying goroutine once the target has no watches left.pkg/sentry/fsimpl/overlay/overlay.go,(*dentry).OnZeroWatches:if d.refs.Load() == 0 { d.fs.renameMu.Lock(); d.checkDropLocked(ctx); d.fs.renameMu.Unlock() }.Walked directories in the overlay normally sit at zero references (they are kept alive by the parent's child map under
renameMu, which is why the drop-list exists), so the condition holds for the parent of a create, and the write lock is requested by the goroutine that holds the read lock. Go's RWMutex never grants that.The
DentryImpl.OnZeroWatchescontract only says no inotify locks may be held by the caller; it does not say anything about filesystem locks, and the overlay calls it from a path holding its own. The gofer filesystem'sOnZeroWatcheshas the same shape throughcheckCachingLockedand its create paths also notify under the read lock, so the same hazard probably exists there; we have not observed it.Sequence that triggers it
IN_ONESHOTon a leaf directory of the overlay (a directory with no cached children, so its dentry sits at zero references) and keeps the inotify fd open.IN_ONESHOTis the only path that marks a watch expired and leaves its removal toWatches.Notify; releasing an inotify instance removes its watches synchronously with no filesystem lock held, so a watcher exiting is not a trigger.mkdirinside it (doCreateAt,renameMu.RLock) or a rename of it (RenameAt,renameMu.Lock, throughInotifyRename/IN_MOVE_SELF).Watch.Notifyfires the one-shot watch and marks it expired; the sameNotifycall runscleanupExpiredWatches -> RmWatch -> OnZeroWatches, which requestsrenameMu.Lockon the goroutine that already holdsrenameMu, and the sandbox is wedged.open(O_CREAT)in the same directory does not trigger it:getChildLockedtakes a parent reference before the notification, soOnZeroWatchesseesrefs != 0and skips the lock.Reproducer inside a container (alpine 3.21 rootfs with python3; the path must be on the overlay-backed rootfs, not under
/tmp, which runsc backs with a tmpfs whoseOnZeroWatchesis a no-op):Results with runsc built from the
gobranch at 1db01ca (release-20260831.0-42), platform systrap,runsc runwith the default--overlay2=root:self(root:memorybehaves the same):mkdirin the watched leaf dirmvof the watched leaf diropen(O_CREAT)in itOn the hung sandbox:
runsc kill -all <id> KILLnever returns (blocked inKernel.Pause, because the task insidemkdiratnever reaches a stop point);runsc kill <id> KILLreturns 0,runsc statethen reportsstopped/pid -1whilerunsc psstill lists the processes, and the init task's exit path is blocked inrunExitMain -> FDTable.DecRef -> overlay.(*dentry).DecRef -> renameMu.Lock. Onlyrunsc delete -force(SIGKILL of the sandbox process) ends it. The stacks match the production dump above frame for frame; the sentry watchdog logsSentry detected 1 stuck task(s)every 45 s and, with the default--watchdog-action=log, does nothing else.Possible fixes
renameMuRUnlockAndCheckDrop, holding a reference on the parent across the call. EveryNotifyissued whilerenameMuis held is exposed, not only the one indoCreateAt.OnZeroWatchesdefer the drop (queue the dentry for the nextrenameMuRUnlockAndCheckDrop) instead of taking the write lock inline.OnZeroWatchesfrom the expired-watch cleanup on its own goroutine and leave theinotify_rm_watchpath synchronous: overlay, gofer: defer dentry drop when Notify removes the last watch #14684, verified against the reproducer above.Environment
--overlay2default; the relevant code is identical in release-20260831.0 and master d94a98a.IN_ONESHOT(racket/src/rktio/rktio_fs_change.c) on the directories it watches (racket/src/expander/eval/collection.rkt), and the build creates directories under them.