Skip to content

overlay: sandbox can never exit after a self-deadlock on renameMu when an inotify Notify under doCreateAt removes an expired watch #14680

Description

@iaroslav-reflection

Description

A container running on the overlay filesystem (--overlay2, default root:self) stops responding to SIGKILL. runsc state reports stopped with pid -1 while runsc ps still lists the workload processes, runsc kill returns success and nothing happens, and the shim's Wait never returns, so containerd and the kubelet retry the stop forever. The sentry itself is healthy: it answers runsc debug --stacks, and the dump shows one goroutine that has requested renameMu for write while the same goroutine already holds it for read. Every later overlay.(*dentry).DecRef, including the ones on the task-exit path, queues behind it, so the exit can never complete.

We have 43 sandboxes in this state, the oldest blocked for over 44 hours at the time of the dump.

Stack (condensed, innermost first, runsc debug --stacks)

goroutine 257515 [sync.RWMutex.Lock, 2656 minutes]:
  overlay.(*renameRWMutex).Lock
  overlay.(*dentry).OnZeroWatches          <- wants renameMu for write
  vfs.(*Inotify).RmWatch
  vfs.(*Watches).cleanupExpiredWatches
  vfs.(*Watches).Notify
  overlay.(*filesystem).doCreateAt         <- holds renameMu for read (deferred unlock)
  overlay.(*filesystem).MkdirAt
  ... mkdir(2)

2 x goroutine [sync.RWMutex.Lock, 2628 minutes]:
  overlay.(*renameRWMutex).Lock
  overlay.(*dentry).DecRef
  vfs.(*FileDescription).DecRef
  mm.(*MemoryManager).DecUsers
  kernel.(*runExitMain).execute            <- the SIGKILL exit path, blocked behind the first goroutine

goroutine 1:   kernel.(*Kernel).WaitExited        (never returns)
goroutine 161: kernel.(*ThreadGroup).WaitExited   (the shim's Wait)

Full dump available on request.

Analysis

Three pieces of code, unchanged between release-20260622.0, release-20260831.0 and master at d94a98a:

  • pkg/sentry/fsimpl/overlay/filesystem.go, doCreateAt: takes fs.renameMu.RLock() with defer fs.renameMuRUnlockAndCheckDrop(ctx, &ds) and calls parent.watches.Notify(...) before returning, so the notification runs under the read lock.
  • pkg/sentry/vfs/inotify.go: when Notify finds watches whose inotify instance has been released, cleanupExpiredWatches calls Inotify.RmWatch, which calls w.target.OnZeroWatches(ctx) synchronously on the notifying goroutine once the target has no watches left.
  • pkg/sentry/fsimpl/overlay/overlay.go, (*dentry).OnZeroWatches: if d.refs.Load() == 0 { d.fs.renameMu.Lock(); d.checkDropLocked(ctx); d.fs.renameMu.Unlock() }.

Walked directories in the overlay normally sit at zero references (they are kept alive by the parent's child map under renameMu, which is why the drop-list exists), so the condition holds for the parent of a create, and the write lock is requested by the goroutine that holds the read lock. Go's RWMutex never grants that.

The DentryImpl.OnZeroWatches contract only says no inotify locks may be held by the caller; it does not say anything about filesystem locks, and the overlay calls it from a path holding its own. The gofer filesystem's OnZeroWatches has the same shape through checkCachingLocked and its create paths also notify under the read lock, so the same hazard probably exists there; we have not observed it.

Sequence that triggers it

  1. A process adds an inotify watch with IN_ONESHOT on a leaf directory of the overlay (a directory with no cached children, so its dentry sits at zero references) and keeps the inotify fd open. IN_ONESHOT is the only path that marks a watch expired and leaves its removal to Watches.Notify; releasing an inotify instance removes its watches synchronously with no filesystem lock held, so a watcher exiting is not a trigger.
  2. Another process produces the first event on that directory: mkdir inside it (doCreateAt, renameMu.RLock) or a rename of it (RenameAt, renameMu.Lock, through InotifyRename / IN_MOVE_SELF). Watch.Notify fires the one-shot watch and marks it expired; the same Notify call runs cleanupExpiredWatches -> RmWatch -> OnZeroWatches, which requests renameMu.Lock on the goroutine that already holds renameMu, and the sandbox is wedged.
  3. open(O_CREAT) in the same directory does not trigger it: getChildLocked takes a parent reference before the notification, so OnZeroWatches sees refs != 0 and skips the lock.

Reproducer inside a container (alpine 3.21 rootfs with python3; the path must be on the overlay-backed rootfs, not under /tmp, which runsc backs with a tmpfs whose OnZeroWatches is a no-op):

mkdir -p /work/w/sub
python3 - <<'EOF' &
import ctypes, time
libc = ctypes.CDLL(None, use_errno=True)
fd = libc.inotify_init()
IN_ALL_EVENTS, IN_ONESHOT = 0x00000fff, 0x80000000
for p in (b"/work/w", b"/work/w/sub"):
    libc.inotify_add_watch(fd, p, IN_ALL_EVENTS | IN_ONESHOT)
time.sleep(3600)          # keep the fd open, never read it
EOF
sleep 1
mkdir /work/w/sub/new     # hangs forever
# or: mv /work/w/sub /work/w/sub2

Results with runsc built from the go branch at 1db01ca (release-20260831.0-42), platform systrap, runsc run with the default --overlay2=root:self (root:memory behaves the same):

binary mkdir in the watched leaf dir mv of the watched leaf dir open(O_CREAT) in it
1db01ca hangs hangs completes
1db01ca + #14684 completes in 10-20 ms, container exits 0 same completes

On the hung sandbox: runsc kill -all <id> KILL never returns (blocked in Kernel.Pause, because the task inside mkdirat never reaches a stop point); runsc kill <id> KILL returns 0, runsc state then reports stopped / pid -1 while runsc ps still lists the processes, and the init task's exit path is blocked in runExitMain -> FDTable.DecRef -> overlay.(*dentry).DecRef -> renameMu.Lock. Only runsc delete -force (SIGKILL of the sandbox process) ends it. The stacks match the production dump above frame for frame; the sentry watchdog logs Sentry detected 1 stuck task(s) every 45 s and, with the default --watchdog-action=log, does nothing else.

Possible fixes

  • Issue the inotify notification after renameMuRUnlockAndCheckDrop, holding a reference on the parent across the call. Every Notify issued while renameMu is held is exposed, not only the one in doCreateAt.
  • Or make the overlay's OnZeroWatches defer the drop (queue the dentry for the next renameMuRUnlockAndCheckDrop) instead of taking the write lock inline.
  • Or, at the vfs level, run OnZeroWatches from the expired-watch cleanup on its own goroutine and leave the inotify_rm_watch path synchronous: overlay, gofer: defer dentry drop when Notify removes the last watch #14684, verified against the reproducer above.

Environment

  • runsc release-20260622.0, platform systrap, --overlay2 default; the relevant code is identical in release-20260831.0 and master d94a98a.
  • containerd 2.2.5 with containerd-shim-runsc-v1, Kubernetes 1.35, Linux 6.12 x86_64.
  • Workload: Racket builds and tests. Racket's runtime adds its inotify watches with IN_ONESHOT (racket/src/rktio/rktio_fs_change.c) on the directories it watches (racket/src/expander/eval/collection.rkt), and the build creates directories under them.

Activity

  1. iaroslav-reflection commented on Sep 10, 2026

    @iaroslav-reflection
    ContributorAuthor

    @ayushr2 ptal 🙏

  2. konstantin-s-bogom commented on Sep 10, 2026

    @konstantin-s-bogom
    Member

    @iaroslav-reflection please sign the CLA

  3. iaroslav-reflection commented on Sep 10, 2026

    @iaroslav-reflection
    ContributorAuthor
  4. added a commit that references this issue on Sep 10, 2026
    204a703
  5. ayushr2 commented on Sep 11, 2026

    @ayushr2
    Collaborator

    @iaroslav-reflection thanks for the detailed report and reproducer.

  6. ayushr2 commented on Sep 11, 2026

    @ayushr2
    Collaborator

    I think this quirk is very specific to overlayfs, because in every other filesystem, when a child is created via doCreateAt(), the child takes a ref on the parent. But looking at overlayfs create path, it passes along the create call to upper layer dentry and doesn't construct an overlayfs child dentry which takes a ref on parent. So by the time parent.watches.Notify() is called, the parent is back to zero refs.

    mkdir inside it (doCreateAt, renameMu.RLock) or a rename of it (RenameAt, renameMu.Lock, through InotifyRename / IN_MOVE_SELF).

    Urgh, yeah it seems like it also impacts rename, because the old dir can then have no refs left on it? I suppose gofer is not impacted because it does not take the rename lock on OnZeroWatches()? It just takes some low level caching mutexes and we don't really destroy the dentry because it has an LRU cache. And overlayfs doesn't have an LRU cache so it tries to immediately destroy.

  7. ayushr2 commented on Sep 11, 2026

    @ayushr2
    Collaborator

    moving discussion to your existing PR

  8. added a commit that references this issue on Sep 11, 2026
    e841704
  9. added a commit that references this issue on Sep 18, 2026
    83ecfd3
  10. added a commit that references this issue on Sep 21, 2026
    f52be0c
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions