Skip to content

[CPU] Eliminate cross-dispatch pack/unpack overhead when tensor.pad intervenes between producer and consumer #24914

Description

@Murari007

Request description

When a tensor.pad operation sits between a producer dispatch's output and a consumer dispatch's encoded input (e.g., convolution activations with spatial padding), the current pipeline creates a cross-dispatch boundary sequence of unpack → pad → repack. This three-operation overhead is avoidable when the padding is constant and symmetric, since the CPU can apply the padding implicitly.

Current behavior:
producer_dispatch (tiled output)
→ unset_encoding (unpack to standard layout)
→ tensor.pad (add spatial halos)
→ set_encoding (repack to tiled layout)
→ consumer_dispatch (tiled input)

Desired behavior:
producer_dispatch (tiled output, unpadded shape)
→ consumer_dispatch (reads tiled unpadded buffer)

Proposed approach:

  1. In SetEncoding, detect symmetric constant tensor.pad on convolution activations and encode the unpadded source instead
  2. Strip kernel-offset reduction addends from the convolution's activation indexing maps to match the smaller operand extent
  3. Make tensor.pad hoistable and fusable with set_encoding in HoistEncodingOps and FuseEncodingOpsIntoDispatchRegions
  4. Add tensor.pad to the "look-through" chain in FusionUtils::getProducerDispatchValueAndOpChain()
  5. Store padding metadata as attributes for the backend to program hardware padding

Eliminates per-layer unpack/repack overhead in CNN inference pipelines where every convolution layer pads its activation input.

What component(s) does this issue relate to?

Compiler

Additional context

No response

Activity

  1. Murari007 commented on Sep 10, 2026

    @Murari007
    ContributorAuthor
  2. ka-snap1 commented on Sep 20, 2026

    @ka-snap1

    Hi! I'm interested in working on this issue. Is anyone already working on an implementation?

    I have C++ experience and am getting familiar with MLIR/IREE. I'd like to start by reproducing the unpack → pad → repack sequence and understanding the relevant encoding and fusion passes. Could you share a minimal MLIR reproducer, or point me to an existing test to start from?

    Also, does the proposed implicit padding support target a specific CPU backend or hardware feature? My local machine is an Intel i7-14700HX (x86_64 with AVX2), so I'd like to confirm what I can validate locally.

    Thanks for any pointers!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions