Request description
When a tensor.pad operation sits between a producer dispatch's output and a consumer dispatch's encoded input (e.g., convolution activations with spatial padding), the current pipeline creates a cross-dispatch boundary sequence of unpack → pad → repack. This three-operation overhead is avoidable when the padding is constant and symmetric, since the CPU can apply the padding implicitly.
Current behavior:
producer_dispatch (tiled output)
→ unset_encoding (unpack to standard layout)
→ tensor.pad (add spatial halos)
→ set_encoding (repack to tiled layout)
→ consumer_dispatch (tiled input)
Desired behavior:
producer_dispatch (tiled output, unpadded shape)
→ consumer_dispatch (reads tiled unpadded buffer)
Proposed approach:
- In SetEncoding, detect symmetric constant tensor.pad on convolution activations and encode the unpadded source instead
- Strip kernel-offset reduction addends from the convolution's activation indexing maps to match the smaller operand extent
- Make
tensor.pad hoistable and fusable with set_encoding in HoistEncodingOps and FuseEncodingOpsIntoDispatchRegions
- Add
tensor.pad to the "look-through" chain in FusionUtils::getProducerDispatchValueAndOpChain()
- Store padding metadata as attributes for the backend to program hardware padding
Eliminates per-layer unpack/repack overhead in CNN inference pipelines where every convolution layer pads its activation input.
What component(s) does this issue relate to?
Compiler
Additional context
No response
Request description
When a
tensor.padoperation sits between a producer dispatch's output and a consumer dispatch's encoded input (e.g., convolution activations with spatial padding), the current pipeline creates a cross-dispatch boundary sequence of unpack → pad → repack. This three-operation overhead is avoidable when the padding is constant and symmetric, since the CPU can apply the padding implicitly.Current behavior:
producer_dispatch (tiled output)
→ unset_encoding (unpack to standard layout)
→ tensor.pad (add spatial halos)
→ set_encoding (repack to tiled layout)
→ consumer_dispatch (tiled input)
Desired behavior:
producer_dispatch (tiled output, unpadded shape)
→ consumer_dispatch (reads tiled unpadded buffer)
Proposed approach:
tensor.padhoistable and fusable with set_encoding in HoistEncodingOps and FuseEncodingOpsIntoDispatchRegionstensor.padto the "look-through" chain inFusionUtils::getProducerDispatchValueAndOpChain()Eliminates per-layer unpack/repack overhead in CNN inference pipelines where every convolution layer pads its activation input.
What component(s) does this issue relate to?
Compiler
Additional context
No response