Channels last: add the common layout pipeline - #21963
Draft
rascani wants to merge 14 commits into
Draft
Conversation
Contributor
Author
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21963
Note: Links to docs will display an error until the docs builds have been completed. ❌ 2 New Failures, 7 Cancelled Jobs, 1 Unrelated FailureAs of commit 42ec007 with merge base b584d82 ( NEW FAILURES - The following jobs have failed:
CANCELLED JOBS - The following jobs were cancelled. Please retry:
FLAKY - The following job failed but was likely due to flakiness present on trunk:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This was referenced Aug 20, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add ToContiguousChannelsLastPass as the shared fixed-point pipeline for
replacing selected operators with channels-last variants and composing existing
data-movement transforms.
Backends can restrict anchors and provide a node-level propagation barrier.
Reporting covers converted anchors and structural-copy counts and bytes; strict
mode rejects skipped anchors, internal copies, and unknown surviving copy sizes.
Un-skip the proof-of-concept model matrix and drive it through the real
pipeline, keeping both the contiguous and the channels-last case sets, and fold
the focused pipeline coverage into the same suite. The channels-last cases pin
that the pipeline leaves an already-channels-last graph alone.
This optimization loop is transitional. Arm's PropagateViewCopyPermuteDown/Up
is a measurably better mover -- on this same matrix it reaches 82 permutes where
this loop stalls at 156, strictly better in 9 cases and worse in none, with the
wins concentrated in LSTM/GRU/attention. The intent is to generalize that pass
out of backends/arm into backends/transforms and retire _optimization_passes()
into it.
It is not adopted yet because it cannot currently be imported outside Arm:
tosa.RESCALE/TABLE/SCATTER resolve lazily and raise, is_swappable raises rather
than returning False on keepdim != True, and Down crosses an explicit
memory_format on a pointwise-tagged clone without remapping it. It also costs
8.08s against 0.27s here on a 522-node MobileNetV2, and its fan-out invariant
has a five-op counterexample. Remove this loop once those are closed.
Keeping it in the meantime is not cosmetic: on quantized MobileNetV2 through the
explicit-layout path it is the difference between 1 cortex_m::transpose and 106,
with identical convolution counts, so the other 105 are pure NHWC copies at
224x224.
The matrix ends at 156 permutes against 142 before the pass.
Local Corstone-300 run of the full Cortex-M suite, the transforms suites, and
lintrunner.
Authored with Codex.