Skip to content

Channels last: add the common layout pipeline - #21963

Draft
rascani wants to merge 14 commits into
gh/rascani/8/headfrom
gh/rascani/9/head
Draft

Channels last: add the common layout pipeline#21963
rascani wants to merge 14 commits into
gh/rascani/8/headfrom
gh/rascani/9/head

Conversation

@rascani

@rascani rascani commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

Add ToContiguousChannelsLastPass as the shared fixed-point pipeline for
replacing selected operators with channels-last variants and composing existing
data-movement transforms.

Backends can restrict anchors and provide a node-level propagation barrier.
Reporting covers converted anchors and structural-copy counts and bytes; strict
mode rejects skipped anchors, internal copies, and unknown surviving copy sizes.

Un-skip the proof-of-concept model matrix and drive it through the real
pipeline, keeping both the contiguous and the channels-last case sets, and fold
the focused pipeline coverage into the same suite. The channels-last cases pin
that the pipeline leaves an already-channels-last graph alone.

This optimization loop is transitional. Arm's PropagateViewCopyPermuteDown/Up
is a measurably better mover -- on this same matrix it reaches 82 permutes where
this loop stalls at 156, strictly better in 9 cases and worse in none, with the
wins concentrated in LSTM/GRU/attention. The intent is to generalize that pass
out of backends/arm into backends/transforms and retire _optimization_passes()
into it.

It is not adopted yet because it cannot currently be imported outside Arm:
tosa.RESCALE/TABLE/SCATTER resolve lazily and raise, is_swappable raises rather
than returning False on keepdim != True, and Down crosses an explicit
memory_format on a pointwise-tagged clone without remapping it. It also costs
8.08s against 0.27s here on a 522-node MobileNetV2, and its fan-out invariant
has a five-op counterexample. Remove this loop once those are closed.

Keeping it in the meantime is not cosmetic: on quantized MobileNetV2 through the
explicit-layout path it is the difference between 1 cortex_m::transpose and 106,
with identical convolution counts, so the other 105 are pure NHWC copies at
224x224.

The matrix ends at 156 permutes against 142 before the pass.

Local Corstone-300 run of the full Cortex-M suite, the transforms suites, and
lintrunner.

Authored with Codex.

[ghstack-poisoned]
@pytorch-bot

pytorch-bot Bot commented Aug 20, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21963

Note: Links to docs will display an error until the docs builds have been completed.

❌ 2 New Failures, 7 Cancelled Jobs, 1 Unrelated Failure

As of commit 42ec007 with merge base b584d82 (image):

NEW FAILURES - The following jobs have failed:

CANCELLED JOBS - The following jobs were cancelled. Please retry:

FLAKY - The following job failed but was likely due to flakiness present on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
[ghstack-poisoned]
@github-actions github-actions Bot added ciflow/trunk module: arm Issues related to arm backend labels Aug 25, 2026
@rascani
rascani changed the base branch from gh/rascani/22/head to gh/rascani/8/head August 25, 2026 04:22
[ghstack-poisoned]
[ghstack-poisoned]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. module: arm Issues related to arm backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant