Walsh-Hadamard as a plain butterfly on both paths, no FFTW plans (closes #89) - #90
Merged
Merged
Conversation
Hadamard.jl's fwht_natural! built an FFTW plan under FFTW's global lock on every per-vector transform!: 170-400 us per vector at 128-4096 dimensions, worse with threads, and the one batched plan the matrix path made (#54) still executed at 1.5-12 us per column. Even a plan cached per structure executes 7-20x slower than log2(n) passes of sums and differences for this shape. So fwht! is that butterfly, with the natural ordering and the 1/n scale kept bit for bit; the matrix path runs it per column under @Batches with the copy inside the batches (a serial copyto! ahead of the loop cost three times the transform); and Hadamard.jl leaves the dependency tree. Tables on #89: 0.5-15 us per vector, 41-232 ns per column. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… in the last pass
The passes h = 1, 2, 4 have inner loops of 1, 2 and 4 elements that never vectorized; they
are now one sweep of 8-point blocks held in a register, three lane shuffles with sign masks.
The passes from h = 8 load and store Vec{8}s, and the 1/n scale rides in the last of them
instead of a sweep of its own. Contiguous Float32/Float64 vectors and column views take this
path; anything else keeps the scalar loops. Bit-identical, 4-5x faster: 0.1-4.6 us per
vector at 128-4096 dimensions, 33-800 ns per column over 64 threads (tables on #89).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #88 (its base is
sq/86-asymmetric-graph, which holdsRandomizedHadamard, the main beneficiary); GitHub retargets it tomainonce #88 merges.The problem (#89):
transform!(hp::HadamardProjection, out, v)went throughHadamard.fwht_natural!, which builds an FFTW plan under FFTW's global planning lock on every call. Per vector that is 170–400 µs at 128–4096 dimensions, and worse with threads, which serialize on the lock; the matrix path's one batched plan (#54) still executed at 1.5–12 µs per column. A plan cached per structure removes the plan and the lock but FFTW's execution of this shape (a 2×2×…×2r2r) stays at 3–54 µs,MEASUREincluded.The change:
Projections.fwht!, the iterative in-place butterfly,log2(n)passes of sums and differences and a scale by1/n, natural (Hadamard) ordering: bit for bit whatfwht_natural!produced, so every encoding built on the projection is unchanged. Bothtransform!paths use it; the matrix path runs one per column under@BATCHES, with the copy fromXinside the batches (a serialcopyto!ahead of the loop cost three times the transform).Hadamard.jlleavesProject.toml.RandomizedHadamardrotationThe Hadamard rotation is now cheaper than the QR one (~40 µs at 384/512-d), as its flops say.
Tests: the butterfly against the Sylvester matrix
H/nat 1–1024 dimensions inFloat32andFloat64,fwht!(fwht!(u)) == u/n, the matrix path equal to the per-column one for any batch size. FullPkg.test()gate green on Julia 1.12 with Aqua. Headroom noted on the issue: the first three passes have inner loops too short to vectorize.🤖 Generated with Claude Code