Skip to content

Walsh-Hadamard as a plain butterfly on both paths, no FFTW plans (closes #89) - #90

Merged
sadit merged 2 commits into
mainfrom
proj/89-hadamard-butterfly
Sep 30, 2026
Merged

sadit merged 2 commits into
mainfrom
proj/89-hadamard-butterfly

Conversation

@sadit

@sadit sadit commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #88 (its base is sq/86-asymmetric-graph, which holds RandomizedHadamard, the main beneficiary); GitHub retargets it to main once #88 merges.

The problem (#89): transform!(hp::HadamardProjection, out, v) went through Hadamard.fwht_natural!, which builds an FFTW plan under FFTW's global planning lock on every call. Per vector that is 170–400 µs at 128–4096 dimensions, and worse with threads, which serialize on the lock; the matrix path's one batched plan (#54) still executed at 1.5–12 µs per column. A plan cached per structure removes the plan and the lock but FFTW's execution of this shape (a 2×2×…×2 r2r) stays at 3–54 µs, MEASURE included.

The change: Projections.fwht!, the iterative in-place butterfly, log2(n) passes of sums and differences and a scale by 1/n, natural (Hadamard) ordering: bit for bit what fwht_natural! produced, so every encoding built on the projection is unchanged. Both transform! paths use it; the matrix path runs one per column under @BATCHES, with the copy from X inside the batches (a serial copyto! ahead of the loop cost three times the transform). Hadamard.jl leaves Project.toml.

dim per vector, one thread RandomizedHadamard rotation per column, 100k columns, 64 threads
128 493 ns (was 167 µs) 565 ns 41 ns (was 1.5 µs)
512 1.9 µs (was 241 µs) 3.1 µs 123 ns (was 5.8 µs)
1024 3.7 µs (was 284 µs) 6.1 µs 232 ns (was 11.9 µs)
4096 15 µs (was 399 µs) 24 µs —

The Hadamard rotation is now cheaper than the QR one (~40 µs at 384/512-d), as its flops say.

Tests: the butterfly against the Sylvester matrix H/n at 1–1024 dimensions in Float32 and Float64, fwht!(fwht!(u)) == u/n, the matrix path equal to the per-column one for any batch size. Full Pkg.test() gate green on Julia 1.12 with Aqua. Headroom noted on the issue: the first three passes have inner loops too short to vectorize.

🤖 Generated with Claude Code

sadit and others added 2 commits September 30, 2026 11:12
Hadamard.jl's fwht_natural! built an FFTW plan under FFTW's global lock on every per-vector
transform!: 170-400 us per vector at 128-4096 dimensions, worse with threads, and the one
batched plan the matrix path made (#54) still executed at 1.5-12 us per column. Even a plan
cached per structure executes 7-20x slower than log2(n) passes of sums and differences for
this shape. So fwht! is that butterfly, with the natural ordering and the 1/n scale kept bit
for bit; the matrix path runs it per column under @Batches with the copy inside the batches
(a serial copyto! ahead of the loop cost three times the transform); and Hadamard.jl leaves
the dependency tree. Tables on #89: 0.5-15 us per vector, 41-232 ns per column.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… in the last pass

The passes h = 1, 2, 4 have inner loops of 1, 2 and 4 elements that never vectorized; they
are now one sweep of 8-point blocks held in a register, three lane shuffles with sign masks.
The passes from h = 8 load and store Vec{8}s, and the 1/n scale rides in the last of them
instead of a sweep of its own. Contiguous Float32/Float64 vectors and column views take this
path; anything else keeps the scalar loops. Bit-identical, 4-5x faster: 0.1-4.6 us per
vector at 128-4096 dimensions, 33-800 ns per column over 64 threads (tables on #89).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@sadit
sadit changed the base branch from sq/86-asymmetric-graph to main September 30, 2026 16:50
@sadit
sadit merged commit f3f25d2 into main Sep 30, 2026
2 checks passed
@sadit
sadit deleted the proj/89-hadamard-butterfly branch September 30, 2026 17:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant