Skip to content

[CUDA] Block-scaled MatMul decode GEMV: tensor cores, packed FP4 decode and M-tiling - #31155

Merged
tianleiwu merged 9 commits into
mainfrom
tlwu/20260730/block_scaled_gemv_decode
Aug 1, 2026
Merged

[CUDA] Block-scaled MatMul decode GEMV: tensor cores, packed FP4 decode and M-tiling#31155
tianleiwu merged 9 commits into
mainfrom
tlwu/20260730/block_scaled_gemv_decode

lintrunner

0d1fd02
Select commit
Loading
Failed to load commit list.
GitHub Advanced Security / lintrunner succeeded Aug 1, 2026 in 2s

No new alerts in code changed by this pull request