[CUDA] MatMulBlockQuantizedFp8Weight: fold the W8A8 activation QDQ into the decode GEMV - #31481
Open
tianleiwu wants to merge 2 commits into
Open
[CUDA] MatMulBlockQuantizedFp8Weight: fold the W8A8 activation QDQ into the decode GEMV#31481tianleiwu wants to merge 2 commits into
tianleiwu wants to merge 2 commits into
Microsoft GitHub Policy Service / license/cla
succeeded
Aug 2, 2026 in 0s
All CLA requirements met.
This check verifies that the author has agreed to a CLA with Microsoft.
Loading