Skip to content

[CUDA] MatMulBlockQuantizedFp8Weight: fold the W8A8 activation QDQ into the decode GEMV - #31481

Merged
tianleiwu merged 2 commits into
mainfrom
tlwu/20260802/fp8_act_qdq_fold
Aug 4, 2026
Merged

[CUDA] MatMulBlockQuantizedFp8Weight: fold the W8A8 activation QDQ into the decode GEMV#31481
tianleiwu merged 2 commits into
mainfrom
tlwu/20260802/fp8_act_qdq_fold

fix(cuda): use ParseEnvironmentVariableWithDefault instead of std::ge…

d828726
Select commit
Loading
Failed to load commit list.
Microsoft GitHub Policy Service / license/cla succeeded Aug 4, 2026 in 0s

All CLA requirements met.

This check verifies that the author has agreed to a CLA with Microsoft.