Environment
- Jetson AGX Thor (sm_110), L4T R38.4 (JetPack 7.1-class), CUDA 13.0
- TensorRT-Edge-LLM 0.10.0 (bb29145)
- TensorRT tested: 10.13.3.9-1+cuda13.0 (JetPack repo) AND 10.14.1.48-1+cuda13.0 (sbsa CUDA network repo) — same failure on both
Repro
Checkpoint: NVFP4 Qwen3.8-27B (ModelOpt NVFP4, group_size 16, MTP extra tensors), e.g. https://huggingface.co/lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL
tensorrt-edgellm-export CKPT OUT --mtp → succeeds (after locally patching two exporter issues: (a) system scipy 1.8 vs numpy 2.x — from numpy import Inf ImportError when the venv uses --system-site-packages; (b) GDN fusion concatenating single-element weight_scale_2/input_scale buffers into a 4-element vector, which TRT's ONNX parser rejects: importTRT_FP4DynamicQuantize: Scale input must be a scalar — fixed by keeping numel==1 buffers unconcatenated in fuse_gdn_input_projections).
llm_build --onnxDir OUT/llm --engineDir ENG --maxBatchSize 1 --maxInputLen 4096 --maxKVCacheCapacity 32768 --maxVerifyTreeSize 4 --specBase →
[ERROR] IBuilder::buildSerializedNetwork: Error Code 1: Internal Error
(Not implemented for node type PLUGIN_V3. In removeEmptyProducersFromSubgraph
at optimizer/myelin/rewrite/removeEmptyTensors.cpp:257) // TRT 10.13
... removeEmptyTensors.cpp:306 // TRT 10.14
- The MTP draft engine (
--specDraft on OUT/mtp_draft) builds fine on both TRT versions. The draft ONNX contains only trt_edgellm::AttentionPlugin x1; the base ONNX additionally has trt_edgellm::gated_delta_net x48 and trt_edgellm::causal_conv1d x48, so the GDN plugins appear to be the PLUGIN_V3 nodes Myelin cannot handle.
- The experimental direct (ONNX-less) builder fails identically at the same Myelin pass, so this is not ONNX-path specific.
Questions
- Which TRT version was Qwen3.8-27B Day-0 support validated against on Thor? Does it require JetPack 7.2 (CUDA 13.2 / TRT 10.16)?
- Is there a builder flag/env to route GDN plugin nodes around the failing Myelin rewrite pass?
Environment
Repro
Checkpoint: NVFP4 Qwen3.8-27B (ModelOpt NVFP4, group_size 16, MTP extra tensors), e.g. https://huggingface.co/lyf/Qwen3.8-27B-Heretic-ARA-NVFP4-MTP-VL
tensorrt-edgellm-export CKPT OUT --mtp→ succeeds (after locally patching two exporter issues: (a) system scipy 1.8 vs numpy 2.x —from numpy import InfImportError when the venv uses--system-site-packages; (b) GDN fusion concatenating single-elementweight_scale_2/input_scalebuffers into a 4-element vector, which TRT's ONNX parser rejects:importTRT_FP4DynamicQuantize: Scale input must be a scalar— fixed by keeping numel==1 buffers unconcatenated infuse_gdn_input_projections).llm_build --onnxDir OUT/llm --engineDir ENG --maxBatchSize 1 --maxInputLen 4096 --maxKVCacheCapacity 32768 --maxVerifyTreeSize 4 --specBase→--specDrafton OUT/mtp_draft) builds fine on both TRT versions. The draft ONNX contains onlytrt_edgellm::AttentionPlugin x1; the base ONNX additionally hastrt_edgellm::gated_delta_net x48andtrt_edgellm::causal_conv1d x48, so the GDN plugins appear to be the PLUGIN_V3 nodes Myelin cannot handle.Questions