Skip to content

🐛 [Bug] execute_engine input checks name neither the engine nor the binding that failed #4678

Description

@SrivastavaKshitij

Bug Description

Every input check in execute_engine reports what disagreed and never which engine it
disagreed in. On a graph with more than one engine that is not enough to locate the problem,
and a refused engine writes none of its outputs, so what usually reaches the caller is an
unrelated error raised by whatever reads one of them.

core/runtime/execute_engine.cpp :: setup_input_tensors validates every input against the
binding it feeds and throws on the first disagreement:

TORCHTRT_CHECK(
    inputs[i].dtype() == binding.expected_type,
    "Expected input tensors to have type " << binding.expected_type
        << ", found type " << inputs[i].dtype());
...
TORCHTRT_CHECK(ctx->setInputShape(name.c_str(), dims), "Error while setting the input shape");

Neither message names the engine. The shape message additionally names neither the binding nor
the shape that was rejected — TensorRT itself reports both, and that report is discarded and
replaced with a fixed string:

IExecutionContext::setInputShape: Error Code 3: API Usage Error (Parameter check failed,
  condition: engineDims.d[i] == dims.d[i]. Static dimension mismatch while setting input shape
  for relu. Set dimensions are [8,17]. Expected dimensions are [8,16].)

torch.ops.tensorrt.set_logging_level(4) does not close the gap for the exception: the engine
name appears only in the LOG_DEBUG line that runs before the checks, so it is absent from
any log captured at the default level and absent from the exception in every case.

On a 14-engine model this made every runtime failure a search. These are all the diagnostics
a caller got:

Expected input tensors to have type Int, found type long int
Error while setting the input shape
CUDA out of memory. Tried to allocate 15.62 GiB.

and, when the refused engine's unwritten outputs were read further on:

a Tensor with 0 elements cannot be converted to Scalar

To Reproduce

docker run --rm --gpus all --ipc=host -v "$PWD":/w -w /w \
        nvcr.io/nvidia/pytorch:26.07-py3 python repro.py

repro.py

import torch
import torch.nn as nn
import torch_tensorrt

ROWS, COLS = 8, 16


class TwoEngines(nn.Module):
    """A graph split in two, so 'which engine?' is a real question."""

    def forward(self, x: torch.Tensor) -> torch.Tensor:
        a = x * 2.0
        b = torch.relu(a)  # pinned to Torch below, which splits the graph in two
        return torch.sigmoid(b + 1.0)


def engine_modules(gm: torch.fx.GraphModule) -> list[tuple[str, torch.nn.Module]]:
    """Returns the (name, module) pairs of ``gm``'s TensorRT submodules, in graph order."""
    return [
        (name, mod)
        for name, mod in gm.named_children()
        if getattr(mod, "serialized_engine", None) is not None
    ]


def message_for(call) -> str:
    """Runs ``call`` and returns the exception message it raises."""
    try:
        call()
    except Exception as exc:  # pylint: disable=broad-except
        return str(exc)
    raise RuntimeError("expected the call to fail, but it succeeded")


def main() -> None:
    model = TwoEngines().eval().cuda()
    x = torch.randn(ROWS, COLS, device="cuda")
    exported = torch.export.export(model, (x,))
    gm = torch_tensorrt.dynamo.compile(
        exported,
        inputs=(x,),
        min_block_size=1,
        pass_through_build_failures=True,
        # The only scaffolding here: it places an engine boundary so the graph holds two
        # engines. The diagnostics being reported are the same with one engine.
        torch_executed_ops={"torch.ops.aten.relu.default"},
    )
    engines = engine_modules(gm)
    print(f"engines in the graph: {[name for name, _ in engines]}")
    assert len(engines) > 1, "expected the partitioner to produce more than one engine"

    name, mod = engines[-1]
    mod.setup_engine()
    # The binding names are reachable only through the engine's __str__; the torchbind class
    # exposes no accessor for a binding's name, dtype or shape.
    binding_names = [b for b in str(mod.engine).split() if b.startswith("relu")]
    print(f"failing engine: {name}, input bindings: {binding_names}")

    cases = {
        "dtype check": lambda: torch.ops.tensorrt.execute_engine(
            [torch.zeros(ROWS, COLS, dtype=torch.int32, device="cuda")], mod.engine
        ),
        "shape check": lambda: torch.ops.tensorrt.execute_engine(
            [torch.zeros(ROWS, COLS + 1, device="cuda")], mod.engine
        ),
    }

    reproduced = False
    for label, call in cases.items():
        msg = message_for(call).strip()
        names_engine = name in msg
        names_binding = any(b in msg for b in binding_names)
        print(f"\n{label} raised:\n    " + msg.replace("\n", "\n    "))
        print(f"  names the engine ({name}): {names_engine}")
        print(f"  names a binding {binding_names}: {names_binding}")
        reproduced = reproduced or not names_engine

    print(f"\nreproduced: {reproduced}")
    if reproduced:
        print(
            "At least one check reports the disagreement without saying which engine it "
            "came from."
        )


if __name__ == "__main__":
    main()

output

engines in the graph: ['_run_on_acc_0', '_run_on_acc_2']
failing engine: _run_on_acc_2, input bindings: ['relu']

dtype check raised:
    [Error thrown at core/runtime/execute_engine.cpp:110] Expected inputs[i].dtype() == expected_type to be true but got false
    Expected input tensors to have type Float, found type int
  names the engine (_run_on_acc_2): False
  names a binding ['relu']: False
ERROR: [Torch-TensorRT] - IExecutionContext::setInputShape: Error Code 3: API Usage Error (Parameter check failed, condition: engineDims.d[i] == dims.d[i]. Static dimension mismatch while setting input shape for relu. Set dimensions are [8,17]. Expected dimensions are [8,16]. In setInputShape at /_src/runtime/api/executionContext.cpp:2295)

shape check raised:
    [Error thrown at core/runtime/execute_engine.cpp:149] Expected compiled_engine->exec_ctx->setInputShape(name.c_str(), dims) to be true but got false
    Error while setting the input shape
  names the engine (_run_on_acc_2): False
  names a binding ['relu']: False

reproduced: True
At least one check reports the disagreement without saying which engine it came from.

Expected behavior

Each check should name the engine and the binding, and the shape check should carry the shape
it rejected alongside the one that was declared. Concretely:

[_run_on_acc_2_engine] input 1 ("relu"): expected type Float, found type Int
[_run_on_acc_2_engine] input 1 ("relu"): rejected shape [8, 17], engine declares [8, 16]

Two related things would help alongside it:

  • The out-of-memory paths in the same function are worth catching for the same reason — a bare
    CUDA out of memory. Tried to allocate 15.62 GiB. says nothing about which engine asked for
    how much.
  • A binding's name, dtype and shape are currently reachable only by parsing the engine's
    __str__; the torchbind class exposes no accessor for them. An accessor would let callers
    diagnose this themselves.

Environment

Build information about Torch-TensorRT can be found by turning on debug messages

  • Pytorch NGC container : 26.07-py3

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions