Skip to content

Aks 2 - #933

Closed
kpouget wants to merge 16 commits into
openshift-psap:mainfrom
kpouget:aks-2
Closed

Aks 2#933
kpouget wants to merge 16 commits into
openshift-psap:mainfrom
kpouget:aks-2

Conversation

@kpouget

@kpouget kpouget commented May 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary by CodeRabbit

  • New Features

    • Support for local hostPath models with explicit resource sizing and hostPath cache mounting.
    • Tensor-parallelism handling with coordinated GPU resources.
    • InfiniBand/AKS support and appending router tolerations.
    • Capture of HTTPRoute routing artifacts for diagnostics.
  • Chores

    • Updated CI/test presets and changed default test flavor.
    • Refined inference-service scheduling, preload, and model configuration defaults.
    • Updated visualization report generation entries.

Review Change Stack

@openshift-ci

openshift-ci Bot commented May 20, 2026

Copy link
Copy Markdown

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by:
Once this PR has been reviewed and has the lgtm label, please assign dagrayvid for approval. For more information see the Code Review Process.

The full list of commands accepted by this bot can be found here.

Details Needs approval from an approver in each of these files:

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

@coderabbitai

coderabbitai Bot commented May 20, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Adds hostpath model support and hostPath HF cache wiring; centralizes tensor-parallelism and vLLM args; introduces InfiniBand/AKS hotfix and router toleration wiring; updates ISVC scheduler/nodeSelector manifests; captures HTTPRoute and benchmark artifacts; reorganizes visualization reports and plotting logic.

Changes

Hostpath Model Support & Infrastructure Configuration

Layer / File(s) Summary
Test Configuration & Model Definitions
projects/llm-d/testing/config.yaml
CI presets extended with flavor/preset reworking, preload configuration updated with InfiniBand settings and extra_images, local hostpath model entries added (llama3-3-70b-local, llama3-2-3b-local, gpt-oss-120-local), default flavor switched to pd-x2-ptp4-px1-dtp4, and inference-service EPP adds router.tolerations, infiniband: null, aks_hotfix_enabled, and hybrid_kv_cache_manager: false.
Prepare script & Tensor Parallelism
projects/llm-d/testing/prepare_llmd.py, projects/llm-d/testing/test_llmd.py
download_single_model skips PVC download for hostpath: sources; add_tensor_parallelism_to_container centralizes TP env/arg injection and GPU resource sizing; flavor/prefill tensor-parallelism functions delegate to the helper.
Hostpath model wiring & vLLM args
projects/llm-d/testing/test_llmd.py
apply_model_configuration(isvc_data, flavor) supports hostpath: sources by setting spec.model.uri, calling apply_hostpath_volume_configuration to mount an hf-cache hostPath and set HF_HUB_CACHE, and replacing vLLM container commands with vllm serve that consume VLLM_ADDITIONAL_ARGS; apply_vllm_args_configuration computes and applies final vLLM args.
InfiniBand & AKS hotfix configuration
projects/llm-d/testing/test_llmd.py
apply_infiniband_configuration applies RDMA/IB or custom resource keys to decode/prefill containers only for pd flavors; apply_infiniband_aks_configuration injects AKS UCX/NIC/NVSHMEM envs, adds IPC_LOCK, optionally validates/mounts vllm-ucx-multiproc-hotfix ConfigMap files, and sets containerd ulimit annotations.
ISVC Manifests & Router Reshape
projects/llm-d/testing/llmisvcs/llmisvc-pd.yaml, projects/llm-d/testing/llmisvcs/llmisvc-simple.yaml, projects/llm-d/testing/test_llmd.py
Removed embedded scheduler command/args in pd router template; simple manifest nodeSelector updated to nvidia.com/gpu.deploy.container-toolkit: "true"; apply_router_configuration appends router scheduler tolerations; reshape_isvc reorders vLLM args/max-model-len → model configuration → InfiniBand/router → extra properties/EPP.
Observability & Visualization Reorganization
projects/llm-d/toolbox/llmd_capture_isvc_state/tasks/main.yml, projects/llm-d/toolbox/llmd_deploy_llm_inference_service/tasks/main.yml, projects/llm-d/toolbox/llmd_run_guidellm_benchmark/tasks/main.yml, projects/llm-d/visualizations/llmd_inference/data/plots.yaml, projects/llm-d/visualizations/llmd_inference/data/reports.yaml, projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py
Ansible tasks capture HTTPRoute YAML and status artifacts; benchmark artifact copy now checks for file existence and retries; plots.yaml had three report entries removed and reports.yaml adds a visualize.llmd_inference.generate list with four report entries; throughput comparisons logic adjusted for platform and excludes simple-tp4-x4 without skip note.

Estimated Code Review Effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly Related PRs

Poem

🐰 I hopped to the hostpath, left PVC behind,
Threads of InfiniBand stitched low-latency kind,
Tensors aligned by a helper so neat,
ISVC tunes tolerations and serves on a seat,
HTTPRoute saved — deployment feels fine!

🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (1 warning, 1 inconclusive)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 78.13% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
Title check ❓ Inconclusive The PR title 'Aks 2' is vague and does not clearly convey the main changes; it lacks descriptive detail about what was actually modified. Revise the title to be more descriptive, such as 'Add AKS and hybrid KV cache manager configurations' or similar, to clearly indicate the primary changes.
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@kpouget

kpouget commented May 20, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 intelligentrouting-flavors guidellm_multiturn_eval llama-70b
/cluster aks-h100
/only test_ci

@psap-forge-bot

Copy link
Copy Markdown

🟢 Test of 'llm-d test test_ci' succeeded after 00 hours 57 minutes 49 seconds. 🟢

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: intelligentrouting-flavors
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: llama-70b

@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 pd-flavors guidellm_multiturn_eval gpt-oss
/cluster aks-h100
/only test_ci

@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss
/cluster aks-h100
/only test_ci

@kpouget

kpouget commented May 21, 2026 •

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 pd-flavors guidellm_multiturn_eval gpt-oss aks_ib
/cluster aks-h100
/only test_ci

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 00 hours 08 minutes 01 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
PR_POSITIONAL_ARG_5: aks_ib

Failure indicator:

/tmp/topsail_202605211779349863/001__llm_d_testing/000__flavor_pd-x2-ptp1-px4-dtp4/000__llmd__deploy_llm_inference_service/FAILURE | [000__llmd__deploy_llm_inference_service] ./run_toolbox.py llmd deploy_llm_inference_service --name=llm-d-pd-x2-ptp1-px4-dtp4 --namespace=kpouget-dev --yaml_file=/tmp/topsail_202605211779349863/001__llm_d_testing/000__flavor_pd-x2-ptp1-px4-dtp4/llmisvc-pd.yaml --> 2


@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss aks_ib
/cluster aks-h100
/only test_ci

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
projects/llm-d/testing/test_llmd.py (1)

760-762: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Apply vLLM args to prefill containers as well.

Hostpath flow later expects VLLM_ADDITIONAL_ARGS in prefill, but this function only populates main. That can fail at runtime for PD hostpath runs when prefill has no max-model-len override.

Suggested fix
 def apply_vllm_args_configuration(isvc_data):
@@
-    # Apply to main container only
+    # Apply to main container
     _apply_vllm_args_to_container_section(isvc_data, 'spec.template.containers', vllm_args, 'main')
+
+    # Apply to prefill container for P/D deployments
+    if 'spec' in isvc_data and 'prefill' in isvc_data['spec']:
+        _apply_vllm_args_to_container_section(isvc_data, 'spec.prefill.template.containers', vllm_args, 'main')
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 760 - 762, The code only
applies vLLM args to the main container; call
_apply_vllm_args_to_container_section for the prefill containers as well so
VLLM_ADDITIONAL_ARGS is populated there (e.g., add a second call like
_apply_vllm_args_to_container_section(isvc_data,
'spec.template.prefill.containers', vllm_args, 'prefill') or the correct path
for prefill containers in this manifest) to ensure prefill gets the
max-model-len override.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 987-1009: The code assumes spec.prefill exists and that each
container has resources->requests/limits initialized; update the logic in
apply_to_containers and the call sites: before modifying a container ensure
container.setdefault('resources', {}) and then
container['resources'].setdefault('limits', {}) and .setdefault('requests', {}),
use those dicts when popping/setting rdma keys, and guard applying to prefill by
checking if 'prefill' in isvc_data['spec'] (only call apply_to_containers on
isvc_data['spec']['prefill']['template']['containers'] when present) while
preserving the existing infiniband_config handling in apply_to_containers.
- Around line 718-725: The code is overwriting existing
container['volumeMounts'] and isvc_data['spec']['template']['volumes'], which
can drop previously defined mounts/volumes; update the logic to merge rather
than replace: for container use a safe-get/ensure pattern (e.g.,
container.setdefault('volumeMounts', []) or check for existing list) and append
hf_cache_mount only if not already present, likewise ensure container['env'] is
initialized (already done) and for the main template use
isvc_data['spec']['template'].setdefault('volumes', []) or merge with the
existing list and append hf_cache_volume if missing so existing mounts/volumes
are preserved (refer to symbols container, hf_cache_mount, isvc_data,
hf_cache_volume, volumeMounts, volumes).

---

Outside diff comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 760-762: The code only applies vLLM args to the main container;
call _apply_vllm_args_to_container_section for the prefill containers as well so
VLLM_ADDITIONAL_ARGS is populated there (e.g., add a second call like
_apply_vllm_args_to_container_section(isvc_data,
'spec.template.prefill.containers', vllm_args, 'prefill') or the correct path
for prefill containers in this manifest) to ensure prefill gets the
max-model-len override.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 2b075169-d39a-46fb-ab08-67cc1d36a76e

📥 Commits

Reviewing files that changed from the base of the PR and between cfc3e8a and 7edd510.

📒 Files selected for processing (10)
  • projects/llm-d/testing/config.yaml
  • projects/llm-d/testing/llmisvcs/llmisvc-pd.yaml
  • projects/llm-d/testing/llmisvcs/llmisvc-simple.yaml
  • projects/llm-d/testing/prepare_llmd.py
  • projects/llm-d/testing/test_llmd.py
  • projects/llm-d/toolbox/llmd_capture_isvc_state/tasks/main.yml
  • projects/llm-d/toolbox/llmd_deploy_llm_inference_service/tasks/main.yml
  • projects/llm-d/visualizations/llmd_inference/data/plots.yaml
  • projects/llm-d/visualizations/llmd_inference/data/reports.yaml
  • projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py
💤 Files with no reviewable changes (2)
  • projects/llm-d/visualizations/llmd_inference/data/plots.yaml
  • projects/llm-d/testing/llmisvcs/llmisvc-pd.yaml

Comment thread projects/llm-d/testing/test_llmd.py
Comment on lines +987 to +1009
def apply_to_containers(containers):
for container in containers:
# Always remove existing rdma/ib first
container['resources']['limits'].pop('rdma/ib', None)
container['resources']['requests'].pop('rdma/ib', None)

# Add InfiniBand resource based on config type
if infiniband_config is True:
container['resources']['limits']['rdma/ib'] = "1"
container['resources']['requests']['rdma/ib'] = "1"
elif isinstance(infiniband_config, str):
# Use custom resource string (e.g., "rdma/shared_ib")
container['resources']['limits'][infiniband_config] = "1"
container['resources']['requests'][infiniband_config] = "1"

# Apply to main template containers
apply_to_containers(isvc_data['spec']['template']['containers'])

# Apply to prefill template containers (for P/D deployments)
if 'prefill' not in isvc_data['spec']:
raise ValueError("Trying to apply infiniband config without a prefill deployment")

apply_to_containers(isvc_data['spec']['prefill']['template']['containers'])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major | ⚡ Quick win

InfiniBand configuration incorrectly assumes prefill exists and resources are pre-initialized.

This will fail on non-PD ISVCs (no spec.prefill) and on containers missing resources.requests/limits.

Suggested fix
     def apply_to_containers(containers):
         for container in containers:
+            container.setdefault('resources', {})
+            container['resources'].setdefault('limits', {})
+            container['resources'].setdefault('requests', {})
             # Always remove existing rdma/ib first
             container['resources']['limits'].pop('rdma/ib', None)
             container['resources']['requests'].pop('rdma/ib', None)
@@
-    if 'prefill' not in isvc_data['spec']:
-        raise ValueError("Trying to apply infiniband config without a prefill deployment")
-
-    apply_to_containers(isvc_data['spec']['prefill']['template']['containers'])
+    if 'prefill' in isvc_data['spec']:
+        apply_to_containers(isvc_data['spec']['prefill']['template']['containers'])
-    if 'annotations' not in isvc_data['spec']['prefill']:
-        isvc_data['spec']['prefill']['annotations'] = {}
-    isvc_data['spec']['prefill']['annotations']['ulimits.nri.containerd.io/container.main'] = ulimit_annotation_value
+    if 'prefill' in isvc_data['spec']:
+        if 'annotations' not in isvc_data['spec']['prefill']:
+            isvc_data['spec']['prefill']['annotations'] = {}
+        isvc_data['spec']['prefill']['annotations']['ulimits.nri.containerd.io/container.main'] = ulimit_annotation_value

Also applies to: 1148-1151

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 987 - 1009, The code
assumes spec.prefill exists and that each container has
resources->requests/limits initialized; update the logic in apply_to_containers
and the call sites: before modifying a container ensure
container.setdefault('resources', {}) and then
container['resources'].setdefault('limits', {}) and .setdefault('requests', {}),
use those dicts when popping/setting rdma keys, and guard applying to prefill by
checking if 'prefill' in isvc_data['spec'] (only call apply_to_containers on
isvc_data['spec']['prefill']['template']['containers'] when present) while
preserving the existing infiniband_config handling in apply_to_containers.

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 00 hours 07 minutes 39 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
PR_POSITIONAL_ARG_5: aks_ib

Failure indicator:

/tmp/topsail_202605211779357343/001__llm_d_testing/000__flavor_pd-x2-ptp2-px2-dtp4/000__llmd__deploy_llm_inference_service/FAILURE | [000__llmd__deploy_llm_inference_service] ./run_toolbox.py llmd deploy_llm_inference_service --name=llm-d-pd-x2-ptp2-px2-dtp4 --namespace=kpouget-dev --yaml_file=/tmp/topsail_202605211779357343/001__llm_d_testing/000__flavor_pd-x2-ptp2-px2-dtp4/llmisvc-pd.yaml --> 2


@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss aks_ib
/cluster aks-h100
/only test_ci
/var tests.llmd.flavors: pd-x2-ptp4-px1-dtp4

@psap-forge-bot

Copy link
Copy Markdown

🟢 Test of 'llm-d test test_ci' succeeded after 00 hours 25 minutes 10 seconds. 🟢

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
PR_POSITIONAL_ARG_5: aks_ib
tests.llmd.flavors: pd-x2-ptp4-px1-dtp4

@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss aks_ib
/cluster aks-h100
/only test_ci
/var tests.llmd.flavors: simple-tp4-x4

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 00 hours 00 minutes 40 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
PR_POSITIONAL_ARG_5: aks_ib
tests.llmd.flavors: simple-tp4-x4

Failure indicator:

/tmp/topsail_202605211779367538/002__plots/FAILURE | An error happened during the visualization post-processing ... (0_matbench_parse.log in /tmp/topsail_202605211779367538/002__plots). Mind that the test that was processed FAILED.


@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss
/cluster aks-h100
/only test_ci
/var tests.llmd.flavors: simple-tp4-x4

@psap-forge-bot

Copy link
Copy Markdown

🟢 Test of 'llm-d test test_ci' succeeded after 00 hours 46 minutes 57 seconds. 🟢

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
tests.llmd.flavors: simple-tp4-x4

@kpouget

kpouget commented May 21, 2026 •

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss aks_ib
/cluster aks-h100
/only test_ci
/var tests.llmd.inference_service.aks_hotfix_enabled: false

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
projects/llm-d/testing/test_llmd.py (1)

743-770: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

apply_vllm_args_configuration drops user vllm_args on the prefill container.

The docstring says "Applies VLLM args to both main container and prefill container (for P/D deployments)", but the implementation only writes to spec.template.containers (main). For PD flavors, the prefill container will therefore not receive --gpu-memory-utilization=0.92, --enable-prefix-caching, --trust-remote-code, --no-disable-hybrid-kv-cache-manager, etc., even though apply_max_model_len_configuration and apply_prefill_tensor_parallelism do target prefill. This is especially impactful for the new hostpath branch, where configure_vllm_command consumes VLLM_ADDITIONAL_ARGS into the prefill command line — those user-configured flags will silently be missing from prefill while present on decode, leading to asymmetric vLLM behavior between P and D and skewed benchmark numbers.

♻️ Suggested fix: also apply to prefill in PD deployments
     logging.info(f"Applying vLLM args: {final_vllm_args}")

     # Apply to main container only
     _apply_vllm_args_to_container_section(isvc_data, 'spec.template.containers', final_vllm_args, 'main')
+
+    # Apply to prefill container if this is a P/D deployment
+    if 'spec' in isvc_data and 'prefill' in isvc_data['spec']:
+        logging.info("P/D deployment detected - applying vLLM args to prefill container")
+        _apply_vllm_args_to_container_section(isvc_data, 'spec.prefill.template.containers', final_vllm_args, 'main')
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 743 - 770,
apply_vllm_args_configuration currently only applies vLLM args to the main
container (via _apply_vllm_args_to_container_section with
'spec.template.containers'/'main') which drops user args for prefill in PD
deployments; update the function to also call
_apply_vllm_args_to_container_section for the prefill container (use the same
final_vllm_args and a container selector like 'spec.template.prefill.containers'
and container name 'prefill' to mirror how apply_max_model_len_configuration and
apply_prefill_tensor_parallelism target prefill), keeping the existing logging
behavior so both main and prefill receive identical VLLM flags.
🧹 Nitpick comments (1)
projects/llm-d/testing/test_llmd.py (1)

1066-1073: 💤 Low value

Stray triple-quoted string in function body.

This block isn't a docstring (the function already has one on line 1022) and isn't assigned to anything — Python evaluates and discards it. It also lists a slightly different set of files than the actual hotfix_files list (interface.py instead of platforms/interface.py, multiproc_executor.py instead of v1/executor/multiproc_executor.py, etc.), so it will drift further from reality over time. Either convert it to a # comment or drop it since the runtime error at line 1058 already produces the correct oc create cm … invocation.

♻️ Suggested cleanup
-    """
-    oc create cm vllm-ucx-multiproc-hotfix  \
-       --from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/envs.py \
-       --from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/interface.py \
-       --from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/multiproc_executor.py \
-       --from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/uniproc_executor.py \
-       --from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/vllm_net_devices.py
-    """
-
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 1066 - 1073, Remove the
stray triple-quoted string block inside the function (the unassigned multiline
literal shown between lines ~1066–1073) because it's not a docstring and
duplicates outdated file names; either delete it entirely or convert it to a
simple # comment if you want to keep the example, and ensure any example matches
the actual hotfix_files list used elsewhere (see hotfix_files and the runtime
error/oc invocation at line ~1058) so the source of truth remains the
hotfix_files variable rather than this dead literal.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 743-770: apply_vllm_args_configuration currently only applies vLLM
args to the main container (via _apply_vllm_args_to_container_section with
'spec.template.containers'/'main') which drops user args for prefill in PD
deployments; update the function to also call
_apply_vllm_args_to_container_section for the prefill container (use the same
final_vllm_args and a container selector like 'spec.template.prefill.containers'
and container name 'prefill' to mirror how apply_max_model_len_configuration and
apply_prefill_tensor_parallelism target prefill), keeping the existing logging
behavior so both main and prefill receive identical VLLM flags.

---

Nitpick comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 1066-1073: Remove the stray triple-quoted string block inside the
function (the unassigned multiline literal shown between lines ~1066–1073)
because it's not a docstring and duplicates outdated file names; either delete
it entirely or convert it to a simple # comment if you want to keep the example,
and ensure any example matches the actual hotfix_files list used elsewhere (see
hotfix_files and the runtime error/oc invocation at line ~1058) so the source of
truth remains the hotfix_files variable rather than this dead literal.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: e4eada77-3ebc-406b-8ac1-23baedcb68fb

📥 Commits

Reviewing files that changed from the base of the PR and between 7edd510 and 093777a.

📒 Files selected for processing (2)
  • projects/llm-d/testing/config.yaml
  • projects/llm-d/testing/test_llmd.py

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 01 hours 12 minutes 50 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
tests.llmd.inference_service.aks_hotfix_enabled: false

Failure indicator:

/tmp/topsail_202605211779391389/001__llm_d_testing/000__flavor_pd-x2-ptp2-px2-dtp4/000__llmd__deploy_llm_inference_service/FAILURE | [000__llmd__deploy_llm_inference_service] ./run_toolbox.py llmd deploy_llm_inference_service --name=llm-d-pd-x2-ptp2-px2-dtp4 --namespace=kpouget-dev --yaml_file=/tmp/topsail_202605211779391389/001__llm_d_testing/000__flavor_pd-x2-ptp2-px2-dtp4/llmisvc-pd.yaml --> 2


@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss aks_ib
/cluster aks-h100
/only test_ci
/var tests.llmd.inference_service.aks_hotfix_enabled: false

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 00 hours 04 minutes 49 seconds. 🔴

• Link to the test results.

• No reports index generated...

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
PR_POSITIONAL_ARG_5: aks_ib
tests.llmd.inference_service.aks_hotfix_enabled: false

Failure indicator: Empty. (See run.log)

@kpouget

kpouget commented May 21, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 llmisvc_pd guidellm_multiturn_eval gpt-oss aks_ib
/cluster aks-h100
/only test_ci
/var tests.llmd.inference_service.aks_hotfix_enabled: false

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 00 hours 07 minutes 38 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: llmisvc_pd
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss
PR_POSITIONAL_ARG_5: aks_ib
tests.llmd.inference_service.aks_hotfix_enabled: false

Failure indicator:

/tmp/topsail_202605211779398438/001__llm_d_testing/000__flavor_pd-x2-ptp2-px2-dtp4/000__llmd__deploy_llm_inference_service/FAILURE | [000__llmd__deploy_llm_inference_service] ./run_toolbox.py llmd deploy_llm_inference_service --name=llm-d-pd-x2-ptp2-px2-dtp4 --namespace=kpouget-dev --yaml_file=/tmp/topsail_202605211779398438/001__llm_d_testing/000__flavor_pd-x2-ptp2-px2-dtp4/llmisvc-pd.yaml --> 2


@kpouget

kpouget commented May 22, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 pd-flavors guidellm_multiturn_eval gpt-oss
/cluster aks-h100
/only test_ci

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (2)
projects/llm-d/testing/test_llmd.py (2)

1202-1206: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Unconditional prefill access will fail if prefill doesn't exist.

Lines 1203-1205 access isvc_data['spec']['prefill']['annotations'] without checking if prefill exists, unlike the guarded access patterns at lines 1091 and 1184. If a malformed pd YAML lacks prefill, this crashes with KeyError before the parent function's more helpful ValueError at line 1030.

Proposed fix to guard prefill access
     # Set annotations for prefill pods (spec.prefill.annotations)
-    if 'annotations' not in isvc_data['spec']['prefill']:
-        isvc_data['spec']['prefill']['annotations'] = {}
-    isvc_data['spec']['prefill']['annotations']['ulimits.nri.containerd.io/container.main'] = ulimit_annotation_value
+    if 'prefill' in isvc_data['spec']:
+        if 'annotations' not in isvc_data['spec']['prefill']:
+            isvc_data['spec']['prefill']['annotations'] = {}
+        isvc_data['spec']['prefill']['annotations']['ulimits.nri.containerd.io/container.main'] = ulimit_annotation_value
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 1202 - 1206, The code
unconditionally indexes isvc_data['spec']['prefill'] causing KeyError if
'prefill' is absent; update the block that sets ulimits.nri annotation to first
ensure 'prefill' exists (e.g., if 'prefill' not in isvc_data['spec']:
isvc_data['spec']['prefill'] = {}) and then ensure annotations exists before
assigning
isvc_data['spec']['prefill']['annotations']['ulimits.nri.containerd.io/container.main']
= ulimit_annotation_value so the code mirrors the guarded access used elsewhere
and avoids KeyError.

1010-1024: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

InfiniBand configuration assumes resources are pre-initialized.

apply_to_containers accesses container['resources']['limits'] and container['resources']['requests'] directly, but these may not exist if tensor parallelism was not configured (e.g., tp_size is None).

Proposed fix to initialize resources safely
     def apply_to_containers(containers):
         for container in containers:
+            container.setdefault('resources', {})
+            container['resources'].setdefault('limits', {})
+            container['resources'].setdefault('requests', {})
             # Always remove existing rdma/ib first
             container['resources']['limits'].pop('rdma/ib', None)
             container['resources']['requests'].pop('rdma/ib', None)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 1010 - 1024, The
apply_to_containers function assumes container['resources']['limits'] and
['requests'] exist; adjust it to safely initialize missing dicts before use:
ensure container has a 'resources' dict and that 'limits' and 'requests' are
dicts (create empty dicts if absent) prior to popping or assigning the rdma
keys, then proceed with the existing logic that removes any prior 'rdma/ib' and
sets the resource based on infiniband_config (True => use 'rdma/ib', str => use
that string).
🧹 Nitpick comments (2)
projects/llm-d/testing/test_llmd.py (2)

1124-1125: 💤 Low value

Use exception chaining when re-raising.

Raising a new exception without chaining loses the original traceback. Use raise ... from e to preserve context.

Proposed fix
         except Exception as e:
-            raise RuntimeError(f"Failed to verify ConfigMap '{cm_name}' in namespace '{namespace}': {e}")
+            raise RuntimeError(f"Failed to verify ConfigMap '{cm_name}' in namespace '{namespace}': {e}") from e
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 1124 - 1125, The except
block that currently does "raise RuntimeError(f\"Failed to verify ConfigMap
'{cm_name}' in namespace '{namespace}': {e}\")" should preserve the original
traceback by using exception chaining; change the re-raise to "raise
RuntimeError(... ) from e" so the original exception is chained to the new
RuntimeError (referencing the cm_name and namespace variables in the error
message).

1157-1166: 💤 Low value

Dead code: add_hotfix_volumes is defined but never called.

The function add_hotfix_volumes (lines 1157-1166) is defined but never used. Instead, add_hotfix_to_containers (lines 1169-1177) performs the same logic and is actually called. Remove the dead function.

Proposed fix to remove dead code
-        # Helper function to add hotfix volume mounts to containers
-        def add_hotfix_volumes(containers, template_volumes):
-            for container in containers:
-                # Add volume mounts
-                if 'volumeMounts' not in container:
-                    container['volumeMounts'] = []
-                container['volumeMounts'].extend(hotfix_mounts)
-
-            # Add volume to template
-            template_volumes.append(hotfix_volume)
-
         # Apply hotfix volumes and complete configuration to containers
         def add_hotfix_to_containers(containers, template_volumes):
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 1157 - 1166, Remove the
dead helper function add_hotfix_volumes which duplicates logic already
implemented and used in add_hotfix_to_containers; delete the entire
add_hotfix_volumes definition (including its loop and the
template_volumes.append(hotfix_volume) line) so only add_hotfix_to_containers
manipulates container['volumeMounts'], hotfix_mounts and
template_volumes/hotfix_volume remain in use.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Duplicate comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 1202-1206: The code unconditionally indexes
isvc_data['spec']['prefill'] causing KeyError if 'prefill' is absent; update the
block that sets ulimits.nri annotation to first ensure 'prefill' exists (e.g.,
if 'prefill' not in isvc_data['spec']: isvc_data['spec']['prefill'] = {}) and
then ensure annotations exists before assigning
isvc_data['spec']['prefill']['annotations']['ulimits.nri.containerd.io/container.main']
= ulimit_annotation_value so the code mirrors the guarded access used elsewhere
and avoids KeyError.
- Around line 1010-1024: The apply_to_containers function assumes
container['resources']['limits'] and ['requests'] exist; adjust it to safely
initialize missing dicts before use: ensure container has a 'resources' dict and
that 'limits' and 'requests' are dicts (create empty dicts if absent) prior to
popping or assigning the rdma keys, then proceed with the existing logic that
removes any prior 'rdma/ib' and sets the resource based on infiniband_config
(True => use 'rdma/ib', str => use that string).

---

Nitpick comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 1124-1125: The except block that currently does "raise
RuntimeError(f\"Failed to verify ConfigMap '{cm_name}' in namespace
'{namespace}': {e}\")" should preserve the original traceback by using exception
chaining; change the re-raise to "raise RuntimeError(... ) from e" so the
original exception is chained to the new RuntimeError (referencing the cm_name
and namespace variables in the error message).
- Around line 1157-1166: Remove the dead helper function add_hotfix_volumes
which duplicates logic already implemented and used in add_hotfix_to_containers;
delete the entire add_hotfix_volumes definition (including its loop and the
template_volumes.append(hotfix_volume) line) so only add_hotfix_to_containers
manipulates container['volumeMounts'], hotfix_mounts and
template_volumes/hotfix_volume remain in use.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 0159a356-3d57-4ad3-92a4-a299d8d5c8e5

📥 Commits

Reviewing files that changed from the base of the PR and between 093777a and 8ae2b7d.

📒 Files selected for processing (4)
  • projects/llm-d/testing/config.yaml
  • projects/llm-d/testing/test_llmd.py
  • projects/llm-d/visualizations/llmd_inference/data/plots.yaml
  • projects/llm-d/visualizations/llmd_inference/data/reports.yaml
💤 Files with no reviewable changes (1)
  • projects/llm-d/visualizations/llmd_inference/data/plots.yaml

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 00 hours 32 minutes 23 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: pd-flavors
PR_POSITIONAL_ARG_3: guidellm_multiturn_eval
PR_POSITIONAL_ARG_4: gpt-oss

Failure indicator:

/tmp/topsail_202605221779436538/001__llm_d_testing/001__flavor_simple-tp4-x4/000__llmd__deploy_llm_inference_service/FAILURE | [000__llmd__deploy_llm_inference_service] ./run_toolbox.py llmd deploy_llm_inference_service --name=llm-d-simple-tp4-x4 --namespace=kpouget-dev --yaml_file=/tmp/topsail_202605221779436538/001__llm_d_testing/001__flavor_simple-tp4-x4/llmisvc-simple.yaml --> 2


@kpouget

kpouget commented May 22, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 pd-flavors guidellm_heterogeneous_eval gpt-oss
/cluster aks-h100
/only test_ci

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 02 hours 02 minutes 49 seconds. 🔴

• Link to the test results.

• No reports index generated...

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: pd-flavors
PR_POSITIONAL_ARG_3: guidellm_heterogeneous_eval
PR_POSITIONAL_ARG_4: gpt-oss

Failure indicator: Empty. (See run.log)

@kpouget

kpouget commented May 22, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 pd-flavors guidellm_heterogeneous_eval gpt-oss
/cluster aks-h100
/only test_ci

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 03 hours 28 minutes 33 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: pd-flavors
PR_POSITIONAL_ARG_3: guidellm_heterogeneous_eval
PR_POSITIONAL_ARG_4: gpt-oss

Failure indicator:

/tmp/topsail_202605221779462981/002__plots/FAILURE | An error happened during the visualization post-processing ... (0_matbench_parse.log in /tmp/topsail_202605221779462981/002__plots). The test that was processed SUCCEEDED.


@kpouget

kpouget commented May 26, 2026

Copy link
Copy Markdown
Contributor Author

/test jump-ci llm-d azure_h100 pd-flavors guidellm_heterogeneous_eval gpt-oss
/cluster aks-h100
/only test_ci
/var tests.llmd.flavors: intelligentrouting-tp4

@psap-forge-bot

Copy link
Copy Markdown

🔴 Test of 'llm-d test test_ci' failed after 01 hours 08 minutes 29 seconds. 🔴

• Link to the test results.

• Link to the reports index.

Test configuration:

PR_POSITIONAL_ARGS: jump-ci
PR_POSITIONAL_ARG_0: jump-ci
PR_POSITIONAL_ARG_1: azure_h100
PR_POSITIONAL_ARG_2: pd-flavors
PR_POSITIONAL_ARG_3: guidellm_heterogeneous_eval
PR_POSITIONAL_ARG_4: gpt-oss
tests.llmd.flavors: intelligentrouting-tp4

Failure indicator:

/tmp/topsail_202605261779778851/002__plots/FAILURE | An error happened during the visualization post-processing ... (0_matbench_parse.log in /tmp/topsail_202605261779778851/002__plots). The test that was processed SUCCEEDED.


@openshift-ci

openshift-ci Bot commented May 26, 2026

Copy link
Copy Markdown

@kpouget: The following test failed, say /retest to rerun all failed tests or /retest-required to rerun all mandatory failed tests:

Test name Commit Details Required Rerun command
ci/prow/jump-ci 19bf445 link true /test jump-ci

Full PR test history. Your PR dashboard.

Details

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository. I understand the commands that are listed here.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🧹 Nitpick comments (1)
projects/llm-d/testing/test_llmd.py (1)

1163-1182: 💤 Low value

Remove unused add_hotfix_volumes helper function.

add_hotfix_volumes (lines 1163-1171) is defined but never called. Only add_hotfix_to_containers (lines 1174-1182) is used at lines 1186 and 1191. These two functions are identical.

🧹 Proposed fix to remove dead code
-        # Helper function to add hotfix volume mounts to containers
-        def add_hotfix_volumes(containers, template_volumes):
-            for container in containers:
-                # Add volume mounts
-                if 'volumeMounts' not in container:
-                    container['volumeMounts'] = []
-                container['volumeMounts'].extend(hotfix_mounts)
-
-            # Add volume to template
-            template_volumes.append(hotfix_volume)
-
         # Apply hotfix volumes and complete configuration to containers
         def add_hotfix_to_containers(containers, template_volumes):
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 1163 - 1182, Remove the
redundant helper add_hotfix_volumes: delete the entire add_hotfix_volumes
function definition (the block that mirrors add_hotfix_to_containers) and keep
only add_hotfix_to_containers; verify there are no remaining references to
add_hotfix_volumes elsewhere and run tests to ensure behavior is unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 1116-1130: The except block currently catches errors raised in the
try (including the intentional RuntimeError when the ConfigMap is missing) and
re-raises a new RuntimeError without preserving the original traceback; update
the handler to preserve the exception chain by using "raise
RuntimeError(f\"Failed to verify ConfigMap '{cm_name}' in namespace
'{namespace}': {e}\") from e" and (optionally) avoid double-wrapping by
re-raising when isinstance(e, RuntimeError) so the original error from the
runtime check on hotfix_files/oc output is not obscured.

In `@projects/llm-d/toolbox/llmd_run_guidellm_benchmark/tasks/main.yml`:
- Around line 138-143: The task "Copy benchmarks.json from PVC to artifacts
directory (with retry)" currently sets retries/delay but uses "when:
file_check_result.rc == 0" so it never retries; change the task to register the
copy command result (e.g., add "register: copy_result"), remove the "when:
file_check_result.rc == 0" guard, and add "until: copy_result.rc == 0" so
Ansible will actually retry the oc exec shell command (the shell line that runs
oc exec ... > "{{ artifact_extra_logs_dir }}/artifacts/results/benchmarks.json")
using the existing retries/delay settings.

In
`@projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py`:
- Line 351: The code uses an unnecessary f-string in the call to html.H4
(html.H4(f"🔧 Simple Flavor Comparison")) which triggers Ruff F541; remove the f
prefix so the literal string is passed directly (change html.H4(f"...") to
html.H4("...")) to eliminate the lint warning and avoid creating an unneeded
formatted string.

---

Nitpick comments:
In `@projects/llm-d/testing/test_llmd.py`:
- Around line 1163-1182: Remove the redundant helper add_hotfix_volumes: delete
the entire add_hotfix_volumes function definition (the block that mirrors
add_hotfix_to_containers) and keep only add_hotfix_to_containers; verify there
are no remaining references to add_hotfix_volumes elsewhere and run tests to
ensure behavior is unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: f0cb6103-b0a3-4911-be6b-657b7ba6d1df

📥 Commits

Reviewing files that changed from the base of the PR and between 8ae2b7d and 5251712.

📒 Files selected for processing (4)
  • projects/llm-d/testing/config.yaml
  • projects/llm-d/testing/test_llmd.py
  • projects/llm-d/toolbox/llmd_run_guidellm_benchmark/tasks/main.yml
  • projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py

Comment on lines +1116 to +1130
try:
result = run.run(f"oc get configmap {cm_name} -n {namespace} --ignore-not-found -o name",
capture_stdout=True, check=True)
if not result.stdout.strip():
# Extract just the filenames from hotfix_files for the error message
filenames = [file_path.split('/')[-1] for file_path in hotfix_files]
file_args = ' \\\n '.join([f"--from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/{filename}" for filename in filenames])

raise RuntimeError(f"Required ConfigMap '{cm_name}' not found in namespace '{namespace}'. "
f"Please create it first:\n\n"
f"oc create cm vllm-ucx-multiproc-hotfix -n {namespace} \\\n"
f" {file_args}")
logging.info(f"Verified ConfigMap '{cm_name}' exists in namespace '{namespace}'")
except Exception as e:
raise RuntimeError(f"Failed to verify ConfigMap '{cm_name}' in namespace '{namespace}': {e}")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

Exception handling wraps its own RuntimeError and loses exception chain.

The try block raises RuntimeError on line 1124 when the ConfigMap is not found. The except Exception on line 1129 catches this same error, wrapping it in another RuntimeError and losing the original traceback. Additionally, raise ... from e should be used to preserve the exception chain.

🐛 Proposed fix
         try:
             result = run.run(f"oc get configmap {cm_name} -n {namespace} --ignore-not-found -o name",
                              capture_stdout=True, check=True)
             if not result.stdout.strip():
                 # Extract just the filenames from hotfix_files for the error message
                 filenames = [file_path.split('/')[-1] for file_path in hotfix_files]
                 file_args = ' \\\n       '.join([f"--from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/{filename}" for filename in filenames])

                 raise RuntimeError(f"Required ConfigMap '{cm_name}' not found in namespace '{namespace}'. "
                                    f"Please create it first:\n\n"
                                    f"oc create cm vllm-ucx-multiproc-hotfix -n {namespace} \\\n"
                                    f"       {file_args}")
             logging.info(f"Verified ConfigMap '{cm_name}' exists in namespace '{namespace}'")
-        except Exception as e:
-            raise RuntimeError(f"Failed to verify ConfigMap '{cm_name}' in namespace '{namespace}': {e}")
+        except RuntimeError:
+            raise  # Re-raise our own RuntimeError as-is
+        except Exception as e:
+            raise RuntimeError(f"Failed to verify ConfigMap '{cm_name}' in namespace '{namespace}': {e}") from e
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
try:
result = run.run(f"oc get configmap {cm_name} -n {namespace} --ignore-not-found -o name",
capture_stdout=True, check=True)
if not result.stdout.strip():
# Extract just the filenames from hotfix_files for the error message
filenames = [file_path.split('/')[-1] for file_path in hotfix_files]
file_args = ' \\\n '.join([f"--from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/{filename}" for filename in filenames])
raise RuntimeError(f"Required ConfigMap '{cm_name}' not found in namespace '{namespace}'. "
f"Please create it first:\n\n"
f"oc create cm vllm-ucx-multiproc-hotfix -n {namespace} \\\n"
f" {file_args}")
logging.info(f"Verified ConfigMap '{cm_name}' exists in namespace '{namespace}'")
except Exception as e:
raise RuntimeError(f"Failed to verify ConfigMap '{cm_name}' in namespace '{namespace}': {e}")
try:
result = run.run(f"oc get configmap {cm_name} -n {namespace} --ignore-not-found -o name",
capture_stdout=True, check=True)
if not result.stdout.strip():
# Extract just the filenames from hotfix_files for the error message
filenames = [file_path.split('/')[-1] for file_path in hotfix_files]
file_args = ' \\\n '.join([f"--from-file=guides/pd-disaggregation/ms-pd/charts/vllm-ucx-multiproc-hotfix/{filename}" for filename in filenames])
raise RuntimeError(f"Required ConfigMap '{cm_name}' not found in namespace '{namespace}'. "
f"Please create it first:\n\n"
f"oc create cm vllm-ucx-multiproc-hotfix -n {namespace} \\\n"
f" {file_args}")
logging.info(f"Verified ConfigMap '{cm_name}' exists in namespace '{namespace}'")
except RuntimeError:
raise # Re-raise our own RuntimeError as-is
except Exception as e:
raise RuntimeError(f"Failed to verify ConfigMap '{cm_name}' in namespace '{namespace}': {e}") from e
🧰 Tools
🪛 Ruff (0.15.14)

[warning] 1129-1129: Do not catch blind exception: Exception

(BLE001)


[warning] 1130-1130: Within an except clause, raise exceptions with raise ... from err or raise ... from None to distinguish them from errors in exception handling

(B904)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/testing/test_llmd.py` around lines 1116 - 1130, The except
block currently catches errors raised in the try (including the intentional
RuntimeError when the ConfigMap is missing) and re-raises a new RuntimeError
without preserving the original traceback; update the handler to preserve the
exception chain by using "raise RuntimeError(f\"Failed to verify ConfigMap
'{cm_name}' in namespace '{namespace}': {e}\") from e" and (optionally) avoid
double-wrapping by re-raising when isinstance(e, RuntimeError) so the original
error from the runtime check on hotfix_files/oc output is not obscured.

Comment on lines +138 to +143
- name: Copy benchmarks.json from PVC to artifacts directory (with retry)
shell: |
oc exec {{ llmd_run_guidellm_benchmark_name }}-copy -n "{{ target_namespace }}" -- cat /results/benchmarks.json > "{{ artifact_extra_logs_dir }}/artifacts/results/benchmarks.json"
retries: 10
delay: 1
when: file_check_result.rc == 0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
rg -n -C2 'Copy benchmarks\.json from PVC|retries:|delay:|until:' projects/llm-d/toolbox/llmd_run_guidellm_benchmark/tasks/main.yml

Repository: openshift-psap/topsail

Length of output: 842


🏁 Script executed:

#!/bin/bash
sed -n '110,170p' projects/llm-d/toolbox/llmd_run_guidellm_benchmark/tasks/main.yml

Repository: openshift-psap/topsail

Length of output: 1983


🌐 Web query:

Ansible task retries delay without until semantics ignored

💡 Result:

In Ansible, the retries and delay parameters are functional components of the until loop mechanism; they have no effect if used in a task without the until keyword [1]. When until is omitted, Ansible ignores these settings, and the task will execute exactly once [2]. Key points regarding this behavior: - Functional Dependency: The retries and delay keywords are specifically designed to control the polling behavior of until [1]. Without until, there is no retry condition for Ansible to evaluate, so the task proceeds as a standard, single-execution task [2]. - Historical Context: Users have frequently requested that retries without until default to "retry until success," but Ansible has maintained the requirement that until must be explicitly defined to enable retry logic [2]. - Best Practices: To implement retry logic correctly, you must register the output of the task and use the until keyword to evaluate that registered variable [3][4]. For example: - name: Example task with retries ansible.builtin.command: /usr/bin/false register: result until: result is succeeded retries: 5 delay: 10 If you need to wait for a condition or a resource to become available without complex retry logic, the ansible.builtin.wait_for or ansible.builtin.wait_for_connection modules are often more appropriate alternatives [5][6].

Citations:


Make the “with retry” task retry by adding an until condition.

In projects/llm-d/toolbox/llmd_run_guidellm_benchmark/tasks/main.yml (lines 138-143), the task sets retries/delay but omits until, so Ansible runs it only once and ignores the retry settings.

Suggested patch
       - name: Copy benchmarks.json from PVC to artifacts directory (with retry)
         shell: |
           oc exec {{ llmd_run_guidellm_benchmark_name }}-copy -n "{{ target_namespace }}" -- cat /results/benchmarks.json > "{{ artifact_extra_logs_dir }}/artifacts/results/benchmarks.json"
+        register: copy_result
         retries: 10
         delay: 1
+        until: copy_result.rc == 0
         when: file_check_result.rc == 0
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@projects/llm-d/toolbox/llmd_run_guidellm_benchmark/tasks/main.yml` around
lines 138 - 143, The task "Copy benchmarks.json from PVC to artifacts directory
(with retry)" currently sets retries/delay but uses "when: file_check_result.rc
== 0" so it never retries; change the task to register the copy command result
(e.g., add "register: copy_result"), remove the "when: file_check_result.rc ==
0" guard, and add "until: copy_result.rc == 0" so Ansible will actually retry
the oc exec shell command (the shell line that runs oc exec ... > "{{
artifact_extra_logs_dir }}/artifacts/results/benchmarks.json") using the
existing retries/delay settings.

baseline_settings = {"model": "llama3.3-70b", "load_shape": load_shape, "flavor": flavor}
# If only one platform, show all simple flavors together
if "platform" not in variables:
header.append(html.H4(f"🔧 Simple Flavor Comparison"))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor | ⚡ Quick win

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
rg -n 'html\.H4\(f"🔧 Simple Flavor Comparison"\)' projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py

Repository: openshift-psap/topsail

Length of output: 141


🏁 Script executed:

#!/bin/bash
set -euo pipefail

FILE="projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py"

# Show the exact surrounding context for the reported line
sed -n '330,370p' "$FILE" | nl -ba -w3 -s':'

# Confirm the exact match count for the specific string
rg -n 'html\.H4\(f"🔧 Simple Flavor Comparison"\)' "$FILE"

# Broader scan: any other html.H4 f-strings without placeholders
rg -n 'html\.H4\(f"' "$FILE" || true

Repository: openshift-psap/topsail

Length of output: 109


🏁 Script executed:

#!/bin/bash
set -euo pipefail

FILE="projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py"

# Show surrounding context with line numbers
awk -v start=330 -v end=370 'NR>=start && NR<=end {printf "%d:%s\n", NR, $0}' "$FILE"

# Confirm the exact match count for the specific string
rg -n 'html\.H4\(f"🔧 Simple Flavor Comparison"\)' "$FILE"

# Broader scan: any other html.H4 f-strings
rg -n 'html\.H4\(f"' "$FILE" || true

Repository: openshift-psap/topsail

Length of output: 2525


Remove unnecessary f-string prefix (Ruff F541)

projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py line 351 uses html.H4(f"🔧 Simple Flavor Comparison") with no {} placeholders—remove the f.

Suggested patch
-                header.append(html.H4(f"🔧 Simple Flavor Comparison"))
+                header.append(html.H4("🔧 Simple Flavor Comparison"))
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
header.append(html.H4(f"🔧 Simple Flavor Comparison"))
header.append(html.H4("🔧 Simple Flavor Comparison"))
🧰 Tools
🪛 Ruff (0.15.14)

[error] 351-351: f-string without any placeholders

Remove extraneous f prefix

(F541)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In
`@projects/llm-d/visualizations/llmd_inference/plotting/throughput_comparisons.py`
at line 351, The code uses an unnecessary f-string in the call to html.H4
(html.H4(f"🔧 Simple Flavor Comparison")) which triggers Ruff F541; remove the f
prefix so the literal string is passed directly (change html.H4(f"...") to
html.H4("...")) to eliminate the lint warning and avoid creating an unneeded
formatted string.

@kpouget kpouget closed this May 27, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant