Skip to content

[NVBUG-6448152][test] TEST ONLY pre-cancellation native-source discriminator#16854

Closed
chienchunhung wants to merge 5 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-midpoint-544199a4
Closed

[NVBUG-6448152][test] TEST ONLY pre-cancellation native-source discriminator#16854
chienchunhung wants to merge 5 commits into
NVIDIA:mainfrom
chienchunhung:codex/nvbug-6448152-native-midpoint-544199a4

Conversation

@chienchunhung

Copy link
Copy Markdown
Collaborator

TEST ONLY — do not merge.

This is one fixed-current-runtime source discriminator for the separate C++ CTX PP forward/device-loop throughput regression. The asynchronous-consensus product change is already merged and is not under review here.

One-factor construction

  • Native source: 544199a47b91b65d295322c6bb1f74dadbcf6a79
  • Native source tree: 5726e87af6bc2267e116eb26c1bd71ef9c5fded6
  • Native source parent: f665e59a7d650672433743bb6233712e04c5a90f
  • This source is three first-parent commits earlier than the valid slow 58d8964d checkpoint from the preceding discriminator. The two intervening commits before 58d8964d are test-waiver and CODEOWNERS changes; 58d8964d is the only intervening product-code change.
  • Current-runtime overlay is limited to the NIXL installer/development pins and all five image tags using build 202607151440-16194; stable patch ID: 7e73140158499de670b40490317c6f77601d06fc.
  • The asynchronous-consensus factor is explicitly a semantically matched pre-cancellation port, not a byte-identical application. This native source predates the cancellation APIs expected by the later factor. The port preserves automatic NIXL/UCX TP1/CP1/PP activation, protocol version 1 with cancellation fixed off (mode 2), immutable terminal-vote publication, and coordinator poll/yield completion. It does not add or execute timeout publication or cancellation handling. Factor stable patch ID: 52adfb97177ae01283e352d36d52b428f2f1bf00.
  • The coordinator implementation, MPI tags/utilities, and coordinator tests are byte-identical to the later validated factor. One CMake list-order normalization was restored after applying the factor and has no net tree effect beyond adding the factor source.
  • The current image, native five-way stage mapping, workload YAML, selector, and timeout are unchanged from the preceding valid slow discriminator.
  • Current main is present only as a no-tree second parent for CI mergeability.

Interpretation

  • At least 1402 output tokens/s places the regression in the following three first-parent commits through 58d8964d, where the only product-code change is the NIXL in-flight-cancellation change.
  • Near 800 output tokens/s excludes those three commits and moves the source bound earlier.
  • Any partial recovery will be quantified.
  • A build/import failure, timeout, failed request, cancellation event, or protocol-mode mismatch is compatibility/censored evidence and is not a throughput result.

Requested evidence

Run only GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, requiring the exact selector, three nodes, twelve tasks, four GPUs per node, --no-container-mount-home, 512/512 successful requests, coordinator activation with protocol mode 2 and cancellation disabled on all four CTX PP ranks, no timeout/cancellation events, clean shutdown, and a valid official throughput metric.

The draft will be closed after terminal evidence is captured; the branch will be preserved.

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
…sensus factor

Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
@chienchunhung

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1"

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61657 [ run ] triggered by Bot. Commit: 015d657 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61657 [ run ] completed with state FAILURE. Commit: 015d657
/LLM/main/L0_MergeRequest_PR pipeline #49864 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Copy link
Copy Markdown
Collaborator Author

Terminal evidence for this TEST ONLY discriminator:

  • At the exact diagnostic head, the workload completed 512/512 requests with 0 failed requests.
  • Benchmark duration was 5257.38 seconds. The official metrics were 797.79 output tok/s, 0.10 req/s, and 13562.48 total tok/s.
  • Protocol mode 2 with in-flight cancellation disabled activated on all four CTX ranks. CTX, GEN, and disaggregated applications shut down cleanly, with no fatal or protocol errors.
  • The terminal Jenkins failure is solely the expected performance-threshold gate after a valid result was recorded.

This is +0.40% versus the later-tree 794.61 result and 51.51% of the historical 1548.84 result. There is no throughput recovery at the native 544199a4 checkpoint, so the following three first-parent commits through 58d8964 are excluded and the slow bound moves earlier.

The diagnostic branch is preserved for reproducibility; this TEST ONLY draft is now complete and will be closed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants