Skip to content

[None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host#16791

Open
hyukn wants to merge 1 commit into
NVIDIA:mainfrom
hyukn:yukunh/prepare-inputs-gettokens-oL
Open

[None][perf] prepare_inputs: avoid O(seq_len) get_tokens(0) marshalling on the host#16791
hyukn wants to merge 1 commit into
NVIDIA:mainfrom
hyukn:yukunh/prepare-inputs-gettokens-oL

Conversation

@hyukn

@hyukn hyukn commented Jul 23, 2026

Copy link
Copy Markdown
Collaborator

Summary

request.get_tokens(0) marshals the whole C++ VecTokens (the entire sequence) into a fresh Python list of boxed ints — O(seq_len) host work that grows linearly with ISL. Three call sites in _prepare_tp_inputs / cuda_graph_runner pay this only to read a length or a small slice, and re-pay it every iteration. In chunked prefill the context loop re-marshals the full prompt for every chunk → O(L²/chunk) per request over the prefill. Normal decode is already immune (it uses get_last_tokens(0), O(1)); the waste is in the prefill/context path and the MTP first-draft paths.

Changes

  • New nanobind binding LlmRequest.get_tokens_range(beam, begin, end) — returns only the [begin, end) window, so nanobind copies just (end-begin) tokens (O(chunk)) instead of the whole VecTokens. Bounds are clamped.
  • context loop (_prepare_tp_inputs): use get_tokens_range(0, begin, end) for the current chunk and get_num_tokens(0) (O(1)) for the prompt length, instead of get_tokens(0) + slice + len(...).
  • first_draft loop (MTP): same — get_num_tokens(0) for the length, get_tokens_range for the last original_max_draft_len+1 tokens.
  • cuda_graph_runner first-draft branch: len(get_tokens(0))get_num_tokens(0).

Output is bit-identicalget_tokens_range returns exactly the same subrange the old get_tokens(0)[begin:end] produced.

Perf — _prepare_inputs host time (DSv4-Pro disagg, c3120, matched A/B, nsys, nvtx_range("_prepare_inputs"))

CTX / prefill worker (non-overlap, host-bound; primary target; iters 400–500, N=80):

_prepare_inputs p50 p90 success
baseline 27.53 ms 43.29 ms
this PR 15.10 ms (−45%) 20.43 ms (−53%) 100% (3000/3000)

GEN / decode worker (overlap on; only the MTP first-draft path hits get_tokens(0), so only the large-L tail moves; iters 6000–6050, N=50):

_prepare_inputs mean p50 p90
baseline 17.84 ms 13.31 ms 32.78 ms
this PR 13.26 ms (−26%) 11.05 ms (−17%) 14.64 ms (−55%)

Each site drops from O(L) to O(1)/O(chunk); a synthetic microbench (bs128) shows each site flat across ISL after the fix (baseline get_tokens ~18 ms/iter at ISL 50K → ~0.05 ms). On CTX the whole _prepare_inputs shifts down (it sits on the exposed critical path). On GEN the median iter barely moves (normal decode is already O(1) via get_last_tokens) while the tail collapses (−55% p90) — the large-accumulated-length first-draft iters. GPU-timeline analysis on the same GEN capture shows the reduction is exposed, not hidden: GEN GPU bubble 25.7% → 20.7% and per-iteration cycle −6.3% with GPU compute unchanged (host win on the local critical path); pooled E2E SSE/throughput stay flat because GEN decode is not this workload's binding constraint. The binding was rebuilt and the new get_tokens_range verified present on the C++ LlmRequest; the full disagg run served at 100% success (functional parity).

🤖 Generated with Claude Code

…ng on the host

`request.get_tokens(0)` marshals the whole C++ VecTokens (entire sequence) into a
fresh Python list of boxed ints -- O(seq_len) host work that grows linearly with
ISL. Three call sites in `_prepare_tp_inputs` / `cuda_graph_runner` pay this only
to read a length or a small slice, and re-pay it every iteration. In chunked
prefill the context loop re-marshals the full prompt for every chunk -> O(L^2/chunk)
per request over the prefill. Normal decode is already immune (get_last_tokens(0),
O(1)); the waste is in the prefill/context and MTP first-draft paths.

Changes:
- New nanobind binding `LlmRequest.get_tokens_range(beam, begin, end)` that copies
  only [begin, end) (O(chunk)) instead of the whole VecTokens.
- context loop and first_draft loop: use get_tokens_range for the chunk and
  get_num_tokens(0) (O(1)) for lengths, instead of get_tokens(0) + slice.
- cuda_graph_runner first-draft branch: len(get_tokens(0)) -> get_num_tokens(0).

Output is bit-identical (get_tokens_range returns the same subrange). Measured on a
DSv4-Pro disagg CTX worker (c3120, non-overlap, CTX nsys): _prepare_inputs p50
27.53 ms -> 15.10 ms (-45%), p90 43.29 -> 20.43 ms; 100% request success. Complements
NVIDIA#16734 (which removes the overlap device-scalar sync, 268 -> 19.96 ms); together they
target the two independent costs in _prepare_inputs.

Signed-off-by: Yukun He <23156053+hyukn@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@hyukn
hyukn force-pushed the yukunh/prepare-inputs-gettokens-oL branch from 460954a to 1916433 Compare July 23, 2026 14:16
@hyukn
hyukn marked this pull request as ready for review July 23, 2026 14:16
@hyukn
hyukn requested review from a team as code owners July 23, 2026 14:16
@hyukn

hyukn commented Jul 23, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61304 [ run ] triggered by Bot. Commit: 1916433 Link to invocation

@coderabbitai

coderabbitai Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Adds a bounded get_tokens_range request binding and updates executor code to use direct token counts or limited token windows for chunked prefill, first-draft requests, and CUDA graph sequence-length calculations.

Changes

Token Range Retrieval

Layer / File(s) Summary
Token range binding
cpp/tensorrt_llm/nanobind/batch_manager/bindings.cpp
Adds get_tokens_range(beam, begin, end), which clamps bounds and returns a half-open token slice.
Model engine token windows
tensorrt_llm/_torch/pyexecutor/model_engine.py
Uses bounded token retrieval for context chunks and first-draft suffixes, and obtains prompt length from the request token count.
CUDA graph sequence sizing
tensorrt_llm/_torch/pyexecutor/cuda_graph_runner.py
Uses get_num_tokens(0) for first-draft sequence-length calculation.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Suggested reviewers: cascade812

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Description check ⚠️ Warning The description explains the change well, but it omits the required Description, Test Coverage, and PR Checklist sections from the template. Add the missing template sections, especially concrete test coverage and a completed PR checklist, and align the headings with the repository template.
✅ Passed checks (4 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title is specific and accurately describes the main performance change in the PR.
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Comment @coderabbitai help to get the list of available commands.

@hyukn

hyukn commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61437 [ run ] triggered by Bot. Commit: 1916433 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61304 [ run ] completed with state ABORTED. Commit: 1916433

Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61437 [ run ] completed with state SUCCESS. Commit: 1916433
/LLM/main/L0_MergeRequest_PR pipeline #49660 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@hyukn

hyukn commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61560 [ run ] triggered by Bot. Commit: 1916433 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61560 [ run ] completed with state SUCCESS. Commit: 1916433
/LLM/main/L0_MergeRequest_PR pipeline #49771 completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

  • Please check the failed tests and fix your PR
  • If you cannot view the failures, ask the CI triggerer to share details
  • Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

@hyukn

hyukn commented Jul 24, 2026

Copy link
Copy Markdown
Collaborator Author

/bot run --disable-fail-fast

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61604 [ run ] triggered by Bot. Commit: 1916433 Link to invocation

@tensorrt-cicd

Copy link
Copy Markdown
Collaborator

PR_Github #61604 [ run ] completed with state SUCCESS. Commit: 1916433
/LLM/main/L0_MergeRequest_PR pipeline #49813 completed with status: 'SUCCESS'

CI Report

Link to invocation

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants