-
Notifications
You must be signed in to change notification settings - Fork 2.6k
All issues
Issue creation is restricted in this repository
- #15044 · laikhtewari opened
on Jun 6, 2026 1 - #3148 · juney-nvidia opened
on Mar 29, 2025 5 - #3124 · juney-nvidia opened
on Mar 27, 2025 11
Issues
is:issue state:open
is:issue state:open
Search results
[Performance]: MAX_UTILIZATION causes a 40.6% output-throughput drop and 9.1x TPOT p99 under KV-cache oversubscription (Llama-3.3-70B-FP8, H200, v1.2.1)
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentKV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferencePytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16827 In NVIDIA/TensorRT-LLM;[Bug]: SA (Suffix Automaton) speculative decoding crashes the executor event loop when a prompt is near
max_seq_len—default_max_tokensis deduced against the draft-inflatedmax_seq_lenbugSomething isn't workingSomething isn't workingInference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Speculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16826 In NVIDIA/TensorRT-LLM;[Bug] Vanilla attention uses incorrect KV indices for layer-specific cache layouts
KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceStatus: Open.#16801 In NVIDIA/TensorRT-LLM;TensorRT-LLM PyTorch backend Qwen3.5-VL parity issue: HF/vLLM greedy output matches, TRT-LLM repeats to max tokens
bugSomething isn't workingSomething isn't workingMultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsPytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesStatus: Open.#16792 In NVIDIA/TensorRT-LLM;[Bug] DSpark speculative decoding: accept length collapses to ~1 at generation batch size > 1 in disaggregated serving
Disaggregated serving<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.<NV>Deploying with separated, distributed components (params, kv-cache, compute). Arch & perf.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16767 In NVIDIA/TensorRT-LLM;gRPC GenerateRequest.max_tokens is required — should be optional with an engine-side default (parity with LLM API)
LLM API<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.<NV>High-level LLM Python API & tools (e.g., trtllm-llmapi-launch) for TRTLLM inference/workflows.Status: Open.#16549 In NVIDIA/TensorRT-LLM;[New Model]: thinkingmachines/Inkling
new modelRequest to add a new modelRequest to add a new modelStatus: Open.#16507 In NVIDIA/TensorRT-LLM;[Bug]: find_input_mm_embeds is order-dependent for mixed full-prefill and partial VLM requests
MultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsStatus: Open.#16460 In NVIDIA/TensorRT-LLM;[Bug]: Preserve per-item processed metadata when computing multimodal token lengths
MultimodalLabel for issues & PRs regarding Multimodal related objectsLabel for issues & PRs regarding Multimodal related objectsStatus: Open.#16459 In NVIDIA/TensorRT-LLM;KV cache connector hangs with Eagle speculative decoding when rewind crosses block boundary
KV-Cache Managementkv-cache management for efficient LLM inferencekv-cache management for efficient LLM inferenceSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16448 In NVIDIA/TensorRT-LLM;[Bug]: Eagle3 speculative decoding silently drops tool_calls for gpt-oss on trtllm-serve
bugSomething isn't workingSomething isn't workingInference runtime<NV>General operational aspects of TRTLLM execution not in other categories.<NV>General operational aspects of TRTLLM execution not in other categories.Pytorch<NV>Pytorch backend related issues<NV>Pytorch backend related issuesSpeculative Decoding<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafter<NV>MTP/Eagle/Medusa/Lookahead/Prompt-Lookup-Decoding/Draft-Target-Model/ReDrafterStatus: Open.#16377 In NVIDIA/TensorRT-LLM;[Performance]: Shape-gated backend selection for FP8 block-scaling GEMM
General perf<NV>Broad performance issues not specific to a particular component<NV>Broad performance issues not specific to a particular componentStatus: Open.#16363 In NVIDIA/TensorRT-LLM;