Skip to content

Investigate llama-server repeated-request failure and semantic corruption on RK3588 dense-balanced hybrid path #4

Description

@KHAEntertainment

Summary

After merging the dense-balanced RKNPU placement fixes from PR #3, the original llama-cli crash path appears resolved on ROCK 5 / RK3588, but a new issue shows up on the long-lived server path.

With the merged dense-balanced hybrid configuration:

  • llama-cli now succeeds repeatedly on the validated post-merge path
  • llama-server starts, loads the same model/config successfully, and serves at least one request
  • a subsequent request causes the connection to drop / server to die
  • the first successful completion also looked semantically corrupted / garbled rather than like sane natural-language output

This issue tracks the new post-merge problem so it can be investigated separately from the earlier packed-weight placement crash.

Environment

Hardware / OS:

  • Radxa ROCK 5C / RK3588
  • Debian 12
  • native DRM render-node path is the correct runtime path on this board
  • /dev/rknpu symlink should remain absent

Repo state:

Build:

  • build-rknpu
  • GGML_RKNPU2=ON
  • llama-cli and llama-server both built successfully

Model:

  • /home/radxa/hf_models/Qwen_Qwen3.5-4B-Q8_0.gguf

What is already validated

The earlier dense-balanced failure mode appears fixed enough for CLI validation:

  • 3/3 llama-cli runs completed successfully with EXIT:0
  • merged placement fix is active in logs:
    • *.ffn_up.weight -> CPU-visible
    • *.ffn_gate.weight -> CPU-visible
    • *.ffn_down.weight -> RKNPU
  • the old ffn_up_gate-35 packed-weight fallback copy crash did not reappear

Relevant routed tensor summary from successful CLI logs:

  • CPU tensors: 395 (~4154 MiB)
  • RKNPU tensors: 32 (~765 MiB)
  • top-level RKNPU routed buffer alloc: 720 MiB
  • top-level CMA alloc still fails, system heap fallback succeeds

Server repro

Server launch:

env HYBRID_MANIFEST=./examples/hybrid-manifests/dense-balanced.json \
    HYBRID_STRICT=1 \
    RKNN_SPLIT_FACTOR=1 \
    RKNPU_B_CACHE_SIZE=1024 \
    RKNPU_CTX_CACHE_SIZE=64 \
    RKNPU_B_MEM_MODE=copy \
    ./build-rknpu/bin/llama-server \
    -m /home/radxa/hf_models/Qwen_Qwen3.5-4B-Q8_0.gguf \
    --host 127.0.0.1 --port 8081 \
    --n-gpu-layers 99 \
    -c 256 -t 1 -b 32 -ub 32 \
    --no-warmup

Observed startup state:

  • /health returns OK
  • /props returns OK
  • server log shows model fully loaded and HTTP server listening

Example request path:

POST /completion
{
  "prompt": "Say hello in five words.",
  "n_predict": 8,
  "temperature": 0.2,
  "stream": false
}

Observed behavior:

  • first request succeeds
  • returned content is garbled / semantically suspicious, for example:
    • bung N -尽 ’
  • a subsequent request fails with connection drop / reset
  • from the client side this showed up as:
    • Remote end closed connection without response
    • then Connection reset by peer

Key log evidence

Server log confirms merged placement fix is active before serving:

  • ffn_gate.weight resolved to CPU with reason: tensor must stay CPU-visible for ggml_fused_up_gate
  • ffn_up.weight resolved to CPU with same reason
  • ffn_down.weight still routed to RKNPU

Server log also shows normal startup:

  • model loaded
  • HTTP server listening on 127.0.0.1:8081
  • first task launched on slot 0

Hypotheses

Most likely categories now:

  1. Long-lived / repeated-request stability bug specific to server mode on the dense-balanced hybrid path
  2. Semantic corruption in the current RKNPU_B_MEM_MODE=copy path even when the process does not immediately crash
  3. A server-only cache / slot / context lifecycle issue that does not show up in one-shot llama-cli runs

Next debugging targets

  • Reproduce with exactly 2 sequential /completion requests against one server process
  • Capture the server log immediately after the second request failure
  • Check whether the server process actually exits or whether the connection dies while the process remains alive
  • Inspect dmesg / kernel logs right after the failed second request
  • Compare semantic quality against CPU-only and/or legacy non-manifest runs on the same short prompts
  • Verify whether the corruption is specific to server mode, repeated requests, or the broader dense-balanced RKNPU_B_MEM_MODE=copy path itself

Why this is separate from the previous issue

The earlier crash was a packed-weight storage contract issue around ggml_fused_up_gate inputs living in packed RKNPU storage. That appears resolved for the validated CLI path. This new issue is about:

  • repeated-request server stability
  • semantic quality of returned text

Those should be investigated separately from the already-merged placement fix.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions