fix: bound root-chain calls to prevent node-wide deadlock#473
Open
pablocampogo wants to merge 2 commits into
Open
fix: bound root-chain calls to prevent node-wide deadlock#473pablocampogo wants to merge 2 commits into
pablocampogo wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fix: bound root-chain calls to prevent node-wide deadlock
Full nodes intermittently wedged until restart: the BLOCK and
BLOCK_REQUEST inboxes filled to capacity and dropped every message
("CRITICAL: Inbox BLOCK queue full"). Both inbox consumers require the
controller mutex, which was being held forever by block processing
stuck in an unbounded root-chain call.
Only full nodes were affected because validators stay on the in-memory
committee cache: their cached root-chain info tracks the live root
height, so LoadCommittee resolves from sub.Info without a network call.
Full nodes repeatedly fall behind and re-sync, crossing checkpoint
heights and certificates with stale root heights, which forces
committee/checkpoint lookups onto the remote HTTP path -- and that path
had no timeout, so one request stuck on a half-open connection (NAT
drop, root-chain RPC restart) held the controller mutex forever.
Two fixes:
Use an http.Client with a 10s timeout for RCSubscription's on-demand
root-chain calls (LoadCommittee, GetCheckpoint, lottery, dex batch,
etc.). A hung request now surfaces as an error and the block is
simply re-requested instead of wedging block processing.
Add a read deadline + ping-refreshed liveness check to the root-chain
subscription websocket. A silently dead connection is now detected
and redialed instead of hanging Listen() forever, which left the
cached root-chain info stale and forced lookups onto the broken
remote path.