[FLINK-39970] Retry incomplete JobManager deployment deletion#1145
Open
stellagai wants to merge 1 commit into
Open
[FLINK-39970] Retry incomplete JobManager deployment deletion#1145stellagai wants to merge 1 commit into
stellagai wants to merge 1 commit into
Conversation
stellagai
marked this pull request as ready for review
June 23, 2026 01:00
Dennis-Mircea
left a comment
Contributor
There was a problem hiding this comment.
Thanks for the PR. The fix itself is correct, but this overlaps heavily with #1138 (FLINK-39953), which targets the same root cause in the same deleteBlocking wait catch and the same shutdownJobManagersBlocking caller.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What is the purpose of the change
Kubernetes deployment deletion waits can fail or time out before the old JobManager deployment is fully removed. The operator currently logs these failures and continues reconciliation, which can submit a replacement cluster while the old deployment is still terminating and result in
AlreadyExistserrors.Brief change log
Verifying this change
This change added and updated tests and was verified with:
JAVA_HOME=/opt/homebrew/opt/openjdk@17/libexec/openjdk.jdk/Contents/Home \ mvn -pl flink-kubernetes-operator -am \ -DskipITs \ -Dtest=AbstractFlinkServiceTest \ -Dsurefire.failIfNoSpecifiedTests=false testTests run: 40, failures: 0, errors: 0, skipped: 0.
Does this pull request potentially affect one of the following parts:
CustomResourceDescriptors: noDocumentation