Skip to content

Tuner autoscaler cannot allocate GPU #96

Description

@Fidasek009

When running the tuner, the autoscaler triggers a request for the worker pod which requires a GPU. If the cluster doesn't provide a GPU in an hour, the worker pod, that is still waiting, suddenly scales down due to the autoscaler config. However the tuner api still reports the pending trials.

The cluster is the root problem since the tuner becomes unusable if it cannot even allocate a GPU within that hour long window.

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingwontfixThis will not be worked on

    Type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions