Lacking BEP resources
Symptoms
If you run a model, and the following happens:
The results are not displayed for a long time.
You receive a timeout exception.
The symptoms indicate that the problem might be not the model but the resources for running it.
Cause
The issue can be due to various reasons:
Unmounted and lost BEP Agents. For details, refer to Issues related to NFS shared resources.
If you have limited resource capacity, your cluster might not even be considered to run ML models. Usually, one model Worker instance requires 4+ GB of memory and 1 CPU.
Solution
First of all, take a look at your BEP cluster and check the following:
Overall resources you have
Resources allocated to different tasks and processes
In case you don't see some part of your resources, it's a symptom of unmounted and lost BEP Agents.
If the resources are not lost but almost all allocated, and there are no free resources to start a new ML task, consider adding a new BEP Agent.
View also: