Using Autoscaler to scale from 0 machines
The cluster-autoscaler project supports Cluster API. With the scale-from-zero enhancement, worker nodes can be scaled down to 0 and provisioned on demand.
Setting up the workload cluster
Add the following annotations to your MachineDeployment to opt in to autoscaling. These are required by the autoscaler to know the min/max bounds when scaling from 0.
apiVersion: cluster.x-k8s.io/v1beta1
kind: MachineDeployment
metadata:
name: "${CLUSTER_NAME}-md-0"
annotations:
cluster.x-k8s.io/cluster-api-autoscaler-node-group-max-size: "5"
cluster.x-k8s.io/cluster-api-autoscaler-node-group-min-size: "0"
Setting up the cluster-autoscaler
-
Clone the autoscaler repository:
git clone https://github.com/kubernetes/autoscaler.git -
Build the autoscaler binary:
cd autoscaler/cluster-autoscaler go build . -
Start the autoscaler:
./cluster-autoscaler \ --cloud-provider=clusterapi \ --v=2 \ --namespace=default \ --max-nodes-total=30 \ --scale-down-delay-after-add=10s \ --scale-down-delay-after-delete=10s \ --scale-down-delay-after-failure=10s \ --scale-down-unneeded-time=5m \ --max-node-provision-time=30m \ --balance-similar-node-groups \ --expander=random \ --kubeconfig=<workload_cluster_kubeconfig> \ --cloud-config=<management_cluster_kubeconfig>
Note: The autoscaler can be run in several ways — see connecting to management and workload clusters for alternatives. A full list of command-line flags is available in the FAQ.
Walkthrough: scale up and scale down
-
Create a workload cluster with 0 worker machines.
-
Apply a sample workload:
apiVersion: apps/v1 kind: Deployment metadata: name: busybox-deployment namespace: default spec: replicas: 1 selector: matchLabels: app: busybox template: metadata: labels: app: busybox spec: containers: - name: busybox image: busybox imagePullPolicy: IfNotPresent command: ["sh", "-c", "echo Running; sleep 3600"] resources: requests: cpu: "0.2" memory: 3G -
Scale the deployment to trigger pending pods:
kubectl scale --replicas=2 deployment/busybox-deployment -
Observe that the second pod is pending (no nodes available yet):
kubectl get pods NAME READY STATUS RESTARTS AGE busybox-deployment-7c87788568-qhqdb 1/1 Running 0 48s busybox-deployment-7c87788568-t26bb 0/1 Pending 0 5s -
On the management cluster, watch the autoscaler provision a new machine:
kubectl get machines NAME CLUSTER PHASE VERSION ibm-powervs-control-plane-smvf7 ibm-powervs Running v1.34.7 ibm-powervs-md-0-6b4d67ccf4-npdbm ibm-powervs Running v1.34.7 ibm-powervs-md-0-6b4d67ccf4-v7xv9 ibm-powervs Provisioning v1.34.7 -
Once the new node joins, both pods should be running:
kubectl get nodes NAME STATUS ROLES AGE VERSION ibm-powervs-control-plane-pgwmz Ready control-plane 92m v1.34.7 ibm-powervs-md-0-n8c6d Ready <none> 42s v1.34.7 ibm-powervs-md-0-qch8f Ready <none> 85m v1.34.7 kubectl get pods NAME READY STATUS RESTARTS AGE busybox-deployment-7c87788568-qhqdb 1/1 Running 0 19m busybox-deployment-7c87788568-t26bb 1/1 Running 0 18m -
Delete the deployment and observe the autoscaler scale the node back down:
kubectl delete deployment/busybox-deployment kubectl get nodes NAME STATUS ROLES AGE VERSION ibm-powervs-control-plane-pgwmz Ready control-plane 105m v1.34.7 ibm-powervs-md-0-qch8f Ready <none> 98m v1.34.7