Node and cluster maintenance on K8s Helm deployment
At some point, you'll need to carry out maintenance on a node (such as for a kernel upgrade, to apply a security patch, to upgrade the operating system, perform hardware maintenance, take a vm snapshot, etc.) which may require a single node or even a cluster-wide shutdown or reboot. It's critical that these events are handled gracefully in a Helm deployment of Swimlane on Kubernetes (K8s).
Definition:
- Node: a virtual or physical server
- Can be a master or a worker
- Cluster: A group of interconnected nodes
Single node maintenance
In brief, maintenance on a single node will include:
- Gracefully shutdown K8s components
- Perform node maintenance
- Restart the node (if necessary)
- Start K8s components again
NOTE: If additional nodes require maintenance then it is recommended to perform the steps below for each node one at a time or refer to the "Multiple nodes maintenance" section for cluster-wide maintenance.
- Connect (SSH) to the node which requires maintenance.
- Drain the node. The kubectl drain operation reschedules all the pods from the node onto another node. To do so, it will 1) tag the node unschedulable in order to prevent new pods from being scheduled (equal to kubectl cordon <node-name>) and 2) evicts or deletes all pods running on the node. For pods with a replica set (like mongo), the pod will be replaced by a new pod that will be scheduled to a new node. Draining a node wonβt cause any downtime as long as Mongoβs resource utilization isnβt so high that it canβt run alongside another mongo pod during the maintenance window. TIP: Use kubectl get nodes -owide to lookup the node name.
NOTE: The flags 'ignore-daemonsets' and 'delete-local-data' deletes data that is not persisted to an actual disk somewhere such as emptyDir. The command above may also need '--force' if there exists pods not managed by ReplicaSet, Job, DaemonSet or StatefulSet.
- Perform the necessary node maintenance (restart if needed)
- Make the node schedulable again
Run the following command to uncordon the node which tags the node to be schedulable again.
- Repeat the process above for any additional nodes that need maintenance work.
Multiple nodes maintenance
It may become necessary to perform maintenance on all the nodes in a cluster together. In such a scenario, follow the steps outlined below to gracefully stop K8s resources, perform the maintenance, and then restart K8s resources. NOTE: This method can also be used if you need to completely shutdown all the nodes (for example to save development cost when using a cloud provider).
Step 1: Shutdown
- Ensure you're on a master node or another machine which has kubectl configured to access the cluster.
- Stop the Swimlane Application TIP: <release-name> : Run the command helm ls to lookup the Helm release name. <chart-name> : will be swimlane/swimlane if working with the chart remotely otherwise swimlane
- Make nodes unschedulable The kubectl cordon <node-name> command will tag the nodes unschedulable and prevent any new pods from being scheduled on other nodes. NOTE: The commands below is for a 6 node cluster. Please update the code to reflect the number of nodes in your cluster.
- Drain the nodes Gracefully terminate all pods running on the node NOTE: The commands below is for a 6 node cluster. Please update the code to reflect the number of nodes in your cluster.
Step 2: Perform maintenance
- Perform the necessary maintenance on all the nodes
- Gracefully shutdown and power down all nodes
Step 3: Startup
- Power up all the nodes
- Ensure you're on a master node or another machine which has kubectl configured to access the cluster
- Make nodes schedulable again The kubectl uncordon <node-name> command will tag the nodes schedulable and ready to take on new pods
- Start the Swimlane Application TIP: <release-name> : Run the command helm ls to lookup the Helm release name. <chart-name> : will be swimlane/swimlane if working with the chart remotely otherwise swimlane
Step 4: Confirmation
Use the following methods to confirm that the cluster and application is back online.
- Run kubectl get nodes
- Run kubectl get pods -A
- Login to Swimlane Application