Pre-Upgrade Health Check
Node Status
- Check that all Nodes are in Ready state
kubectl get nodes -owide
- If any nodes are in NotReady, SchedulingDisabled, etc then it would be recommended to review and fix them before upgrading.
Pod Status
- Check that all pods are in Running state and the pod values are matching 1/1, 3/3, etc.
kubectl get pods -A
- If any pods are not running i.e, 0/1, 2/3, CrashLoopBackoff, Error, etc then it would be advised to review the issue and fix it prior to upgrading
Cluster Status
- Check that Kubernetes is up and running and healthy.
kubectl cluster-info
- If Kubernetes indicates that it is not healthy then review and establish where the issue is prior to upgrading.
etcd health
- etcd needs to be in a stable state prior to upgrading any issues
for pod in $(kubectl get pods -l component=etcd -n kube-system \
-o jsonpath='{.items[*].metadata.name}')
do
echo "### etcd pod : ${pod} ###"
kubectl -n kube-system exec ${pod} -- /bin/sh \
-c "ETCDCTL_API=3 etcdctl \
--cert=/etc/kubernetes/pki/etcd/server.crt \
--key=/etc/kubernetes/pki/etcd/server.key \
--cacert=/etc/kubernetes/pki/etcd/ca.crt \
endpoint health"
done
- Should any other states be listed on the output then it would be advised to investigate the logs of the pod to establish the issue
Hangfire Status
- Check that Hangfire is running and there is no large amount of Enqueued Jobs
- It would be recommended to grab a screenshot of the Hangfire Dashboard prior to upgrading to check the amount processed for Deleted Jobs
Logging Status
- Check the Logging section within the UI for Error Messages
- It would be recommended to grab a screenshot of the Logging section to ensure that any post-upgrade issues are correctly identified
Velero Status and Backups
- Ensure a backup is taken prior to upgrading should any issues arise during upgrade to be able to perform a restore.
velero get backups
Only required for versions < 10.11.x
Rook-Ceph Health Status
- Check the status of Rook-Ceph for any Health_Warn messages:
kubectl -n rook-ceph exec deploy/rook-ceph-operator -- ceph status
- Check the status of Rook-Ceph for Health Information:
kubectl -n rook-ceph exec deploy/rook-ceph-operator -- ceph health detail
- If Rook-Ceph health is not in working health state then the upgrade should not be attempted and Rook-Ceph should be healthy prior to attempting.
MongoDB Records Amount
- Establishing the amount of Swimlane Records the customer has:
use Swimlanedb.Records.find().count()
- Establishing the amount of SwimlaneHistory Records the customer has:
use SwimlaneHistorydb.Records.find().count()
MongoDB Collection Sizes
- Please refer to Collect Mongo DB Collection and Index Sizes Across All DatabasesCollect Mongo DB Collection and Index Sizes Across All Databasesο»Ώ
Download the ConfigValues File If kotsadm goes in CrashLoopBackOff while upgrading, its good to collect below information while have pre-upgrade checks call.
- Downloaded kots plugin (Replace:Β APP_NAMESPACE )
kubectl kots download APP_SLUG -n APP_NAMESPACE --dest ./manifests --overwrite --decrypt-password-values
with the namespace on the cluster where you installed your application.APP_SLUG with the slug of the application Example:
kubectl kots download swimlane-platform -n default --dest ./manifests --overwrite --decrypt-password-values