Basic Swimlane Platform Installer (TPI) Troubleshooting Checklist
The following is a list of initial steps when troubleshooting SPI issues:
- Ensure Swimlane is running inside the Kubernetes cluster:
# Swimlane Health Check
# Web Service Endpoints
export NS=<SWIMLANE-NAMESPACE>
$ kubectl -n $NS get ep sw-web -o jsonpath='{range.subsets[*].addresses[*]}{.ip}{"\n"}{end}' | while read -r i; do
echo "Testing /nginx-health on $i: "
curl -ks https://${i}:4443/nginx-health
echo
done
# Api Service Endpoints
$ kubectl -n $NS get ep sw-api -o jsonpath='{range.subsets[*].addresses[*]}{.ip}{"\n"}{end}' | while read -r i; do
echo "======================"
echo "sw-api pod $i"
echo "======================"
echo "Test /heath endpoint:"
curl http://${i}:5000/health
echo -e "\\n"
echo "Test /settings/version endpoint:"
curl http://${i}:5000/settings/version
echo -e "\\n"
done
- Ensure firewalld and UFW is disabled. If it is suppose to be enabled then verify as well.
systemctl status firewalld
systemctl status ufw
- Ensure Containerd and Kubelet are running
- Containerd Run:
systemctl status docker
- If the service is in a failed state search the service log for errors
journalctl -fu containerd
- Kubelet. Run
systemctl status kubelet
- If the service is in a failed state search the service log for errors
journalctl -fu kubelet
- Check cluster node health. Run:
kubectl get nodes -owide
- All nodes should have show Ready in the STATUS column.
- If any nodes are not showing as Ready, run kubectl describe node NODENAME and look for relevant issues/errors under Conditions and Events.
- Check pods status. Run:
kubectl get pods -A -owide | grep -v Completed
- Investigate further if any pods have STATUS not equal to βRunningβ OR STATUS is equal to βRunningβ but not all containers within the pod are running (e.g. 0/1, 0/2, 1/2).
- Check the pod logs for issues with
kubectl logs --all-containers PODNAME -n NAMESPACE
- Check the pod describe output for issues with
kubectl describe pod PODNAME -n NAMESPACE
- Check for errors in the events. Run
kubectl get events --sort-by=.metadata.creationTimestamp
- Ensure the infrastructure is configured according to the Swimlane install guide.
The following is a list of resources to check and logs/information to retrieve:
- Request that a support bundle is provided. Follow the below steps in order to generate the Support Bundle:
- See How to generate a Support Bundle for SPI using the Admin Console or CLI via kubelet. This bundle includes both Swimlane and Kubernetes cluster information.
- Use the /usr/local/bin/kubectl-support_bundle binary. This bundle includes both Swimlane and Kubernetes cluster information.
/usr/local/bin/kubectl-support_bundle secret/default/kotsadm-swimlane-platform-supportbundle
- Generate a support bundle when Kubernetes cluster is completely down. This bundle only includes Kubernetes cluster information.use the /usr/local/bin/kubectl-support_bundle binary with an alternative url:
/usr/local/bin/kubectl-support_bundle --interactive=false https://raw.githubusercontent.com/replicatedhq/troubleshoot-specs/main/host/cluster-down.yaml
- In the Ticket, specify which method was used: a) Admin Console b) kubectl cli (secret/default/kotsadm-swimlane-platform-supportbundle) c) binary (secret/default/kotsadm-swimlane-platform-supportbundle) d) binary (https://raw.githubusercontent.com/replicatedhq/troubleshoot-specs/main/host/cluster-down.yaml)
- In addition to the support bundle, get the status and logs of the Docker and Kubelet services
for SERVICES in containerd kubelet;
do echo --- $SERVICES --- ;
systemctl is-active $SERVICES ;
systemctl is-enabled $SERVICES ;
echo "";
done
for SERVICES in containerd kubelet;
do echo --- $SERVICES --- ;
systemctl status $SERVICES ;
systemctl status $SERVICES ;
echo "";
done
journalctl -u kubelet --since "3 hour ago" > /tmp/kubelet-<node-name>-log.txt
journalctl -u containerd --since "3 hour ago" > /tmp/docker-<node-name>-log.txt
- If a support bundle is unable to be generated (the issue is occurring before the SPI is successfully installed and running), or the support bundle generation fails, then run the following commands and provide the output:
kubectl get pods -A -owide | grep -v Running | grep -v Completed
- If any pods are listed:
- Include the pod logs with
kubectl logs --all-containers PODNAME -n NAMESPACE
- Include the pod describe output with
kubectl describe pod PODNAME -n NAMESPACE
kubectl get pods -A -owide
kubectl get deploy -A
kubectl get sts -A
kubectl get pvc -A
kubectl get pv -A
kubectl get svc -A
kubectl get httpproxy -A
kubectl get nodes
- If any nodes have a status other than Ready, also include the describe output of the node with
kubectl describe node NODENAME.
kubectl get events --sort-by=.metadata.creationTimestamp
- Get the status and logs of the Docker and Kubelet services
for SERVICES in containerd kubelet;
do echo --- $SERVICES --- ;
systemctl is-active $SERVICES ;
systemctl is-enabled $SERVICES ;
echo "";
done
for SERVICES in containerd kubelet;
do echo --- $SERVICES --- ;
systemctl status $SERVICES ;
systemctl status $SERVICES ;
echo "";
done
journalctl -u kubelet --since "3 hour ago" > /tmp/kubelet-<node-name>-log.txt
journalctl -u containerd --since "3 hour ago" > /tmp/docker-<node-name>-log.txt
- Provide the following:
- OS and version
- Infrastructure setup:
- Instance provider:
- Bare-metal, VM, AWS instance, Azure instance, GCP instance, etc.
- Instance sizes:
- Memory, cpu, disk/partitions
- Load balancer type and set up.