Pre-upgrade: checks and mitigations
Reference for every check pre-upgrade-standalone.sh scores. The terminal prints FAIL and WARN with a Fix: line. Passed checks are written to the summary unless you pass --verbose.
- Script: validation/pre-upgrade-standalone.sh
- When it runs: on a node or jumpbox with kubectl, before the change window
- FAIL means: do not start the upgrade. Fix the item, then re-run the same script.
- WARN means: review before you start. A warning alone does not block the script exit code.
- SKIP means: the check did not apply on this host (missing path, jumpbox, or not required for this hop).
- NOTE means: recorded in the summary only. It does not change the pass/fail count shown as FAIL or WARN.
Results land in /tmp/pre_upgrade/pre_upgrade_<host>_<timestamp>/. Attach the sibling .tar.gz when you open a case.
Re-run after each fix:
Some checks run only when you pass a flag, only on a Linux node, or only for a given target version. Those conditions are called out on each check.
ο»Ώ
Index
Check label | Typical result | Runs when |
|---|---|---|
CURRENT VERSION | FAIL | Version was not detected or is not X.Y.Z |
ALREADY CURRENT | PASS | Detected version is already at or past the target |
ALLOWED HOP / ALLOWED HOP (INTERMEDIATES REQUIRED) | PASS or FAIL | A current and target version are known |
LICENSE FILE | FAIL | --license-file was passed |
LICENSE ARCHIVED | PASS or FAIL | --license-file was passed |
LICENSE CHANNEL | PASS, WARN, or FAIL | --license-file was passed |
BUNDLE PATH | FAIL | --bundle-dir was passed and the path is missing |
BUNDLE ARCH | PASS or FAIL | --bundle-dir was passed |
AIRGAP ISOLATION / AIRGAP EGRESS | PASS or WARN | Connectivity is offline or airgap |
AIRGAP IMAGE LOAD | NOTE | Connectivity is offline or airgap |
HOST PROXY | PASS or WARN | Always |
NO_PROXY COMPLETE | PASS or FAIL | An HTTP(S) proxy is set in the shell |
BACKUP PATH EXISTS | FAIL | --backup-path was passed and the directory is missing |
BACKUP PATH WRITABLE (root) | PASS or FAIL | --backup-path was passed |
BACKUP PATH UID 1001 | PASS, WARN, or FAIL | --backup-path was passed and root can write |
Operating system | PASS or WARN | Script is running on a Linux node |
FIPS mode | PASS or WARN | Linux node |
CPU CORES (>= 8) | PASS or FAIL | Linux node |
MEMORY (>= 32GB) | PASS or FAIL | Linux node |
KUBELET SERVICE | PASS or FAIL | Linux node |
CONTAINERD SERVICE | PASS or FAIL | Linux node |
HEADROOM / and related paths | PASS, WARN, or FAIL | Linux node, and the path exists |
CRI API v1 | SKIP | No crictl and no containerd socket |
CRI / CONTAINERD | PASS or NOTE | crictl is present |
KUBELET CRI CONFIG | PASS or WARN | /var/lib/kubelet/config.yaml exists |
KERNEL VS TARGET K8s | FAIL | --target-k8s is 1.33+ and the kernel is 4.18 or older |
RHEL 8 SUPPORT | WARN or FAIL | /etc/redhat-release is RHEL 8 |
NOFILE / INOTIFY | PASS or FAIL | Embedded cluster and target is 26.3.3 or later |
HOST RESOURCE CHECKS | SKIP | Script is not on a Linux node |
KUBECTL | FAIL | kubectl is not on PATH |
K8S NODES (All Ready) | PASS or FAIL | kubectl works |
NODE PRESSURE | PASS, FAIL, or SKIP | kubectl works |
CLUSTER STATUS | PASS or WARN | kubectl works |
POD STATUS | PASS or FAIL | kubectl works |
STALE COMPLETED JOBS | PASS or WARN | kubectl works |
FAILED JOBS | FAIL | A Job in the app namespace has failed > 0 |
MIGRATION / MONITOR JOBS | NOTE | Named migration or monitor Jobs exist |
TLS SECRET gotenberg-cert / notifications-api-cert | PASS or WARN | kubectl works |
ETCD HEALTH | PASS, FAIL, or SKIP | kubectl works |
VELERO BSL | PASS, WARN, or FAIL | kubectl works |
VELERO LAST SNAPSHOT | PASS, WARN, FAIL, or NOTE | kubectl works |
RABBITMQ | SKIP | No Running RabbitMQ pod |
RABBITMQ QUEUES | PASS or WARN | A Running RabbitMQ pod was found |
RABBITMQ SPREAD | PASS, FAIL, or NOTE | A Running RabbitMQ pod was found |
RABBITMQ PARTITION | FAIL | cluster_status mentions a partition or store error |
MONGODB | WARN | No Running mongo-0 / sw-mongo-0 |
MONGODB VERSION | PASS or WARN | Mongo pod is Running |
MONGODB FCV (7.0) | PASS, FAIL, or SKIP | Mongo pod is Running |
MONGODB RECORD COUNTS | PASS or NOTE | Mongo pod is Running |
MONGO REPLICA SET | PASS or FAIL | Replica-set status was read |
MONGO MEMBER COUNT | FAIL | HA pod count does not match rs.status() |
FCV VS TARGET MONGO 8 | PASS or FAIL | A hop in the path installs Mongo 8 |
MONGO RETENTION / PRUNE JOB | NOTE | Mongo pod is Running |
The script also writes upgrade_plan.txt when the hop chain resolves. That file is the install sequence (Mongo, Kubernetes, KOTS, Velero, FCV, installer command). It is not a separate pass/fail check.
These helpers exist in the file and are not scored on a normal run: custom CA detection, online egress to kurl.sh / Replicated, required-directory free-space minimums, and the HA join-order reminder. Disk usage percentage checks (local_pct_check) are defined and not called. Mongo StatefulSet podManagementPolicy is copied into sts_drift.txt and is not scored.
ο»Ώ
Version and upgrade path
ο»Ώ
CURRENT VERSION
Result: FAIL
What it checks: The running KOTS app version. The script reads apps.kots.io status.currentRelease.versionLabel, then kubectl kots get apps, then the image tag on swimlane-web, turbine-api, or swimlane-api. A label such as 26.2.3_1177 is treated as 26.2.3.
Why it matters: Without a current version the hop chain cannot be built. Starting from the wrong version skips a required intermediate.
Mitigation:
- In the KOTS Admin Console, open Version History and copy the installed version (ignore a _sequence suffix).
- Or run: kubectl get apps.kots.io -A
- Re-run with an explicit version: sudo ./pre-upgrade-standalone.sh --current-version 25.5.1 --target-version 26.4.1
- If the value is not three numbers (X.Y.Z), correct it and re-run. Do not start the upgrade until this check passes or ALREADY CURRENT / ALLOWED HOP passes instead.
ο»Ώ
ALREADY CURRENT
Result: PASS
What it checks: Parsed current version is greater than or equal to the target (default Turbine 26.4.1, Swimlane 10.24.0).
Mitigation: None. Do not start an upgrade to a version you are already on. If the detected version is wrong, re-run with --current-version and --target-version.
ALLOWED HOP / ALLOWED HOP (INTERMEDIATES REQUIRED)
Result: PASS when a documented chain exists. FAIL when the target is missing from the hop table, the target is not newer than current, a prerequisite is missing, or the resolver loops.
What it checks: The embedded hop table (product version, minimum previous version, Mongo, Kubernetes, KOTS, Velero, required Mongo FCV, and whether a release is paused). Paused releases (today: Turbine 25.4.1) are replaced with a safe same-train release when one exists.
Why it matters: Jumping versions (for example 25.3.2 straight to 25.5.1, or 25.0.8 straight to 25.3.2) leaves Mongo, Kubernetes, or KOTS on a combination the next installer does not support.
What you will see on PASS: A path such as 25.5.1 -> 26.1.4 -> 26.2.3 -> 26.3.3 -> 26.4.1, plus a short plan. The full plan is upgrade_plan.txt. If there are more than two versions in the chain, the label is ALLOWED HOP (INTERMEDIATES REQUIRED). Use every hop. Do not jump to the last version.
Mitigation when it FAILs:
- Confirm current and target with --current-version and --target-version.
- Follow the documented intermediate for that product. Turbine index: https://docs.swimlane.com/turbine-upgrade-guide . Swimlane: https://docs.swimlane.com/10x-upgrade-guides .
- If the guide says to pause a release (Turbine 25.4.1), use the replacement the script prints, not the paused build.
- Re-run. Do not start the installer until ALLOWED HOP passes.
- On each hop, before Deploy, delete completed Jobs (see STALE COMPLETED JOBS) and take a fresh Velero backup.
ο»Ώ
Optional inputs
These run only when the matching flag is set.
LICENSE FILE
Flag: --license-file PATH
Result: FAIL if the file is missing.
Mitigation:
- Copy the current .yaml license onto this host.
- Re-run with --license-file /path/to/license.yaml.
- Do not start the install or upgrade with an archived or wrong-channel license.
ο»Ώ
LICENSE ARCHIVED
Result: FAIL if the file contains archived: true or isArchived: true. PASS otherwise.
Why it matters: An archived license returns 403 during install or upgrade.
Mitigation:
- Request a current license from Swimlane.
- Replace the file and re-run. Do not proceed while this check fails.
ο»Ώ
LICENSE CHANNEL
Result: FAIL when --license-channel is set and channelName / channel in the file does not match. WARN if the channel field cannot be parsed. PASS when a channel is present and matches (or no expected channel was passed).
Mitigation:
- Open the license YAML and read channelName.
- Use a license for the channel you are installing (stable vs unstable).
- Re-run with --license-channel set to that name. Do not proceed on FAIL.
ο»Ώ
BUNDLE PATH / BUNDLE ARCH
Flag: --bundle-dir PATH
What it checks: The directory exists, and its file names are not the wrong CPU architecture for this node (x86_64 vs aarch64 / arm64 / amd64).
Why it matters: The wrong package installs an aarch64 kubelet on an x86_64 node (or the reverse). The node then cannot start pods.
Mitigation:
- Confirm architecture: uname -m.
- Stage the matching online or airgap bundle for that architecture.
- Re-run with --bundle-dir pointed at that directory. Do not install a bundle that fails this check.
ο»Ώ
AIRGAP ISOLATION / AIRGAP EGRESS
Runs when: --connectivity offline or airgap, or when https://kurl.sh is not reachable and connectivity was not set.
What it checks: Whether this host can open https://registry.k8s.io, quay.io, ghcr.io, and sigstore.dev (4 second timeout).
- PASS (AIRGAP ISOLATION): the host cannot reach that registry. Expected for airgap.
- WARN (AIRGAP EGRESS): the host can reach a public registry.
Why it matters: An airgap install must not pull pause or Wolfi images from the internet. A node that can reach a public registry may still fail later if those images were never loaded locally.
Mitigation:
- Load the full upgrade bundle, including pause and both -wolfi and plain tags, onto every node before Deploy.
- Treat the warning as a reminder, not as permission to pull from the internet during the window.
- Re-run if you change proxy or firewall rules and need a fresh isolation result.
ο»Ώ
AIRGAP IMAGE LOAD
Result: NOTE only.
Mitigation: Before Deploy, confirm the airgap images from this bundle are present on every node. A missing pause or -wolfi tag causes ImagePullBackOff after KOTS reports success.
HOST PROXY / NO_PROXY COMPLETE
What it checks: HTTP_PROXY / HTTPS_PROXY (any case) in this shell, and whether NO_PROXY / no_proxy mentions cluster.local, an RFC1918 or 100. range, and localhost.
- PASS HOST PROXY: no HTTP(S) proxy in this shell.
- WARN HOST PROXY: a proxy is set.
- FAIL NO_PROXY COMPLETE: a proxy is set and NO_PROXY is missing cluster DNS, pod/service CIDR, or localhost.
Why it matters: Incomplete NO_PROXY sends in-cluster traffic out through the proxy. API calls, image pulls, and Mongo/RabbitMQ connections then fail mid-upgrade.
Mitigation:
- Put proxy settings in the product or installer config (se.yaml or KOTS). Do not rely on a one-off kubectl patch. An upgrade will overwrite it. Do not use the Beta proxy section on current Turbine.
- When a proxy is required, set NO_PROXY to include the pod and service CIDRs, .svc.cluster.local, node names and IPs, localhost, and the in-cluster registry.
- Re-run in the same shell you will use for the installer. Do not start while NO_PROXY COMPLETE fails.
ο»Ώ
BACKUP PATH EXISTS / BACKUP PATH WRITABLE (root) / BACKUP PATH UID 1001
Flag: --backup-path PATH
What it checks: The directory exists, root can create a file there, and UID 1001 (Restic) can create a file there.
Why it matters: Velero backups fail when the NFS or local path is mounted but not writable by Restic. An upgrade you cannot restore is a failed window.
Mitigation:
- Mount the NFS or local backup path before the window. Confirm it with a write, not only with mount.
- Fix the export and permissions so UID 1001 can write. Root write access is not enough.
- If the script warns that it could not switch to UID 1001, confirm that user can write before you continue: touch /path/to/backup/.probe && rm -f /path/to/backup/.probe runuser -u '#1001' -- touch /path/to/backup/.probe && rm -f /path/to/backup/.probe
- Re-run with the same --backup-path. Do not start while either write check fails.
ο»Ώ
Host (Linux node only)
If /proc/meminfo is missing, the script records HOST RESOURCE CHECKS as SKIP and tells you to run it on each node. Do that before the window. A jumpbox-only run does not prove CPU, RAM, disk, kubelet, or containerd.
Operating system
Result: PASS for Red Hat, Rocky, Ubuntu, AlmaLinux, Oracle Linux, or CentOS Stream. WARN for any other NAME in /etc/os-release.
Mitigation: Confirm the OS and version are supported for the Turbine or Swimlane release you are moving to. Ubuntu 20.04 support ends after Swimlane 10.24.0. RHEL 8 is covered separately by RHEL 8 SUPPORT. Re-run after any OS change.
FIPS mode
What it checks: /proc/sys/crypto/fips_enabled and, when present, update-crypto-policies --show.
- WARN: FIPS is enabled (1).
- PASS: FIPS is off, or the sysctl file is not on this kernel. Off is normal for most Turbine sites.
Why it matters: FIPS can break kURL, containerd, and Mongo TLS during the upgrade.
Mitigation:
- Confirm with the customer whether this site is required to run FIPS.
- If FIPS is on by mistake, disable it using the OS vendorβs procedure and reboot before the window.
- If FIPS is required, do not start until Support has confirmed this release and this installer path are validated for that policy.
- Re-run on every node.
ο»Ώ
CPU CORES (>= 8)
What it checks: nproc, or processor lines in /proc/cpuinfo. Default minimum is 8 (MIN_CPU_CORES).
Why it matters: Undersized nodes stall or run out of memory while new images start.
Mitigation:
- Confirm with nproc.
- Add vCPU on every node that will run the control plane or the application. For a 3-node HA cluster, size every control-plane node.
- Re-run. Do not start while this fails.
ο»Ώ
MEMORY (>= 32GB)
What it checks: MemTotal in /proc/meminfo, rounded to the nearest gigabyte. Default minimum is 32 (MIN_MEM_GB).
Why it matters: Below 32 GB, Postgres and other components time out during the upgrade. A 32 GB VM can report 31 GB because of kernel rounding. If free -h shows about 32 GB, confirm the hypervisor size before you treat 31 GB as a real shortfall.
Mitigation:
- On the same node: free -h awk '/MemTotal/ {printf "%.1f GB\n", $2/1024/1024}' /proc/meminfo
- Increase the node or VM to at least 32 GB. For HA, do this on every control-plane node.
- Reboot only if the platform requires it for the new size to appear.
- Re-run. Do not start the installer or KOTS Deploy until Memory passes.
ο»Ώ
KUBELET SERVICE
Result: PASS when systemctl is-active --quiet kubelet. FAIL otherwise.
Mitigation:
- systemctl status kubelet and journalctl -u kubelet -e.
- Start or repair kubelet so the node is Ready.
- Re-run. Do not upgrade a node whose kubelet is down.
ο»Ώ
CONTAINERD SERVICE
Result: PASS when systemctl is-active --quiet containerd. FAIL otherwise.
Mitigation:
- systemctl status containerd and journalctl -u containerd -e.
- Start or repair containerd. Image pulls and new pods fail until it is healthy.
- Re-run before you continue.
ο»Ώ
HEADROOM / , /var , /var/lib/containerd , /var/openebs , /var/lib/kubelet
What it checks: Filesystem use on each path that exists. Defaults: WARN at 75% (HEADROOM_WARN_PCT), FAIL at 84% (HEADROOM_FAIL_PCT). A missing path is SKIP.
Label | Path |
|---|---|
HEADROOM / | / |
HEADROOM /var (containerd blobs) | /var |
HEADROOM /var/lib/containerd | /var/lib/containerd |
HEADROOM /var/openebs | /var/openebs |
HEADROOM /var/lib/kubelet | /var/lib/kubelet |
Why it matters: The upgrade unpacks new image layers. A volume already at or over 84% used has blocked preflight retries. OpenEBS filling up takes Mongo and other data disks offline mid-hop.
Mitigation:
- See what is full: df -h / /var /var/lib/containerd /var/openebs /var/lib/kubelet
- Free space before the window: crictl rmi --prune journalctl --vacuum-size=500M
- On /var/openebs, also look for unused PVCs and unbounded Mongo collections (see MONGO RETENTION / PRUNE JOB). Expanding the volume is safer than deleting data you cannot identify. Past recoveries have needed on the order of +300 GB on /var/openebs.
- Re-run. WARN means prune now so the hop does not cross 84% in the middle of the window. FAIL means do not start until usage is under 84%.
ο»Ώ
CRI API v1 / CRI / CONTAINERD / KUBELET CRI CONFIG
What it checks: crictl info against containerd, and whether /var/lib/kubelet/config.yaml mentions a CRI v1 endpoint, RuntimeRequestTimeout, or container-runtime-endpoint.
- SKIP CRI API v1: no crictl and no /run/containerd/containerd.sock.
- PASS CRI / CONTAINERD: crictl info looks like containerd.
- NOTE CRI / CONTAINERD: crictl info did not look healthy.
- WARN KUBELET CRI CONFIG: kubelet config exists but does not mention those CRI settings.
- PASS KUBELET CRI CONFIG: those settings are present.
Why it matters: On existing Kubernetes, kubelet must speak containerd CRI API v1. dockershim will not run this productβs current nodes.
Mitigation:
- crictl info should succeed and name containerd.
- Review /var/lib/kubelet/config.yaml and point kubelet at the containerd CRI v1 socket, not dockershim.
- Restart kubelet only after the endpoint is correct, then confirm the node returns to Ready.
- Re-run. Clear the NOTE or WARN before an existing-Kubernetes upgrade.
ο»Ώ
KERNEL VS TARGET K8s
Runs when: --target-k8s is set to 1.33 or newer and uname -r is 4.x with minor 18 or lower (or any kernel older than 5). If the kernel is new enough, this check is not printed.
Why it matters: Kubernetes 1.33+ does not run on kernel 4.18. Embedded hops from 26.1.3 upward land on those Kubernetes versions.
Mitigation:
- uname -r on every node.
- Upgrade the kernel and OS so each node meets the target Kubernetes version. Per-node kubelet, containerd, and kernel must match.
- Re-run with the same --target-k8s. Do not start this hop while the check fails.
ο»Ώ
RHEL 8 SUPPORT
Runs when: /etc/redhat-release contains release 8.
- FAIL if the target version is 26.1.3 or newer.
- WARN if the target is older. The current hop may be allowed, and later hops will block.
Why it matters: Embedded Kubernetes 1.33+ drops kernel 4.18, which is what RHEL 8 ships. Later 26.1.x, 26.2.3, 26.3.3, and 26.4.1 embedded hops are not supported on RHEL 8.
Mitigation:
- Plan a move to RHEL 9 on every embedded node before the first blocked hop.
- If this check already FAILs, upgrade the OS first. Do not run the 26.1.3+ installer on RHEL 8.
- Re-run on every node after the OS change.
ο»Ώ
NOFILE / INOTIFY
Runs when: deployment is embedded and the target version is 26.3.3 or newer.
What it checks:
Setting | Required |
|---|---|
containerd LimitNOFILE | 1048576 |
fs.inotify.max_user_watches | 524288 |
fs.inotify.max_user_instances | 8192 |
Why it matters: Kubernetes 1.35-era nodes exhaust inotify watches and file descriptors during image load and pod start. The 26.3.3 upgrade guide documents the host limits.
Mitigation:
- On every node, apply the 26.3.3 host limits (99-kubernetes.conf and the containerd override) from the 26.3.3 upgrade guide.
- sysctl --system, then restart containerd and kubelet.
- Confirm: systemctl show containerd -p LimitNOFILE sysctl fs.inotify.max_user_watches fs.inotify.max_user_instances
- Re-run. Do not Deploy 26.3.3 or later until this passes on every node.
ο»Ώ
Cluster
ο»Ώ
KUBECTL
Result: FAIL when kubectl is not on PATH. The script then stops. Host checks above may already have run.
Mitigation: Install kubectl or run from a host that can reach the API, then re-run. The rest of the cluster checks do not run until this passes.
K8S NODES (All Ready)
What it checks: Every nodeβs status column is Ready.
Mitigation:
- kubectl get nodes -o wide
- Uncordon a drained node, or replace a dead node, until every node is Ready.
- Re-run. Do not upgrade with a NotReady node.
ο»Ώ
NODE PRESSURE
What it checks: Node conditions DiskPressure, MemoryPressure, and PIDPressure. FAIL if any is True. SKIP if conditions could not be read.
Why it matters: Pressure during the hop evicts pods and causes image-pull failures.
Mitigation:
- kubectl describe nodes and read the pressure condition message.
- Free disk or memory (see the HEADROOM checks) or add capacity.
- Re-run after every condition is False.
ο»Ώ
CLUSTER STATUS
Result: PASS when kubectl cluster-info output contains βrunningβ. WARN otherwise.
Mitigation: Open ${OUT_DIR}/cluster_info.txt. Run kubectl cluster-info dump and confirm the API server is healthy before you continue.
POD STATUS
What it checks: Pods in all namespaces. A pod is healthy when it is Completed/Succeeded, or Running/Completed with ready count equal to desired count. Anything else is listed in unhealthy_pods.txt.
Mitigation:
- kubectl get pods -A | grep -v Running | grep -v Completed
- Bring CrashLoop, 0/1, and Error pods back to Running. Do not delete data-plane pods (Mongo, Postgres, etcd, RabbitMQ) to clear the check.
- Re-run. Do not start while unhealthy pods remain.
ο»Ώ
STALE COMPLETED JOBS
What it checks: Jobs in the app namespace with status.successful=1.
- WARN: one or more successful Jobs remain.
- PASS: none remain. The summary still says to delete them again immediately before Deploy.
Why it matters: Completed Jobs are immutable. The next apply fails when the new spec cannot replace them.
Mitigation:
- Immediately before each Deploy: kubectl delete jobs --field-selector=status.successful=1 -n <namespace>
- Re-run if you want a clean warning count. Deleting them again right before Deploy is still required.
ο»Ώ
FAILED JOBS
Result: FAIL only when at least one Job in the app namespace has status.failed > 0. No PASS line when there are none.
Mitigation:
Delete failed Jobs before apply, then re-run. Do not start while this check fails.
MIGRATION / MONITOR JOBS
Result: NOTE when any of these Jobs exist: correlation-migrations, swimlane-tenant-migrations, turbine-identity-migrations, content-lib-monitor.
Mitigation: Delete those Jobs before Deploy so the new spec can apply. They are called out separately because leaving them in place blocks the upgrade even when they previously succeeded.
TLS SECRET gotenberg-cert / TLS SECRET notifications-api-cert
Result: PASS if the secret exists in the app namespace. WARN if it does not.
Why it matters: Target manifests may mount these secrets. A missing secret becomes CreateContainerConfigError after Deploy.
Mitigation:
- Confirm whether the next hopβs manifest still creates the secret.
- If the running release already depends on it, restore or apply the secret before Deploy.
- Re-run. A WARN here is a review item: do not Deploy if you have confirmed the next version mounts a secret that is absent.
ο»Ώ
ETCD HEALTH
What it checks: etcdctl endpoint health inside each kube-system pod labeled component=etcd, using the in-pod etcd certificates. Output is etcd_health.txt.
- SKIP: no etcd pods (typical of some managed or external control planes).
- PASS: every memberβs output contains healthy and not unhealthy.
- FAIL: any member is unhealthy.
Why it matters: Losing quorum during an upgrade takes the API server down. There is no safe way to continue a Kubernetes hop from an unhealthy etcd.
Mitigation:
- Read etcd_health.txt in the output directory.
- kubectl -n kube-system logs <etcd-pod> for each unhealthy member.
- Restore quorum so every member reports healthy. If you are not sure the member can rejoin, stop and contact Support before removing members or restoring a snapshot.
- Re-run. Do not start the upgrade while any member is unhealthy.
ο»Ώ
VELERO BSL
What it checks: backupstoragelocations.velero.io across all namespaces.
- PASS: a location is Available or Ready.
- FAIL: locations exist and none are Available.
- WARN: no BackupStorageLocation custom resources.
Why it matters: An upgrade that cannot be restored is a failed window.
Mitigation:
- kubectl get backupstoragelocations.velero.io -A -o wide
- Fix the bucket, credentials, or restic repository until phase is Available. Also confirm the backup path UID 1001 check if you passed --backup-path.
- If Velero is not installed, confirm some other tested restore path before you continue. The WARN does not block the script, and it does mean you currently have no Velero restore.
- Re-run. Do not start on FAIL.
ο»Ώ
VELERO LAST SNAPSHOT
What it checks: velero get backups when the CLI exists, otherwise Backup custom resources.
- PASS: a Completed or Available backup exists. Take a new one before the window anyway.
- FAIL: the CLI exists and no backup is Completed.
- WARN: Backup CRs exist but the CLI was not used to prove one is Completed.
- NOTE: neither the CLI nor Backup CRs were found.
Mitigation:
Wait until the backup is Completed. If it stays InProgress or Failed, check restic locks and ownership (UID 1001). Re-run. Do not start on FAIL. On PASS, still create a fresh backup at the start of the window.
ο»Ώ
RabbitMQ
Skipped entirely (RABBITMQ) when no Running RabbitMQ pod exists in the app namespace. The script prefers a Running rabbitmq-server-* pod, then any Running rabbitmq pod whose name does not contain operator.
RABBITMQ QUEUES
What it checks: rabbitmqctl list_queues. WARN when any queue has more than 100 messages. PASS otherwise.
Mitigation: Drain the backlog (let consumers catch up, or correct the consumer that is stuck) before the window. A deep queue during a restart delays playbooks and can look like an upgrade failure. Re-run and confirm no queue is over 100 messages.
RABBITMQ SPREAD
What it checks: When more than one RabbitMQ pod exists, they must not all sit on the same node.
- FAIL: multiple members, one node (the summary calls this fake HA).
- PASS: multiple pods on more than one node.
- NOTE: a single RabbitMQ pod, which is expected on a single-node install.
Mitigation: Spread rabbitmq-server pods across control-plane nodes (pod anti-affinity or an explicit move), wait until they are Running, and re-run. Do not upgrade an HA cluster whose brokers all share one node.
RABBITMQ PARTITION
Result: FAIL only when rabbitmqctl cluster_status mentions partition, a Khepri error, or an Mnesia error. No PASS line when the status is clean.
Why it matters: A split-brain broker drops or duplicates messages across the restart the upgrade will trigger.
Mitigation:
- Read rabbitmq_cluster.txt in the output directory.
- Heal the partition so cluster_status lists the running nodes and does not show a split-brain. Stop and contact Support if you are choosing which side of a partition to keep.
- Re-run. Do not continue while this check fails.
ο»Ώ
MongoDB
MONGODB warns when no Running mongo-0 or sw-mongo-0 pod exists in the app namespace. Fix that before you trust any later Mongo result:
The script authenticates as Admin using the password from mongo-admin, swimlane-sw-mongo-admin, or mongodb-admin, and tries mongosh then mongo, with TLS and invalid certificates allowed.
MONGODB VERSION
Result: PASS when db.version() looks like major.minor.... WARN when it cannot be parsed.
Mitigation: Exec into the pod and run db.version(). Compare it to the Mongo version printed on the next hop in upgrade_plan.txt. A WARN here means the script could not see the version, so FCV checks below may also be wrong. Re-run after mongosh works.
MONGODB FCV (7.0)
Runs when: any hop after the current version has a featureCompatibilityVersion in the hop table (Turbine hops that require 7.0 before the Mongo 8 family, and Swimlane rows that list an FCV). SKIP when this chain does not require it.
Result: PASS when the parameter output contains 7.0. FAIL otherwise.
Why it matters: FCV must be set to 7.0 only after Mongo is already 7.0.x, and before Mongo 8. Leaving FCV at 6.0 and installing Mongo 8 can leave the replica set unable to start.
Mitigation:
- Confirm a PRIMARY exists and no member is RECOVERING (see MONGO REPLICA SET).
- Confirm db.version() is already 7.0.x. Do not set FCV while Mongo is still 6.x.
- On the primary only: db.adminCommand({ setFeatureCompatibilityVersion: "7.0" }) On Mongo 7.0.6 and later you may need:
- Read FCV again. Re-run the script. Do not start the Mongo 8 hop until this check passes.
Confirm by hand:
ο»Ώ
MONGODB RECORD COUNTS
Result: PASS when Swimlane.Records or SwimlaneHistory.Records returns a number. NOTE when the count fails. Tenant data may live in SwimlaneAccountDB_* instead, which this check does not count.
Mitigation: No change is required for a NOTE. Record the counts from mongo_info.txt before the window so you can compare them after the hop. Contact Support if a PASS count is far from what the customer expects and you have not yet started.
MONGO REPLICA SET
What it checks: rs.status().members[].stateStr.
- FAIL: any member is RECOVERING, DOWN, or REMOVED.
- PASS: the status contains PRIMARY and the failure cases below did not match.
Mitigation:
- On the primary, rs.status() and rs.printSecondaryReplicationInfo().
- Free disk if a member is stuck because the volume is full (HEADROOM /var/openebs).
- Bring RECOVERING or removed members back, and wait out oplog lag, until rs.status() shows a healthy PRIMARY and the expected secondaries.
- Stop and contact Support before you force a reconfig or rs.stepDown you cannot undo.
- Re-run. Do not continue while this fails.
ο»Ώ
MONGO MEMBER COUNT
Result: FAIL when there are at least three pods matching mongo- and rs.status() lists fewer than three members.
Why it matters: The StatefulSet can look like HA while the replica set is missing a member. A later step-down then has no majority.
Mitigation: Add the missing replica-set members so the set matches the HA topology. Confirm rs.status().members.length equals the Running mongo pod count, then re-run.
FCV VS TARGET MONGO 8
Runs when: a later hop in the chain lists a Mongo version starting with 8..
- PASS: FCV output contains 7.0.
- FAIL: target Mongo is 8.x and FCV is not 7.0.
Mitigation: Same as MONGODB FCV (7.0). Set FCV to 7.0 only after db.version() is already 7.0.x. Do not hop to Mongo 8 while this fails. FCV 6.0 into Mongo 8 can brick the set.
MONGO RETENTION / PRUNE JOB
Result: NOTE. The script lists collection names in SwimlaneTurbineEngine matching Prune or PlaybookRun, and tells you to confirm Hangfire PruneBackgroundJobPlaybookRunJobs last succeeded.
Why it matters: Unbounded playbook-run history fills OpenEBS during the upgrade.
Mitigation: In Hangfire, confirm PruneBackgroundJobPlaybookRunJobs succeeded recently. If the job is failing, fix it and let it prune before the window, then re-check HEADROOM on /var/openebs.
ο»Ώ
After the script finishes
The last lines are the counts and one of:
- Do not start the upgrade yet. Fix every FAIL and re-run.
- No blockers. Take a fresh backup, delete completed Jobs, then follow the upgrade path. Review every WARN before you start.
Exit code 1 means at least one FAIL. Exit code 0 can still include WARN and NOTE items.