In some cases, velero backups or restores fail without clear error messages. Use these steps to gather logs, diagnose issues, and understand the root cause.
Prerequisite
Set the namespace. For embedded cluster, the value should be "default"
export NS=<your-namespace>
1. Collect Velero debug logs
To generate detailed logs for a specific backup or restore:
velero get backup
velero debug --backup <backup-name>
velero get restore
velero debug --restore <restore-name>
This creates a bundle-<date>.tar.gz file containing logs and metadata.
2. Capture a support bundle
Gather the clusterβs state after failure:
kubectl support-bundle --interactive=false secret/${NS}/kotsadm-swimlane-platform-supportbundle
kubectl support-bundle -n $NS --interactive=false https://raw.githubusercontent.com/replicatedhq/troubleshoot-specs/main/host/default.yaml
3. Look for out-of-memory (OOM) events
On each node:
dmesg -T | grep -i oom
dmesg -T | egrep -i 'killed process'
journalctl -k | grep -i 'killed process'
From any node, check specific pods:
kubectl describe pod swimlane-tools-0 -n $NS | grep -i oom
kubectl get pod swimlane-tools-0 -n $NS -o jsonpath="{.status.containerStatuses[*].lastState.terminated.reason}"
kubectl describe pod swimlane-sw-mongo-0 -n $NS | grep -i oom
kubectl get pod swimlane-sw-mongo-0 -n $NS -o jsonpath="{.status.containerStatuses[*].lastState.terminated.reason}"
4. Check node and container memory configuration
On each node:
kubectl describe node <node-name> | grep -i memory
journalctl -u kubelet | tail -100 > kubelet.log
The journalctl -u kubelet command helps uncover kubelet-level issues
From any node, check resource limits for mongo and tools containers:
kubectl get sts swimlane-sw-mongo -n $NS -o jsonpath="{.spec.template.spec.containers[*].resources}" && echo
kubectl get sts swimlane-tools -n $NS -o jsonpath="{.spec.template.spec.containers[*].resources}" && echo
5. Check for large MongoDB collections
Run inside Mongo shell:
Swimlane 10.x
kubectl exec -it -n $NS swimlane-sw-mongo-0 -- mongosh -u Admin -p --authenticationDatabase admin --tls --tlsAllowInvalidCertificates admin
NOTE: For older versions, change "mongosh" to "mongo"
kubectl exec -it -n $NS mongo-0 -- mongosh -u Admin -p --authenticationDatabase admin --tls --tlsAllowInvalidCertificates admin
Then identify top collections by data and index size using the provided JavaScript scripts below:
Large MongoDB collections
let collections = [];
db.getMongo().getDBNames().forEach(function(dbName) {
const currentDB = db.getSiblingDB(dbName);
currentDB.getCollectionInfos().forEach(function(collInfo) {
if (collInfo.type === "collection" && !collInfo.name.startsWith("system.")) {
const stats = currentDB.getCollection(collInfo.name).stats();
collections.push({
db: dbName,
collection: collInfo.name,
sizeBytes: stats.size,
sizeMB: (stats.size / (1024 * 1024)).toFixed(2),
sizeGB: (stats.size / (1024 * 1024 * 1024)).toFixed(2)
});
}
});
});
collections.sort((a, b) => b.sizeBytes - a.sizeBytes);
print("Top 10 Largest Collections by Data Size:");
collections.slice(0, 10).forEach(function(item, index) {
print(`${index + 1}. ${item.db}.${item.collection} - ${item.sizeBytes} bytes | ${item.sizeMB} MB | ${item.sizeGB} GB`);
});
let indexStats = [];
db.getMongo().getDBNames().forEach(function(dbName) {
const currentDB = db.getSiblingDB(dbName);
currentDB.getCollectionInfos().forEach(function(collInfo) {
if (collInfo.type === "collection" && !collInfo.name.startsWith("system.")) {
const stats = currentDB.getCollection(collInfo.name).stats();
indexStats.push({
db: dbName,
collection: collInfo.name,
indexBytes: stats.totalIndexSize,
indexMB: (stats.totalIndexSize / (1024 * 1024)).toFixed(2),
indexGB: (stats.totalIndexSize / (1024 * 1024 * 1024)).toFixed(2)
});
}
});
});
indexStats.sort((a, b) => b.indexBytes - a.indexBytes);
print("Top 10 Collections by Index Size:");
indexStats.slice(0, 10).forEach(function(item, index) {
print(`${index + 1}. ${item.db}.${item.collection} - ${item.indexBytes} bytes | ${item.indexMB} MB | ${item.indexGB} GB`);
});
6. Check disk space and dump folder content
Inside the swimlane-tools pod:
kubectl exec -n $NS $(kubectl get pod -n $NS -l app=swimlane-tools -o name) -- ls -lh /dump
On each node, check disk usage:
7. Additional log collection
To improve verbosity during troubleshooting use the --log-level argument. The available values (in increasing verbosity) are:
- error β Logs only critical errors
- warn β Logs warnings and errors
- info β Default; logs general operational messages
- debug β Logs detailed debug information
- trace β Logs everything, including low-level internal operations (very verbose)
You can set this in the Velero deployment like this:
kubectl edit deployment/velero -n velero
# add argument: --log-level debug under spec.containers.args
spec:
containers:
- name: velero
args:
- server
- --log-level=debug
This will provide deeper insight.
8. Additional recommendations
- Capture screenshots of error messages in the UI
- Note the failure timestamp and backup or restore names
- For restore, ensure the backup being used is complete