Troubleshooting and Upgrades¶
Overview¶
Diagnose production Jenkins: failed builds, agent issues, Pipeline replay, plugin problems, performance symptoms, LTS upgrade guides, safe restart, and rollback.
Upgrades are operational procedures — not surprise Friday clicks.
This is a core tutorial in Module 16 · Troubleshooting and Upgrades of the REBASH Academy Jenkins for Cloud & DevOps Engineers series — written for Cloud, DevOps, Platform, and SRE engineers.
Prerequisites¶
- Completed prior modules in this track where linked in frontmatter
- Git and Docker for lab workflows
- Running Jenkins LTS from Installing Jenkins LTS when a live controller is required
Learning Objectives¶
By the end of this tutorial, you will be able to:
- Triage a failed build from console logs and stage views
- Use Pipeline replay appropriately (and know its limits)
- Outline a plugin regression bisect approach
- Plan an LTS upgrade with backup and rollback
Architecture¶
This topic’s control points and relationships are shown below.
Theory¶
What it is¶
Console logs are the first artefact; Pipeline Steps and stage views narrow the failing step. Replay re-runs with edited script (admin capability — audit it). Agent issues show as “offline”, “launch failed”, or queue stuck. Plugin problems often appear after updates — binary search disable/update. LTS upgrades follow published upgrade guides; backup JENKINS_HOME, upgrade test, then production, with a restore rollback path.
Why it matters¶
Mean time to recovery depends on disciplined triage. Unplanned upgrades couple plugin binaries to controller core. SRE-owned runbooks beat heroics.
How it works¶
- Capture build URL, Git SHA, agent label, and failing step.
- Reproduce on a non-prod controller when possible.
- Check agent connectivity and executor availability.
- For regressions, identify last good plugin/core versions.
- Upgrade along LTS: read LTS upgrade guides, snapshot, change window, smoke tests.
System admin topics: Troubleshooting style handbook pages.
Key concepts and comparisons¶
| Symptom | First checks |
|---|---|
| Queue forever | Agents online? labels? executors? |
| Git checkout fail | Credentials, SSH host keys |
| Groovy CPS errors | Script security, library versions |
| Slow UI | Heap, plugins, build records retention |
| Upgrade step | Why |
|---|---|
| Backup | Rollback reality |
| Test controller | Plugin surprise |
| Smoke Pipelines | Real validation |
| Safe restart | Clean load |
Common pitfalls¶
- Replaying production with ad-hoc scripts and never committing the fix.
- Upgrading core and 50 plugins simultaneously with no notes.
- Deleting
JENKINS_HOMEbackups after “successful” upgrade day-of. - Ignoring agent clock skew and disk-full errors.
Hands-on Lab¶
Create a workspace for this tutorial.
Focus: triage a broken Jenkinsfile and draft an LTS upgrade runbook
Step 1 – Primary exercise¶
cat > Jenkinsfile.broken << 'EOF'
pipeline {
agent any
stages {
stage('Oops') {
steps {
sh 'exit 1'
}
}
}
}
EOF
cat > Jenkinsfile.fixed << 'EOF'
pipeline {
agent any
stages {
stage('Ok') {
steps {
sh 'echo healthy; exit 0'
}
}
}
}
EOF
# Local triage simulation
bash -c 'sh -c "exit 1"' || echo "captured-failure-exit-$?"
diff -u Jenkinsfile.broken Jenkinsfile.fixed || true
cat > upgrade-runbook.md << 'EOF'
# LTS upgrade runbook
1. Read upgrade guide for target LTS
2. Snapshot volume / backup JENKINS_HOME
3. Upgrade TEST controller; run smoke Pipelines
4. Change window: upgrade PROD; safe restart
5. Smoke: login, agent, sample Multibranch, deploy dry-run
6. Rollback: restore snapshot if smoke fails
EOF
grep -E 'Snapshot|Rollback|TEST' upgrade-runbook.md
Step 2 – Support bundle note¶
cat > support-notes.md << 'EOF'
# When asking for help
- Jenkins version (LTS)
- Plugin list excerpt
- Sanitised console log
- Whether still reproducible after replay
# Prefer support/admin monitors over pasting secrets
EOF
grep version support-notes.md
Final cleanup¶
# Keep ~/rebash-jenkins/ for later tutorials; stop Compose only if you are done with the controller
# docker compose -f ~/rebash-jenkins/module-02/docker-compose.yml down # optional; omit -v to keep JENKINS_HOME
Validation¶
- Lab commands run under
~/rebash-jenkins/module-16/ - You can explain each Theory section in your own words
- You used current Jenkins LTS / Pipeline practices where they apply
- You can describe one production failure mode for this topic
Code Walkthrough¶
Production practice for Troubleshooting and Upgrades always combines:
- Inspect before you change (status, plan, logs, dry-run)
- Prefer reversible, documented changes (Git, Jenkinsfile, JCasC)
- Capture evidence (console logs, plan artefacts) for handovers
- Prefer current LTS and supported plugins over legacy shortcuts
- Least privilege — escalate credentials only when required
Keep runbooks short enough to follow under pressure. Automate checks; keep humans for judgement.
Security Considerations¶
- Treat Jenkins credentials and cloud tokens as privileged — never commit them
- Keep builds off the built-in node; isolate untrusted pull requests
- Prefer short-lived auth (OIDC-style patterns, scoped RBAC) over long-lived keys
- Validate blast radius before apply/deploy/delete operations
- Collect audit logs; limit who can administer the controller
Common Mistakes¶
Fix only via Replay
Commit Jenkinsfile/library fixes; replay is for diagnosis, not configuration management.
Big-bang upgrades
Split core and plugin changes; use a test controller.
No rollback path
A backup you have never restored is a hope, not a plan.
Best Practices¶
- Encode Troubleshooting and Upgrades changes as code and review them in pull requests
- Prefer Jenkins LTS and pinned agent/tool versions
- Keep builds off the controller; use labelled agents
- Least privilege for credentials and cluster/cloud access
- Destroy or stop lab resources; keep
~/rebash-jenkins/notes for the track
Troubleshooting¶
| Symptom | Likely cause | Fix |
|---|---|---|
| Job stuck in queue | No matching agent/label or executors busy | Check nodes, labels, and executor counts |
| Checkout / SCM failure | Credentials, URL, or permissions | Verify credential ID and repository access |
| Pipeline CPS / script error | Syntax, sandbox, or library mismatch | Read error line; validate Jenkinsfile; pin library version |
| Plugin / UI broken after update | Incompatible plugin set | Restore backup; disable suspect plugin on test controller |
| Disk full on agent/controller | Workspaces or old builds | Clean workspaces; trim build retention |
Summary¶
Troubleshooting and Upgrades is essential for Cloud and DevOps engineers operating Jenkins. Practise the lab until the inspection and change path is muscle memory, then continue the track.
Interview Questions¶
- How do you triage a red Pipeline build?
- When is Pipeline replay appropriate?
- How do you isolate a bad plugin update?
- What is your LTS upgrade checklist?
- Agent offline — what do you check first?
Sample answer — question 2
Use replay to test a hypothesis quickly, then commit the real fix to SCM; control who has replay permission.
Sample answer — question 4
Read the upgrade guide, backup, upgrade test, smoke, production change window, smoke again, keep rollback snapshots.