Production Shell Scripting¶
Overview¶
Production scripts survive retries, concurrent cron, hostile input, and tired operators. Build those properties in deliberately.
This is Tutorial 17 in Module 17: Production Shell Scripting of the REBASH Academy Shell Scripting for DevOps Engineers series — written for Linux administrators, DevOps engineers, SREs, and platform engineers who automate production hosts with Bash.
Prerequisites¶
- Error Handling, Logging, and Debugging
- Bash 4.2+ on Linux (WSL2/VM/cloud)
Learning Objectives¶
By the end of this tutorial, you will be able to:
- Apply the core ideas of “Production Shell Scripting” in a real ops script
- Use
set -euo pipefailas the production default - Use quoted expansions and clear stderr diagnostics
- Produce meaningful exit codes for automation consumers
- Debug behaviour with
bash -xwhen something fails - Relate this topic to day-to-day Linux admin and DevOps work
Architecture¶
Ops scripts sit between humans/automation and system tools. This topic’s control points are shown below.
Theory¶
What it is¶
Production shell scripting is the set of practices that make Bash safe to run unattended at scale: static analysis with ShellCheck, idempotent behaviour under re-runs, secure handling of paths and secrets, structured logging with bounded retries, lock files against overlap, and clear configuration via environment or flags. It is less about clever one-liners and more about boring, reviewable contracts that Continuous Integration (CI) and operators can rely on.
Why it matters¶
A script that works once on a laptop can still destroy data on the second cron tick or leak credentials in process lists. Production standards catch those classes of failure before merge: ShellCheck finds quoting bugs; locks stop overlapping backups; retries distinguish flaky networks from hard errors; secrets stay out of git and world-readable files. Teams that skip this layer accumulate fragile glue that only the original author dares to touch.
How it works¶
Run shellcheck script.sh in CI and treat quoting and cd warnings as defects. Disable rules only with a justified inline directive. Design for idempotency: create users only if missing, use mkdir -p, enable services when needed — declare desired state rather than blind mutation.
Secure defaults: no eval, no unquoted rm -rf $var, no secrets on argv, path allow-lists before destructive operations, and least privilege. Log structured messages; retry transient network errors with sleep or backoff, a hard attempt cap, and escalation after exhaustion. Guard scheduled work with locks:
Accept configuration from environment files, flags, or drop-in directories. Keep secrets out of the repository; integrate with Ansible, systemd EnvironmentFile=, or a secrets store. Prefer small scripts that call well-tested tools over re-implementing complex logic in Bash.
Key concepts¶
| Practice | Production role |
|---|---|
| ShellCheck in CI | Catch quoting and common foot-guns early |
| Idempotency | Safe re-runs from cron and retries |
| Security | No eval; allow-listed paths; least privilege |
| Retries | Bounded backoff for transient failures |
flock | One active run for critical jobs |
| External config | Env/flags; secrets outside git |
Common pitfalls¶
- Merging scripts that fail ShellCheck with silenced warnings and no rationale
- Non-idempotent steps that fail loudly on every scheduled re-run
- Putting tokens in command-line arguments visible via
ps - Infinite retry loops that mask a permanent outage
- Skipping locks so two backups write to the same destination
Hands-on Lab¶
Create a workspace for this tutorial.
Focus: shellcheck clean script; flock; retry helper; idempotent mkdir/user check
Step 1 – Production patterns¶
cat > prod.sh << 'EOF'
#!/usr/bin/env bash
set -euo pipefail
lock=$PWD/prod.lock
exec 9>"$lock"
flock -n 9 || { echo "locked" >&2; exit 0; }
retry() {
local n=0 max=3
until "$@"; do
n=$((n + 1))
(( n >= max )) && return 1
sleep 1
done
}
mkdir -p state
[[ -d state ]] || exit 4
retry true
echo "idempotent ok"
EOF
chmod +x prod.sh
./prod.sh
command -v shellcheck >/dev/null && shellcheck prod.sh || echo 'shellcheck optional'
Final step – Cleanup note¶
Validation¶
- Lab commands run under
~/rebash-shell/lab17/ - You can explain each Theory heading in your own words
- Failure path exits non-zero and prints diagnostics to stderr (where applicable)
- You can relate this topic to a real DevOps or Linux admin task
Code Walkthrough¶
Production Bash for Production Shell Scripting always combines:
- A clear shebang (
#!/usr/bin/env bash) - Strict mode near the top (
set -euo pipefail) from Module 2 onward - Quoted expansions and explicit tests
- Functions with
localfor reusable behaviour - Documented exit codes and stderr logging
Keep scripts short enough to review in a single merge request. When logic grows (complex JSON APIs, heavy state), hand off to Python and keep Bash as the launcher.
Security Considerations¶
- Treat all external input (args, files, env) as untrusted until validated
- Never log secrets; prefer masked CI variables and secret stores
- Prefer least privilege — do not require root for file-local tasks
- Avoid
evaland unquoted expansions in destructive commands - Validate paths stay under an allow-listed root before
rmor overwrite
Common Mistakes¶
Skipping strict mode
Cron and CI hide failures that an interactive terminal would show. Fix: start with set -euo pipefail from Module 2 onward.
Unquoted path expansions
Spaces and globs rewrite your command line. Fix: always "$path" / "$@".
Assuming interactive PATH
Aliases and fancy PATH entries disappear under schedulers. Fix: set PATH or use absolute paths.
Best Practices¶
- One purpose per script; compose with functions or small binaries
- Log to stderr; reserve stdout for data or RESULT lines
- Idempotent behaviour where scheduling may overlap
- Pair every new script with a failing-path test you actually run
- Run ShellCheck in CI before merging automation
Troubleshooting¶
| Symptom | Likely cause | Fix |
|---|---|---|
| Works in terminal, fails in cron | PATH / cwd / env | Fingerprint env; set PATH |
unbound variable | set -u | Provide defaults or export vars |
| Pipeline “succeeds” incorrectly | Missing pipefail | set -o pipefail |
[[ unexpected operator | Running under sh/dash | Fix shebang to Bash |
Summary¶
Production Shell Scripting is a core skill for Linux admins and DevOps engineers automating real hosts and pipelines. Practise the lab until the failure path is as familiar as the happy path, then continue the track.
Interview Questions¶
- How does this topic show up in production Linux administration or CI?
- What failure mode appears if you ignore quoting or strict mode here?
- How would you test this behaviour under a minimal cron-like environment?
- When would you move this logic out of Bash into Python or another tool?
- What exit code contract would you document for teammates?
Sample answer — question 2
Unquoted expansions and missing pipefail create silent or partial failures — especially under cron — that look healthy in monitoring until data is wrong.
Related Tutorials¶
- Shell Scripting for DevOps Engineers – Category Overview
- Error Handling, Logging, and Debugging (previous)
- Troubleshooting Shell Scripts (next)
- Learning Paths