Logging and Debugging¶
Overview¶
Emit useful logs at the right levels (including structured JSON for aggregators), read tracebacks quickly, and use pdb when a failing path is unclear.
print is fine in Module 2 labs. Production automation needs logging: levels, correlation fields, no secrets, and stderr by default so stdout stays for data.
Complete OOP first. Diagrams use Excalidraw only.
This is a core tutorial in Module 10 · Logging & Debugging of the REBASH Academy Python for Cloud & DevOps Engineers series — written for Cloud, DevOps, Platform, and SRE engineers.
Prerequisites¶
Required¶
Learning Objectives¶
By the end of this tutorial, you will be able to:
- Configure
loggingwith a basic stderr handler - Choose DEBUG / INFO / WARNING / ERROR / CRITICAL
- Emit structured key=value or JSON log lines
- Read a traceback to the failing line
- Break into
pdbon a stubborn bug - Avoid logging secrets
Architecture¶
This topic’s control points and relationships are shown below.
Theory¶
logging basics¶
import logging
import sys
logging.basicConfig(
level=logging.INFO,
format="%(asctime)s %(levelname)s %(name)s %(message)s",
stream=sys.stderr,
)
log = logging.getLogger("inventory")
log.info("start env=%s", "lab")
Use %s lazy formatting (not f-strings in hot DEBUG loops) so expensive messages are skipped when filtered.
Levels¶
| Level | Ops meaning |
|---|---|
| DEBUG | Detail for authors |
| INFO | Normal milestones |
| WARNING | Degraded but continuing |
| ERROR | Failed operation |
| CRITICAL | About to exit / page |
Structured logging¶
Prefer stable fields for Loki/CloudWatch/ELK:
log.info("check_done name=%s ok=%s duration_ms=%d", name, ok, ms)
# or json.dumps({"event": "check_done", "name": name, "ok": ok})
Tracebacks¶
Unread tracebacks waste incidents. Read bottom frame first (your file:line), then causes. logging.exception("...") inside except attaches the traceback.
pdb¶
Commands: l list, n next, s step, c continue, p expr, q quit. Remove breakpoints before committing.
Debugging techniques¶
- Reproduce with a fixture
- Add one INFO/DEBUG at the boundary
- Binary-search the failing branch
pdbonly when state is unclear- Fix + add a regression test (Module 22)
Hands-on Lab¶
Focus: practise the core workflow for Logging and Debugging
mkdir -p ~/rebash-python/module-10
cd ~/rebash-python/module-10
python3 -m venv .venv
source .venv/bin/activate
Step 1 – Configure a module logger¶
cd ~/rebash-python/module-10
source .venv/bin/activate
cat > app_log.py << 'EOF'
#!/usr/bin/env python3
from __future__ import annotations
import logging
import sys
def setup(level: str = "INFO") -> logging.Logger:
logging.basicConfig(
level=getattr(logging, level.upper(), logging.INFO),
format="%(levelname)s %(name)s %(message)s",
stream=sys.stderr,
)
return logging.getLogger("healthcheck")
def main() -> int:
log = setup("INFO")
log.info("event=start")
try:
int("nope")
except ValueError:
log.exception("event=parse_failed field=replicas")
return 1
return 0
if __name__ == "__main__":
raise SystemExit(main())
EOF
python app_log.py || true
Step 2 – Structured fields¶
python - <<'PY'
import logging, sys
logging.basicConfig(level=logging.INFO, stream=sys.stderr, format="%(message)s")
log = logging.getLogger("inv")
for host, ok in [("web-a", True), ("web-b", False)]:
log.info("event=host_check host=%s ok=%s", host, str(ok).lower())
PY
Step 3 – Traceback reading¶
Note the bottom frame pointing at 1 / 0.
Step 4 – Optional pdb (interactive)¶
# Run only if you can interact with the terminal:
# python - <<'PY'
# x = 1
# breakpoint()
# print(x)
# PY
echo "Use breakpoint() locally when you need a REPL on a frozen value"
Validation¶
- Lab commands run under
~/rebash-python/module-10/ - You can explain each Theory section in your own words
- You used modern tooling where it applies to this topic
- You can describe one production failure mode for this topic
Code Walkthrough¶
Production practice for Logging and Debugging always combines:
- Inspect before you change (status, plan, logs, dry-run)
- Prefer reversible, documented changes (Git, IaC, drop-ins, version pins)
- Capture evidence (command output, pipeline logs) for handovers
- Prefer current tools and APIs over legacy shortcuts
- Least privilege — escalate credentials only when required
Keep runbooks short enough to follow under pressure. Automate checks; keep humans for judgement.
Security Considerations¶
- Treat credentials and tokens for python as privileged — never commit them
- Prefer short-lived auth (OIDC, roles, SSO) over long-lived keys
- Validate blast radius before apply/deploy/delete operations
- Restrict who can approve production changes
- Collect audit logs; limit who can read sensitive traces
Common Mistakes¶
Skipping fundamentals for Logging and Debugging
Validate assumptions against the Theory section and official docs before changing production.
Treating lab defaults as production-ready for Logging and Debugging
Lab shortcuts (open security groups, admin roles, skip approvals) must not ship unchanged.
Changing production without a rollback path
Always know how to revert (previous artefact, prior release, state rollback, DNS failback).
Best Practices¶
- Encode Logging and Debugging changes as code and review them in pull requests
- Pin versions (images, modules, actions, provider plugins)
- Separate environments with clear promotion gates
- Alert on symptoms with runbooks attached
- Destroy lab resources; tag everything with owner and expiry where possible
Troubleshooting¶
| Symptom | Likely cause | What to do |
|---|---|---|
| No log output | Level too high / wrong logger | Set root/basicConfig; check name |
| Logs on stdout | Handler stream | Use stream=sys.stderr |
| Secret leak | Logged env/headers | Redact; never log Authorization |
Summary¶
loggingto stderr with levels and stable fieldsexception()on unexpected failures- Tracebacks: read the bottom frame first
pdbfor hard state; tests for regressions
Interview Questions¶
- How does Logging and Debugging show up when operating Cloud or production platforms?
- What would you check first if this area misbehaves in production?
- Which modern tools or APIs replace older equivalents here?
- What security control should accompany this capability?
- How would you automate verification of this topic in CI?
Sample answer — question 2
Start with blast radius and recent changes, gather evidence (logs, status, plan/diff), then fix forward with a known rollback path — not guesswork.