Production Helm Practices¶
Overview¶
Apply SemVer to charts, structure an enterprise chart repo, and complete a production readiness checklist.
Production charts are boring: small templates, documented values, pinned deps, OCI publish, CI lint/template, GitOps deploy. Treat chart version bumps as release engineering.
This is a core tutorial in Module 11 · Production Helm of the REBASH Academy Helm for Kubernetes Engineers series — written for Cloud, DevOps, Platform, and SRE engineers.
Prerequisites¶
- Modules 7–10
Learning Objectives¶
By the end of this tutorial, you will be able to:
- Version charts with SemVer
- Separate library vs app charts
- Document values (README / values schema)
- Complete an excellence checklist
Architecture¶
This topic’s control points and relationships are shown below.
Theory¶
What it is¶
Production Helm practices are the engineering habits that make charts operable at scale: Semantic Versioning (SemVer) for chart packages, clear ownership of library vs application charts, documented values (README and optionally values.schema.json), reproducible publishes to OCI, CI gates (lint + template), and GitOps deploy with tested rollback. Production charts are deliberately boring — small templates, explicit defaults, pinned dependencies.
| Practice | Outcome |
|---|---|
SemVer chart version | Consumers know when upgrades are breaking |
| Documented values | Teams configure without reading every template |
| OCI publish | Immutable artefacts beside images |
| CI lint/template | Broken charts never merge |
| Rollback drill | Confidence when upgrades fail |
Why it matters¶
A clever template that only one author understands becomes an outage multiplier. Enterprises need charts that pass review, publish cleanly, and upgrade safely across dozens of services. Treating chart releases as release engineering — changelog, version bump, artefact publish — aligns Helm with how you already ship container images.
How it works¶
A practical production loop:
- Change templates/values in a PR; CI runs
helm lintandhelm templatefor each env values file. - Bump
versioninChart.yamlusing SemVer (MAJOR for breaking values/template contracts, MINOR for compatible features, PATCH for fixes). - Update
appVersionwhen the bundled application release changes (informational clarity). helm packageand push to an OCI registry with an immutable reference.- GitOps Applications consume the new chart version / digest and env values.
- Verify rollback path in staging before promoting.
Structure repos so library charts (shared helpers) version independently from application charts. Keep values schemas honest — required keys fail fast in CI rather than at apply time.
Key concepts and comparisons¶
| Chart type | Contains | Consumers |
|---|---|---|
| Application | Workload manifests | GitOps apps / product teams |
| Library | Helpers only | Other charts via dependencies |
| Umbrella | Mostly dependencies | Platform compositions (use sparingly) |
Excellence means you can answer: What changed in 1.4.0? Which values are required? How do we roll back?
Common pitfalls¶
- Bumping only
appVersionwhen templates changed — consumers track chartversion. - One mega-chart for every microservice — prefer composeable units.
- Undocumented required values discovered only in production.
- Publishing mutable
latestchart tags in OCI for production promotion.
Hands-on Lab¶
Create a workspace for this tutorial.
Focus: Pin versions, use values files, and verify before upgrade
Step 1 – Prepare production-style install inputs¶
kubectl create namespace rebash-helm
helm create prod-demo
# Pin chart version in Chart.yaml (already 0.1.0) and freeze image tag via values
cat > prod-values.yaml <<'EOF'
replicaCount: 2
image:
tag: "1.27-alpine"
resources:
requests:
cpu: 50m
memory: 64Mi
EOF
helm lint prod-demo
helm template demo ./prod-demo -n rebash-helm -f prod-values.yaml > /tmp/prod-render.yaml
grep -E 'replicas:|image:|cpu:' /tmp/prod-render.yaml | head -n 20
Step 2 – Upgrade with atomic flag and check history¶
helm upgrade --install demo ./prod-demo -n rebash-helm -f prod-values.yaml --atomic --timeout 2m
helm -n rebash-helm history demo
kubectl -n rebash-helm get deploy -o wide
Final step – Cleanup note¶
helm uninstall demo -n rebash-helm --ignore-not-found || true
kubectl delete namespace rebash-helm --ignore-not-found
# Workspace kept for notes; remove with: rm -rf "$(pwd)" when finished
Validation¶
- Lab commands run under
~/rebash-helm/module-11/ - You can explain each Theory section in your own words
- You used modern tooling where it applies to this topic
- You can describe one production failure mode for this topic
Code Walkthrough¶
Production practice for Production Helm Practices always combines:
- Inspect before you change (status, plan, logs, dry-run)
- Prefer reversible, documented changes (Git, IaC, drop-ins, version pins)
- Capture evidence (command output, pipeline logs) for handovers
- Prefer current tools and APIs over legacy shortcuts
- Least privilege — escalate credentials only when required
Keep runbooks short enough to follow under pressure. Automate checks; keep humans for judgement.
Security Considerations¶
- Treat credentials and tokens for helm as privileged — never commit them
- Prefer short-lived auth (OIDC, roles, SSO) over long-lived keys
- Validate blast radius before apply/deploy/delete operations
- Restrict who can approve production changes
- Collect audit logs; limit who can read sensitive traces
Common Mistakes¶
Bumping only appVersion when templates changed — consumers track chart version.
Validate assumptions against the Theory section and official docs before changing production.
One mega-chart for every microservice — prefer composeable units.
Lab shortcuts (open security groups, admin roles, skip approvals) must not ship unchanged.
Changing production without a rollback path
Always know how to revert (previous artefact, prior release, state rollback, DNS failback).
Best Practices¶
- Encode Production Helm Practices changes as code and review them in pull requests
- Pin versions (images, modules, actions, provider plugins)
- Separate environments with clear promotion gates
- Alert on symptoms with runbooks attached
- Destroy lab resources; tag everything with owner and expiry where possible
Troubleshooting¶
| Symptom | Likely cause | Fix |
|---|---|---|
| Auth / permission denied | Wrong identity, policy, or scope | Check caller identity, roles, and least-privilege policies |
| Timeout / no route | Network, DNS, security group, or endpoint | Trace path, DNS, and allow-lists before retrying |
| Drift / unexpected plan | Manual change or wrong state/workspace | Reconcile desired vs actual; avoid click-ops on managed resources |
| Pipeline/job red | Flaky step, cache, or missing secret | Read failing step logs; bisect recent workflow/config changes |
| Cost spike | Idle load balancer, NAT, oversized compute | Inventory billable resources; stop/delete labs promptly |
Summary¶
Production Helm Practices is essential for Cloud and DevOps engineers working with helm. Practise the lab until the inspection and change path is muscle memory, then continue the track.
Interview Questions¶
- Which Helm flags help safer production upgrades?
- How should chart and image versions be pinned?
- What belongs in a release checklist before upgrading prod?
- How do you limit blast radius of a bad chart upgrade?
- Why keep values structured per environment rather than one mega file?
Sample answer — question 2
--atomic, timeouts, and staged environments help upgrades fail cleanly. Pin chart version and image tags, review helm diff/template output, and ensure PDBs exist for the app.
Sample answer — question 4
Use canary namespaces, smaller replica changes, and fast rollback. Separate prod pipelines with approvals so a bad values edit cannot silently ship.