Artifacts, Caches, and Dependencies¶
Overview¶
Configure artifacts: and cache: so a build job produces a shareable package, a test job consumes it, and dependency installs reuse a keyed cache without treating cache as a correctness guarantee.
Artefacts are job outputs GitLab stores and can pass to later jobs or attach to merge requests (packages, JUnit reports, coverage). Cache speeds installs by restoring files between pipelines; it is best-effort and must not replace pinned lockfiles. Reports (junit, coverage_report, and similar) surface quality signals in the MR UI.
This is a core tutorial in Module 7 · Artifacts & Cache of the REBASH Academy GitLab CI/CD for Cloud & DevOps Engineers series — written for Cloud, DevOps, Platform, and SRE engineers.
Prerequisites¶
Learning Objectives¶
By the end of this tutorial, you will be able to:
- Declare
artifacts:pathswithexpire_in - Share artefacts across jobs with
dependencies/needs - Configure a dependency cache with a lockfile-based
key - Distinguish artefacts from cache and from container images
- Attach a simple report artefact for the MR widget
Architecture¶
This topic’s control points and relationships are shown below.
Theory¶
What it is¶
GitLab CI separates what you ship between jobs from what you hope to reuse for speed:
| Mechanism | Purpose | Guaranteed? |
|---|---|---|
artifacts: | Persist outputs for later jobs / download / MR | Yes (within retention) |
cache: | Speed dependency or build dirs | No — may miss or be empty |
| Reports | Structured UI feedback (tests, coverage) | Via artifacts:reports |
Artefacts can expire (expire_in), be limited to certain jobs (dependencies), or ride DAG edges (needs). Cache keys often hash package-lock.json, poetry.lock, or go.sum so installs invalidate when dependencies change.
Why it matters¶
Slow pipelines waste runner minutes and delay review. Wrong cache design causes flaky “works on my branch” builds when a stale cache hides a missing lockfile. Artefacts make builds auditable: the same tarball or binary that passed tests can be promoted. Reports turn CI into a review surface instead of a log scrape.
How it works¶
- A build job writes files under a path (for example
dist/) and lists them inartifacts:pathswithexpire_in. - A test job declares
needs: [build](ordependencies) so GitLab downloads those artefacts beforescript. - An optional cache restores
node_modules/or.cache/pipusing a key derived from the lockfile; the job still runs install if the cache is cold. - Test jobs may set
artifacts:reports:junitso failures appear on the merge request. - Expired artefacts disappear; cache may be cleared by runners or policy — never assume either is permanent storage.
Prefer short retention for intermediate binaries; keep release artefacts on tags or push immutable images to a registry.
Key concepts and comparisons¶
| Concern | Artefacts | Cache |
|---|---|---|
| Correctness | Part of the pipeline contract | Optimisation only |
| Cross-job | Explicit download | Same key / same project scope |
| Security | Can contain secrets — scope carefully | Same risk if you cache credentials |
| UI | Downloadable; reports in MR | Opaque speed-up |
Common pitfalls¶
- Treating cache as a substitute for committing lockfiles.
- Omitting
expire_inand filling storage with huge trees. - Caching the entire workspace including secrets or
.env. - Expecting artefacts without listing paths or after a failed job (
when: on_successdefault). - Confusing job artefacts with Container Registry images.
Hands-on Lab¶
Create a workspace for this tutorial.
Focus: configure cache keys and job artifacts with expire_in
Step 1 – Build pipeline with cache + artifacts¶
mkdir -p src && echo 'print("build")' > src/main.py
cat > .gitlab-ci.yml << 'EOF'
stages: [deps, build, test]
variables:
PIP_CACHE_DIR: "$CI_PROJECT_DIR/.cache/pip"
deps:
stage: deps
image: python:3.12-alpine
cache:
key: {files: [requirements.txt]}
paths: [.cache/pip]
script:
- echo pytest > requirements.txt
- pip install -r requirements.txt
artifacts:
paths: [requirements.txt]
expire_in: 1 day
build:
stage: build
image: python:3.12-alpine
needs: [deps]
script:
- mkdir -p dist && cp -r src dist/ && echo "$CI_COMMIT_SHA" > dist/REVISION
artifacts:
paths: [dist/]
expire_in: 3 days
test:
stage: test
image: python:3.12-alpine
needs: [build]
script: ["test -f dist/REVISION", "cat dist/REVISION"]
EOF
Step 2 – Verify cache and artifact paths¶
Final step – Cleanup note¶
Validation¶
- Lab commands run under
~/rebash-gitlab/module-07/ - You can explain each Theory section in your own words
- You used modern tooling where it applies to this topic
- You can describe one production failure mode for this topic
Code Walkthrough¶
Production practice for Artifacts, Caches, and Dependencies always combines:
- Inspect before you change (status, plan, logs, dry-run)
- Prefer reversible, documented changes (Git, IaC, drop-ins, version pins)
- Capture evidence (command output, pipeline logs) for handovers
- Prefer current tools and APIs over legacy shortcuts
- Least privilege — escalate credentials only when required
Keep runbooks short enough to follow under pressure. Automate checks; keep humans for judgement.
Security Considerations¶
- Treat credentials and tokens for gitlab as privileged — never commit them
- Prefer short-lived auth (OIDC, roles, SSO) over long-lived keys
- Validate blast radius before apply/deploy/delete operations
- Restrict who can approve production changes
- Collect audit logs; limit who can read sensitive traces
Common Mistakes¶
Treating cache as a substitute for committing lockfiles.
Validate assumptions against the Theory section and official docs before changing production.
Omitting expire_in and filling storage with huge trees.
Lab shortcuts (open security groups, admin roles, skip approvals) must not ship unchanged.
Changing production without a rollback path
Always know how to revert (previous artefact, prior release, state rollback, DNS failback).
Best Practices¶
- Encode Artifacts, Caches, and Dependencies changes as code and review them in pull requests
- Pin versions (images, modules, actions, provider plugins)
- Separate environments with clear promotion gates
- Alert on symptoms with runbooks attached
- Destroy lab resources; tag everything with owner and expiry where possible
Troubleshooting¶
| Symptom | Likely cause | Fix |
|---|---|---|
| Auth / permission denied | Wrong identity, policy, or scope | Check caller identity, roles, and least-privilege policies |
| Timeout / no route | Network, DNS, security group, or endpoint | Trace path, DNS, and allow-lists before retrying |
| Drift / unexpected plan | Manual change or wrong state/workspace | Reconcile desired vs actual; avoid click-ops on managed resources |
| Pipeline/job red | Flaky step, cache, or missing secret | Read failing step logs; bisect recent workflow/config changes |
| Cost spike | Idle load balancer, NAT, oversized compute | Inventory billable resources; stop/delete labs promptly |
Summary¶
Artifacts, Caches, and Dependencies is essential for Cloud and DevOps engineers working with gitlab. Practise the lab until the inspection and change path is muscle memory, then continue the track.
Interview Questions¶
- What is the difference between cache and artifacts in GitLab CI?
- A later job cannot find a file produced earlier — how do you diagnose it?
- When should cache keys include a lockfile hash?
- What security issue arises from caching world-writable directories?
- How does expire_in help control storage cost?
Sample answer — question 2
Confirm the producer declared artifacts paths, the consumer needs/dependencies includes that job, and paths match. Cache is best-effort and must not be the only way to pass build outputs.
Sample answer — question 4
Treat caches as untrusted acceleration: pin keys, avoid storing secrets, and keep artifact retention short.