Skip to content

Repository files navigation

cnpe-exam-prep

OpenSSF Scorecard OpenSSF Baseline

An unofficial lab for the CNCF Certified Cloud Native Platform Engineer (CNPE) exam.

The exam is hands-on and covers five domains, so reading docs is not enough. This repo builds a working internal developer platform on kind and then proves each piece actually functions. Not "the pod is Running", but "a repo pushed to Gitea produced an Argo CD Application that deployed itself, and reverting drift works".

Every layer installs on its own. That matters on a laptop, because running all of it at once will saturate the CPU.

Curriculum

The lab is the machinery; curriculum/ is the study plan that drives it. It covers every competency in the official curriculum PDF across 29 evening-sized sections. Each section pairs concepts with exercises against this lab, and each exercise ends with a command that proves the thing worked.

It is published at cnpe.rbstp.dev. It is a self-contained site rather than a pile of markdown, so it also runs off disk: open curriculum/index.html in a browser (no server, no build step, file:// is fine) or run make study.

Code blocks have copy buttons, / jumps to any section by name, tool or concept, and exercises tick off as you verify them. Thirteen interactive figures can be driven directly, and drill mode replays all 257 self-check questions as flashcards weighted toward what you miss. CNPE Quest turns the whole map into a role-playing game: five regions for the five domains, a town per section whose people teach the theory and the commands, a trial built from the section's own self-check cards, and a dungeon holding one of the lab's faults, fought with real kubectl, argocd, flux, tkn and crossplane commands against a simulated cluster. Progress lives in that browser's local storage, and Export and Import on the dashboard move it between browsers or machines. On the hosted site you can also sign in with GitHub from the header to keep a copy across machines: opt-in, off by default, identity-only (the OAuth scope is empty), and signed out the console never touches the network. See docs/progress-sync.md.

The study pages use a reading-first black-and-gold theme, dark by default with a full light theme (t). Read / Practice / Recall and the exercise index jump directly through the original continuous lesson; Contents provides the same access on mobile. Use f for focus view and Aa for larger reading text. Text size and theme stay local to the browser, outside progress sync. The quest uses the same page shell, navigation, charcoal-and-gold windows and dark/light switch, with an original 16-bit RPG presentation: pixel-art towns, staged battles, scene transitions, optional synthesized sound, click-to-travel and a quest journal (q). Desktop stats sit beside the game, while mobile uses a compact strip. It respects reduced motion and keeps quick help collapsed below the game.

The dashboard maps every official competency to a section and to the make validate check or exercise that demonstrates it. Two 15-task mock exams (paper 1, paper 2) share no tasks, score separately, and come with grading commands and a 120-minute clock. A weak-spots panel splits your drill accuracy by domain and names the one to spend evenings on.

make site stages the same directory exactly as it is published to GitHub Pages, including a single-file console at cnpe.rbstp.dev/console.html with the fonts inlined: one URL to hand over, or to save and read offline. See docs/deploy-pages.md for the deploy, the DNS and the settings it needs.

1. Platform architecture and infrastructure (15%) networking, right-sizing, storage, multi-tenancy, cost
2. GitOps and continuous delivery (25%) GitOps fundamentals, Argo CD, Flux, Tekton, progressive delivery, troubleshooting
3. Platform APIs and self-service (25%) the platforms white paper, CRDs, operators, Argo Workflows, Crossplane, kro + Backstage
4. Observability and operations (20%) Prometheus, alerting, dashboards + logs, tracing, DORA metrics, incident response
5. Security and policy enforcement (15%) RBAC + secrets, policy engines, PSS, audit + SBOM, mTLS + SPIRE, pipeline security

What it builds

Layer Tools Exam domain
up kind (3 nodes, 2 zones), Cilium, Gateway API, metrics-server, VPA, local registry, API audit logging Platform architecture (15%)
gitea Gitea on the kind network, resolvable in-cluster as gitea.lab through CoreDNS GitOps (25%)
gitops Argo CD, Argo Rollouts, Argo Workflows, Flux GitOps (25%)
cicd Tekton Pipelines/Triggers/Dashboard, Trivy Operator GitOps, Security
api Crossplane v2, provider-kubernetes, provider-helm, CloudNativePG, kro Platform APIs (25%)
obs Prometheus, Grafana, OpenTelemetry, Jaeger, Loki + Alloy, OpenCost Observability (20%)
sec Kyverno, OPA Gatekeeper, Sealed Secrets, External Secrets, Pod Security Standards, quotas Security (15%)
spire SPIFFE/SPIRE, workload identity Security (15%)
mesh Second cluster with Istio ambient and Flagger Security (15%)
portal Backstage on the host, with a software template that publishes to Gitea Platform APIs (25%)

Versions as tested: kind 0.33.0, Helm 4.2.2, Cilium 1.20.1, Argo CD 10.9.0 (chart), Crossplane 2.4.0 (chart), kube-prometheus-stack 88.5.3, Istio 1.30.3, Flux 2.9.4. K8S_IMAGE pins Kubernetes 1.37.0; the curriculum's September output was captured on 1.36.1, the pin in force before that bump. Only Kubernetes is pinned, by K8S_IMAGE in lab.env; scripts/01-tools.sh installs kind and the other CLIs from their latest release, so those numbers record what was tested rather than what you will get.

Only the Kubernetes node image is pinned, by digest, in lab.env. Helm charts and the Tekton manifests float on purpose, so a fresh install gets whatever is current and the versions above will drift. That is the right trade for exam prep, because chart values and API versions moving under you is the thing the exam actually tests. When something breaks, kubectl api-resources | grep <tool> and kubectl explain <kind> are the fix. If you want reproducibility instead, pinning --version in helmi (scripts/lib.sh) is a one-line change.

Tools

Everything the lab installs, with a link to each project. The CNCF exam tool list is broad, so I aimed for coverage of it rather than picking favorites.

Cluster and networking: Kubernetes · kind · Cilium · Hubble · Gateway API · cloud-provider-kind · metrics-server · Vertical Pod Autoscaler

Packaging and templating: Helm · Kustomize

Git and GitOps: Gitea · Argo CD · Argo Rollouts · Argo Workflows · Flux · Flagger

CI and supply chain: Tekton · Trivy · Trivy Operator · Cosign · skopeo

Platform APIs and self-service: Crossplane · kro · CloudNativePG · Kubebuilder · Backstage

Observability and cost: Prometheus · Alertmanager · Grafana · OpenTelemetry · Jaeger · Loki · Alloy · OpenCost · kubectl-cost

Security, policy and identity: Kyverno · OPA Gatekeeper · Sealed Secrets · External Secrets Operator · SPIFFE/SPIRE · Istio · Linkerd (alternative to Istio, MESH=linkerd make mesh)

Shell tooling installed by make tools: k9s · stern · kubectx · mise (pins Node 24 for Backstage)

Hardware and build time

I built and tested this on a 2019 Intel MacBook Pro, Core i9-9880H (8 cores, 16 threads), 32 GB RAM, running Omarchy 4 (Arch). It is not fast hardware for this. The CPU is the bottleneck, never the memory.

Step Time Notes
make host ~1 min Needs sudo
make tools ~5 min Mostly downloads
make up ~4 min Cluster ready in under 3
make gitea ~1 min
make gitops ~5 min
make cicd ~6 min Trivy scans every workload on install
make api ~6 min Crossplane packages pull slowly
make obs ~10 min The heaviest layer
make sec ~5 min
make spire ~3 min
make mesh ~6 min Builds a second cluster
make portal ~20 min Backstage pulls a large npm tree

Around 70 minutes for everything, or about 25 minutes for a useful subset (make core cicd api).

At rest the full stack uses roughly 21 GB of RAM across two clusters, 85 pods and 20 Helm releases, plus about 19 GB of Docker images and volumes. Load average sits near 18 during the observability install and settles to about 3 afterwards.

If you are on the same chassis, install mbpfan so the fans ramp before the CPU throttles:

yay -S mbpfan-git && sudo systemctl enable --now mbpfan
watch -n2 'grep MHz /proc/cpuinfo | head'   # ~800 MHz means you are throttling

Quick start

git clone https://github.com/rbstp/cnpe-exam-prep.git && cd cnpe-exam-prep
cp lab.env.example lab.env      # set GITEA_PASS
make host                       # sysctls, docker, kernel limits. needs sudo
make tools                      # every CLI into ~/.local/bin
make core                       # cluster + git server + Argo CD/Flux/Rollouts/Workflows
make cicd api obs sec spire     # the rest, one at a time
make validate                   # 71 functional checks
make urls                       # where everything is

Run the layers one at a time and watch make status in between. Running two at once on this hardware makes both slower.

What needs root, and what it touches

More than just make host, so it is worth knowing before you run any of it:

  • make host adds you to the docker group (root-equivalent), writes /etc/sysctl.d/99-cnpe-lab.conf and a systemd drop-in for docker, and creates /etc/docker/daemon.json only if absent. It restarts dockerd when the limits drop-in changes, which bounces every container on the machine, and it warns before doing so.
  • make gitea appends one line to /etc/hosts mapping gitea.lab, and only when the entry is missing or the IP moved.
  • make up starts cloud-provider-kind under sudo -b so LoadBalancer Services get real IPs. It needs root to bind ports 80 and 443, it records its PID, and it skips itself with a warning if sudo would prompt.
  • make down stops that process by the recorded PID.
  • make portal runs sudo npm i -g yarn only if yarn is missing.

make tools installs every CLI into ~/.local/bin and needs sudo only for the pacman packages. Three upstream installers are piped to a shell unpinned (crossplane, istioctl, linkerd), which is how those projects document installation, but read them first if that bothers you.

Targets

make host            Kernel limits, docker, thermal advice (run once, needs sudo)
make tools           Install every CLI into ~/.local/bin
make up              Create the cluster: kind + Cilium + LB + metrics + VPA + registry
make gitea           Local git server, seeded repos, CoreDNS entry
make gitops          Argo CD, Argo Rollouts, Argo Workflows, Flux
make cicd            Tekton Pipelines/Triggers/Dashboard, Trivy Operator
make api             Crossplane, CloudNativePG, kro, kubebuilder hints
make obs             Prometheus, Grafana, OTel, Jaeger, Loki+Alloy, OpenCost
make sec             Kyverno, Gatekeeper, sealed/external secrets, PSS
make spire           SPIFFE/SPIRE workload identity
make mesh            Second cluster + Istio ambient (MESH=linkerd works) + Flagger
make portal          Scaffold Backstage on the host
make core            Minimum useful lab (~4 GB)
make full            Everything on the main cluster (~14 GB)
make validate        Functionally verify every layer (FAST=1 to skip probes)
make grade           Run a mock exam's grading block from its page (EXAM=1|2)
make study           Open the CNPE study console (curriculum) in a browser
make site            Stage the study console exactly as Pages publishes it, into ./_site
make typecheck       Type-check the console's JS via JSDoc (tsc --noEmit, nothing compiled)
make fix-cp-metrics  Expose control-plane metrics on an existing cluster
make urls            Every UI, its URL/port-forward, and credentials
make forward         Start a background port-forward for every UI
make forward-stop    Kill all port-forwards started by 'make forward'
make status          Clusters, endpoints, unhealthy pods, host load
make break           Inject a random fault, then diagnose it under time pressure
make break-answer    Reveal the last injected fault
make break-fix       Auto-diagnose and repair whatever 'make break' injected
make down-gitops     Remove one layer from the running cluster ('make gitops' puts it back)
make down-cicd       ... and the same for cicd, api, obs, sec, spire and mesh
make down            Delete both clusters (keeps git history + registry)
make nuke            Delete everything including Gitea data

Each section page of the study console ends with the make down-<layer> commands for whatever the next section does not need, so a laptop only ever runs the layers the section in front of you uses. A layer's CRDs stay behind; nothing else does.

make validate

The part I care about most. It checks behavior, not pod status.

── Domain 3: Platform APIs & self-service
  ✓ XRD established                                True
  ✓ example XR reconciled (Ready)                  True
  ✓ XR actually created its namespace              team-c

── NetworkPolicy enforcement (the thing kindnet fakes)
  ✓ egress blocked by NetworkPolicy                curl exit=28 (denied, correct)

── Portal: Backstage golden path
  ✓ golden-path template installed
  ✓ Applications auto-generated from git           1

──────────────────────────────────────────────
  PASS 71   FAIL 0   SKIP 0
──────────────────────────────────────────────

Lab is fully functional.

make urls

A trimmed sample. The real output also prints the port-forward command for each service, the generated passwords, and the second cluster.

SERVICE           URL                      ACCESS
───────────────── ──────────────────────── ──────
Argo CD           http://172.18.0.10:80    admin / generated, printed locally
Argo Rollouts     http://172.18.0.12:3100  no auth
Argo Workflows    http://172.18.0.16:2746  no auth (server authMode)
Tekton Dashboard  http://172.18.0.11:9097  no auth
Grafana           http://172.18.0.9:80     admin / admin (3000 is Backstage)
Prometheus        http://172.18.0.13:9090  no auth
Jaeger            http://172.18.0.14:16686 no auth
OpenCost          http://172.18.0.15:9090  UI on /
Alertmanager      http://localhost:9093    needs a port-forward
Hubble UI         http://localhost:12000   needs a port-forward

Backstage portal (runs on the HOST, not in the cluster)
  Portal           http://localhost:3000     start: cd <clone>/portal && yarn start
                   backend API on :7007      (needs Node 24; mise.toml pins it)
  Golden path      Create -> "Golden path service" -> publishes to gitea org 'services'
                   Argo CD then generates an Application automatically

Always-on (no port-forward needed)
  Gitea            http://gitea.lab:3000   (also http://localhost:3001)
  OCI registry     localhost:5001   (push: docker push localhost:5001/demo:v1)

Grafana really is admin / admin, set by scripts/50-observability.sh. The things with no UI are worth knowing about:

API audit log    docker exec -it cnpe-control-plane tail -f /var/log/kubernetes/audit.log | jq .
Rollouts TUI     kubectl argo rollouts dashboard
Hubble CLI       cilium hubble port-forward &  then: hubble observe
Cost CLI         kubectl cost namespace --opencost --show-all-resources
Compliance       kubectl get clustercompliancereports,vulnerabilityreports,sbomreports -A
SPIRE identities kubectl -n spire exec sts/spire-server -c spire-server -- \
                   /opt/spire/bin/spire-server entry show
Golden path apps kubectl -n argocd get applicationset,applications

Real LoadBalancer IPs need cloud-provider-kind, which wants root to bind ports 80 and 443. Skip it and use make forward instead if you would rather not run a root process.

make break

Incident response is a third of the observability domain and the hardest thing to practice alone. make break injects one fault at random from a scenario library, prints the kind of ticket a platform team actually receives, and starts you on a clock. Target is 7 minutes for the workload group and 10 for the platform ones, which is roughly exam pace.

The workload group is the classic broken-pod drill in team-a: image, probe, resources, rbac, quota, netpol and config. Each fails differently, and some are invisible in kubectl get pods. The rbac one only shows up under kubectl auth can-i --as=..., and the netpol one strips the DNS egress rule off the tenant policy, leaving a Running pod that cannot resolve anything. That last one is worth understanding: NetworkPolicies are additive allow-lists, so you cannot break DNS by adding a restrictive policy. You have to remove the rule that allowed it.

The other four groups break the platform tooling itself, which is where the exam puts most of its weight:

  • gitops points an Argo CD Application at a git ref that does not exist, suspends a Flux Kustomization and then deletes the workload it manages, or aims a Rollouts canary analysis at a Prometheus address that resolves to nothing.
  • cicd removes the Task a Tekton Pipeline references, or revokes the RBAC that keeps an EventListener alive.
  • apis revokes a Crossplane provider's ClusterRoleBinding under a fresh XR, or pauses an XR so spec changes stop propagating.
  • security ships a Kyverno Deny policy nobody asked for, or flips the namespace to PSS restricted while stripping the pod's securityContext.

Where a healthy starting state is what makes the symptom realistic, the scenario is built healthy first, then broken. The GitOps scenarios deploy into their own drill-gitops namespace so they never fight the ApplicationSet exercise over the same objects.

make break                     # random fault from the whole library
DOMAIN=gitops make break       # random fault from one domain
FAULT=flux-suspend make break  # drill one specific fault
make break-answer              # reveal what was injected, and why it broke
make break-fix                 # auto-diagnose, repair, and explain the evidence

Domains are workload, gitops, cicd, apis and security. Injecting a new fault first heals and removes whatever the previous drill left behind, so you can chain drills without cleaning up in between.

make break-fix detects faults from cluster state rather than reading the answer file, and prints the evidence it matched on. Use it to reset, or to check your own diagnosis after you have had a go.

Credentials

Nothing sensitive is committed. lab.env is gitignored, so copy the template first:

cp lab.env.example lab.env

GITEA_PASS is the only value you need to set. It is the admin password for a throwaway Gitea container that only listens on localhost, so it is not a real secret, but it does not belong in git either.

Everything else is generated at runtime and gitignored:

  • .gitea-token, an API token created by make gitea, mode 600
  • .gitea-info, a connection summary
  • The Argo CD admin password, random per install, read it with make urls
  • portal/, the scaffolded Backstage app, which contains its own generated config

Things worth knowing if you build something similar

These cost me real time, and none of them are obvious from the upstream docs.

Use Cilium, not kindnet. kindnet does not enforce NetworkPolicy, so every network-policy exercise silently passes and you learn nothing. make validate proves enforcement with an egress probe that must fail.

Turn on API server audit logging. Generating audit trails is an explicit exam competency and almost nobody practices it. One gotcha: the apiserver is started with --audit-policy-file=/etc/kubernetes/audit/policy.yaml, so the mounted directory must contain a file named exactly policy.yaml. A missing audit policy stops kube-apiserver from starting at all, which looks like a cluster that never boots.

kubeadm binds control-plane metrics to localhost. kube-controller-manager, kube-scheduler and etcd all listen on 127.0.0.1, so Prometheus cannot scrape them and Grafana's control-plane dashboards stay empty. Fix it in the kind config with bind-address: 0.0.0.0 and listen-metrics-urls, plus a KubeProxyConfiguration patch for kube-proxy. make fix-cp-metrics retrofits a running cluster.

A namespaced Crossplane v2 XR cannot compose cluster-scoped resources. provider-kubernetes serves Object as cluster-scoped in kubernetes.crossplane.io and namespaced in kubernetes.m.crossplane.io. Mixing them gives you cannot apply cluster scoped composed resource. The namespaced variant also needs a ClusterProviderConfig rather than a ProviderConfig.

Argo CD's Gitea SCM generator needs an organization, not a user. It lists repos through the org API, so repos owned by a user account return error listing repos: not found. It also wants a token with write:repository, write:organization and read:issue, and each missing scope only shows up as a runtime ApplicationSet error. Set cloneProtocol: https too, or {{ .url }} resolves to an SSH URL and Argo CD fails with SSH agent requested but SSH_AUTH_SOCK not-specified.

Run the service mesh on a second cluster. Istio ambient and Cilium both rewrite the dataplane. Debugging that interaction teaches you nothing about the exam, and a broken mesh should not cost you your GitOps state.

Gateway API v1.5 and later blocks older CRDs. It ships a ValidatingAdmissionPolicy that rejects any Gateway API CRD before v1.5.0. cloud-provider-kind embeds an older bundle and installs it at startup, so it gets denied and its service controller dies, and no LoadBalancer ever gets an IP. Start it with --gateway-channel=disabled.

Backstage wants Node 22 or 24, not any even LTS. The app create-app scaffolds declares "node": "22 || 24", so 20 is even and still refused. Arch ships 26, and create-app fails on it, so mise.toml pins 24 for this directory. create-app also has no --name flag and always prompts, so it hangs when scripted, and the .yarnrc.yml it generates sets npmMinimalAgeGate: 3d, which can reject one of Backstage's own fresh dependencies.

Trivy's node-collector needs a toleration. It is pinned to each node by nodeSelector but does not tolerate the control-plane taint, so that node silently produces no compliance report. nodeCollector.tolerations and trivyOperator.scanJobTolerations are separate chart keys and you need both.

The local registry needs two names and a CoreDNS entry. Pods pull localhost:5001/... through a containerd mirror, but anything that runs inside a pod and talks to the registry itself (kaniko pushing, Kyverno fetching signatures) needs kind-registry:5000, and pods cannot resolve that name until CoreDNS is taught it. The cluster script wires both names into containerd and the registry IP into CoreDNS. Kyverno also needs features.registryClient.allowInsecure=true, or every image verification dies on server gave HTTP response to HTTPS client.

cosign v3 signs in a format Kyverno does not read yet. Its default is the new bundle format (a sha256-<digest> tag); Kyverno's ImageValidatingPolicy looks for the legacy sha256-<digest>.sig. Sign with --use-signing-config=false --new-bundle-format=false --tlog-upload=false and verification works. "no signatures found" on an image you definitely signed is a format-skew symptom, not a key problem.

Not included

No cloud provider. Everything runs locally, which means no real cloud Crossplane providers and no managed services. That is deliberate, because the exam gives you a cluster and not an AWS account.

License

MIT. See LICENSE.

Reference

About

Unofficial lab for the CNCF Cloud Native Platform Engineer (CNPE) exam: a full internal developer platform on kind: Argo CD, Flux, Crossplane, Tekton, Backstage, Cilium, Kyverno, OpenTelemetry. Layered make targets, tuned for Arch/Omarchy.

Topics

Resources

Security policy

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages