Control Plane, Nodes and etcd
The API server is the front door and validation boundary. etcd durably stores API objects. The scheduler assigns unscheduled Pods to nodes; controller managers run reconciliation loops. On each node, kubelet makes the PodSpec real, the container runtime runs containers, and kube-proxy or an eBPF dataplane implements Service traffic. A component map matters because each failure leaves a different signature.
Analogy: Kubernetes is a thermostat, not a remote control. You declare the temperature; independent controllers measure reality and keep acting. Debugging means finding which sensor, rule or actuator prevents convergence.
Read the object as evidence
kubectl get RESOURCE NAME -n NAMESPACE -o yaml
kubectl describe RESOURCE NAME -n NAMESPACE
kubectl get events -n NAMESPACE --sort-by=.metadata.creationTimestamp
kubectl explain RESOURCE.spec
Do not memorize these as a ritual. The first command exposes desired and observed state, the second connects conditions and events, the third supplies a timeline, and the fourth checks the server's schema. Compare what the controller was asked to do with what it reports doing. Then test the narrowest hypothesis.
Scenario: Existing Pods may keep serving during an API-server outage, but new scheduling and most changes stop. Losing etcd is different: the cluster loses its authoritative state. Backups, quorum and tested restores are therefore control-plane concerns, not ordinary application backups.
Match symptoms to components
Use kubectl get --raw='/readyz?verbose' when API access works but control-plane health is suspect, and inspect kubectl get nodes plus system Pods before blaming an application. Scheduler trouble leaves Pods Pending with FailedScheduling events. Kubelet or runtime trouble leaves node conditions and container-start failures. A CNI failure often produces sandbox-creation events; an etcd or API failure affects writes and controller progress broadly. This symptom map keeps the investigation at the correct layer.
Warning: Restarting every control-plane component destroys timing evidence and can turn a partial impairment into an outage. Capture component health, events, logs and leader state first.
Scenario: Deployments stop progressing across every namespace while running traffic still succeeds. That cluster-wide scope points toward the control plane, not sixty unrelated application defects.
Production reasoning
Ask four questions: Who owns this object? What dependency must become ready next? Which controller reports the blocking condition? What evidence would disprove my current theory? This prevents symptom-driven changes. Record the context, namespace, object generation, image digest and recent rollout before mutation; a recreated Pod may erase the evidence you needed.
Warning: Running is not the same as ready, healthy, durable or correct. Kubernetes status is layered. Confirm the application-level outcome as well as the object state.
Goal: Put this model into practice in the Kubernetes lab namespace-basics. Open/labs/kubernetes, choosenamespace-basics, predict the failure path before changing anything, then use the simulator'scheckcommand to validate the finished state.
Deliberate practice
Before the lab, write the expected object relationship and the first three commands you will run. Afterward, explain why the fix converged and name one tempting change that would only mask the symptom. Repeat using an explicit namespace and a structured output format. This prediction-observation-explanation loop is what turns command familiarity into production judgment.
45-minute investigation
- Map (5 min): draw the owner-to-child chain and mark every namespace, selector, identity and dependency involved.
- Predict (5 min): write one expected status condition, one likely event and one log or metric signal before opening the lab.
- Observe (10 min): collect YAML, describe output and ordered events. Do not mutate state. Record which observation disproves your first theory.
- Repair (15 min): make the smallest declarative correction, watch the responsible controller converge, and verify the user-facing path rather than stopping at
Running. - Stress (10 min): change one relevant constraint - replica count, label, readiness, resource value or placement rule - predict the outcome, observe it, then restore the known-good declaration.
Tip: Keep a short incident note with symptom, evidence, hypothesis, change, result. Across four sections this produces a reusable runbook instead of a pile of remembered commands.