Architecture · Version 1.0.0 · Reviewed 2026-08-02
Kubernetes Architecture Review Specialist
Make a defensible decision about kubernetes architecture boundary review and kubernetes failure-mode modeling with evidence, explicit trade-offs, and a verification plan.
4 method steps
4 documented failure modes
4 diagnostic checks
7 quality gates
Reviews architecture boundaries, operating assumptions, and failure behavior in Kubernetes using workload manifests, Services, policies, events, and cluster topology and pod states, endpoint membership, scheduler events, and resource telemetry, with explicit attention to readiness, requests, or policy disagreeing with runtime behavior and hiding the true failure layer.
₹299 one-time
Get this skill archive
What it checks first
Kubernetes Architecture Review Specialist reviews architecture boundaries, operating assumptions, and failure behavior in Kubernetes using workload manifests, Services, policies, events, and cluster topology and pod states, endpoint membership, scheduler events, and resource telemetry, with explicit attention to readiness, requests, or policy disagreeing with runtime behavior and hiding the true failure layer. Use it when the work involves Kubernetes architecture boundary review, Kubernetes failure-mode modeling, Kubernetes architecture decision record.
- The quality attribute that actually constrains the design: latency, consistency, availability, cost, or compliance.
- The critical path and the number of network hops on it.
- Where state lives and who owns it, since ownership ambiguity becomes a correctness problem.
- The failure behavior of every dependency: fail open, fail closed, or degrade.
Example task
Input
Apply the architecture review specialist to our Kubernetes system before the next production change. We can provide workload manifests, Services, policies, events, and cluster topology; the main concern is readiness, requests, or policy disagreeing with runtime behavior and hiding the true failure layer.
Expected output
Map workload lifecycle, scheduling, service discovery, policy, and nodes before choosing components. The first design risk to test is readiness, requests, or policy disagreeing with runtime behavior and hiding the true failure layer. Compare only options that preserve the stated invariant, then record load assumptions, rollback, ownership, and the signal that would reverse the decision.