SkillVaultskills Browse all 1,000+ skills

Reliability · Version 1.3.0 · Reviewed 2026-08-02

Capacity Headroom Planner

Reduce production risk in growth modeling and failure-domain sizing with evidence, explicit trade-offs, and a verification plan.

4 method steps 4 documented failure modes 4 diagnostic checks 7 quality gates

Turns growth, seasonality, and failure-domain requirements into a defensible capacity plan with explicit saturation limits.

₹299 one-time

Get this skill archive

Install in your AI coding tool

SkillVault packages this skill in the open Agent Skills format for five leading coding tools, with a raw SKILL.md fallback for every other compatible IDE or agent.

See the complete graphical installation and usage guide

What this skill helps you do

  • Growth modeling
  • Failure-domain sizing
  • Saturation limit analysis

How Capacity Headroom Planner works

You provide

Growth forecast, failure-domain count, and current utilization

It inspects

The genuinely binding resource, not the obvious one

It decides

Headroom derived from surviving a domain loss

You verify

Simulated domain failure keeps survivors below saturation

What it checks first

Capacity Headroom Planner turns growth, seasonality, and failure-domain requirements into a defensible capacity plan with explicit saturation limits. Use it when the work involves Growth modeling, Failure-domain sizing, Saturation limit analysis.

  1. User-visible impact and error-budget consumption rather than component health.
  2. Saturation signals — queue depth, pool utilization, connection counts — near the onset.
  3. Whether the system recovered on its own, which indicates saturation rather than corruption.
  4. The blast radius and what boundary should have contained it.

Failure modes it recognizes

  • Retry amplification turning a partial failure into a total outage.
  • A shared dependency creating correlated failure across supposedly independent services.
  • Slow resource exhaustion invisible until a hard limit is crossed.
  • A rollback blocked by an incompatible migration.

Answers it will reject

  • Treating the trigger as the root cause, which stops the analysis before the fragility is identified.
  • Adding a runbook step where a boundary would remove the failure mode.
  • Measuring availability as a mean, which hides regional and tenant-level outages.

Decision rules it applies

  • Stabilize user impact before completing diagnosis.
  • Bound every retry with a budget, jitter, and a circuit breaker.
  • Prefer removing a failure mode over detecting it faster.

Evidence it asks for

  • Record time-to-detect, time-to-mitigate, and time-to-resolve separately.
  • Quantify impact in customer terms: failed requests, affected accounts, duration.
  • Verify recovery with the same signal that detected the failure.

The method inside

  1. Extract decisions, facts, and unresolved questions needed for growth modeling.
  2. Organize failure-domain sizing around the reader's next decision or action rather than the source order.
  3. Draft saturation limit analysis with source traceability and no invented behavior.
  4. Run a completeness, consistency, audience, and actionability review before returning the artifact.

Deliverables

  • Growth modeling assessment
  • Failure-domain sizing decision and action plan
  • Saturation limit analysis verification checklist

Evidence requirements

  • User-visible symptoms and SLO impact
  • Timeline, telemetry, deploys, and dependency state
  • Current mitigations and operational constraints

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

How much headroom should we keep for our API tier and how do we justify it to finance?

Expected output

Headroom is not a preference, it is arithmetic from your failure domain. If you run three zones and must survive losing one, steady-state utilization cannot exceed roughly 66 percent before a zone loss saturates the survivors. Establish the true bottleneck resource first, since sizing on CPU when you are connection-bound produces the wrong number...

Boundaries and compatibility

Ideal for

  • Growth modeling: produce a decision or artifact grounded in supplied evidence.
  • Failure-domain sizing: produce a decision or artifact grounded in supplied evidence.
  • Saturation limit analysis: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Replacing incident command authority
  • Calling a trigger the root cause without a causal chain

Agent compatibility

  • GitHub Copilot Agent Skills
  • Cursor Agent Skills
  • Claude Code Skills
  • OpenAI Codex Skills
  • JetBrains Junie Skills

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.