SkillVaultskills Browse all 1,000+ skills

Infrastructure · Version 1.6.0 · Reviewed 2026-08-02

Linux Performance Investigator

Review and harden CPU saturation analysis and I/O latency diagnosis with evidence, explicit trade-offs, and a verification plan.

4 method steps 4 documented failure modes 4 diagnostic checks 7 quality gates

Builds evidence-driven Linux performance investigations using CPU, memory, disk, network, scheduler, pressure, and eBPF signals.

₹199 one-time

Get this skill archive

Install in your AI coding tool

SkillVault packages this skill in the open Agent Skills format for five leading coding tools, with a raw SKILL.md fallback for every other compatible IDE or agent.

See the complete graphical installation and usage guide

What this skill helps you do

  • CPU saturation analysis
  • I/O latency diagnosis
  • Memory-pressure investigation

How Linux Performance Investigator works

You provide

Manifests, plans, and current runtime topology

It inspects

Reversibility and blast radius for CPU saturation analysis

It decides

A I/O latency diagnosis change staged by risk

You verify

Platform-native health check after each stage

What it checks first

Linux Performance Investigator builds evidence-driven Linux performance investigations using CPU, memory, disk, network, scheduler, pressure, and eBPF signals. Use it when the work involves CPU saturation analysis, I/O latency diagnosis, Memory-pressure investigation.

  1. Whether a change is reversible, and specifically whether it replaces or mutates a stateful resource.
  2. Blast radius: the number of environments, regions, and workloads a change touches at once.
  3. Identity and permission scope of the executing principal.
  4. Drift between declared and actual state.

Failure modes it recognizes

  • An immutable-attribute change forcing replacement of a stateful resource.
  • A change applied to all environments simultaneously with no canary.
  • Over-broad permissions granted to make a deployment succeed and never narrowed.
  • Manual changes creating drift that the next apply silently reverts.

Answers it will reject

  • Approving a plan from summary counts rather than reading the replacement lines.
  • Suppressing drift detection to silence noise, which disables reconciliation.
  • Granting administrative rights as a debugging shortcut.

Decision rules it applies

  • Any stateful replacement requires a tested backup and restore path before approval.
  • Roll out by blast radius: one non-critical target, then one zone, then the fleet.
  • Grant the narrowest permission that completes the task, with an expiry.

Evidence it asks for

  • Diff the plan in machine-readable form and classify every action.
  • Verify the rollback path by executing it in a non-production environment.
  • Confirm post-change health with a platform-native check, not an assumption.

The method inside

  1. Define the measured baseline and user-visible target for CPU saturation analysis.
  2. Attribute the dominant cost or latency mechanism affecting I/O latency diagnosis.
  3. Rank memory-pressure investigation changes by expected impact, confidence, effort, and regression risk.
  4. Validate under representative load and retain guardrail metrics that detect a shifted bottleneck.

Deliverables

  • CPU saturation analysis assessment
  • I/O latency diagnosis decision and action plan
  • Memory-pressure investigation verification checklist

Evidence requirements

  • Infrastructure code or configuration
  • Runtime topology and environment constraints
  • Plan, events, policies, and failure symptoms

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Load average is high but CPU utilization is low. What data should I collect before changing instance size?

Expected output

High load includes tasks waiting in uninterruptible I/O. Check pressure stall information, disk latency, blocked-task stacks, and network filesystem health before assuming compute shortage...

Boundaries and compatibility

Ideal for

  • CPU saturation analysis: produce a decision or artifact grounded in supplied evidence.
  • I/O latency diagnosis: produce a decision or artifact grounded in supplied evidence.
  • Memory-pressure investigation: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Applying infrastructure changes without approval
  • Assuming cloud access or live resource visibility

Agent compatibility

  • GitHub Copilot Agent Skills
  • Cursor Agent Skills
  • Claude Code Skills
  • OpenAI Codex Skills
  • JetBrains Junie Skills

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.