SkillVaultskills Browse all 1,000+ skills

Reliability · Version 1.2.0 · Reviewed 2026-08-02

BigQuery Observability Design Specialist

Reduce production risk in BigQuery service-level signal design and BigQuery diagnostic telemetry mapping with evidence, explicit trade-offs, and a verification plan.

4 method steps 6 documented failure modes 5 diagnostic checks 7 quality gates

Designs low-noise signals that expose user impact and causal mechanisms in BigQuery using table partitioning, clustering, SQL, reservations, and scheduled jobs and bytes processed, stage timelines, slot use, shuffle, and spill, with explicit attention to unpruned scans or high-cardinality shuffle turning a small result into large cost.

₹199 one-time

Get this skill archive

Install in your AI coding tool

SkillVault packages this skill in the open Agent Skills format for five leading coding tools.

What this skill helps you do

  • BigQuery service-level signal design
  • BigQuery diagnostic telemetry mapping
  • BigQuery actionable alert definition

How BigQuery Observability Design Specialist works

You provide

Schema, query plans, and the real access pattern

It inspects

Plan accuracy and lock behavior for BigQuery service-level signal design

It decides

A BigQuery diagnostic telemetry mapping change weighed against write cost

You verify

Re-measured plan with buffer reads and timing compared

What it checks first

BigQuery Observability Design Specialist designs low-noise signals that expose user impact and causal mechanisms in BigQuery using table partitioning, clustering, SQL, reservations, and scheduled jobs and bytes processed, stage timelines, slot use, shuffle, and spill, with explicit attention to unpruned scans or high-cardinality shuffle turning a small result into large cost. Use it when the work involves BigQuery service-level signal design, BigQuery diagnostic telemetry mapping, BigQuery actionable alert definition.

  1. The actual query plan with real row counts, not the estimated plan or the query text alone.
  2. Whether the workload is read-heavy, write-heavy, or mixed, since the correct design differs sharply.
  3. Transaction boundaries and duration, because long transactions block vacuum and hold locks.
  4. Index coverage relative to both the filter and the sort, since satisfying one but not the other still costs a sort.
  5. Connection pool behavior, as pool exhaustion presents as database slowness while the database is idle.

Failure modes it recognizes

  • An index that serves the predicate but not the ordering, forcing a full sort for a small LIMIT.
  • A long-running transaction preventing vacuum and causing gradual bloat and plan degradation.
  • Implicit type casting on a join or filter column silently disabling index use.
  • Connection pool exhaustion from long-held connections, appearing as a database problem.
  • A write-heavy table with excessive indexes where insert cost dominates the workload.
  • Statistics stale after a bulk load, so the planner chooses a plan for a table size that no longer exists.

Answers it will reject

  • Adding an index per slow query until write amplification becomes the new bottleneck.
  • Tuning configuration parameters before examining the plan for the dominant query.
  • Interpreting `EXPLAIN` without `ANALYZE`, which reports estimates and proves nothing.
  • Increasing pool size to fix latency caused by lock contention, which adds waiters rather than capacity.

Decision rules it applies

  • Optimize the query that dominates total time, not the one that feels slowest in isolation.
  • Order composite index columns by equality first, then range or sort last.
  • Keep transactions short and never hold one open across an external call.
  • Create and drop indexes concurrently on live tables, accepting the longer build for the absent lock.

Evidence it asks for

  • `EXPLAIN (ANALYZE, BUFFERS)` to compare estimated with actual rows and attribute I/O.
  • Rank queries by cumulative execution time rather than by single-execution latency.
  • Monitor the oldest open transaction and lock wait counts as standing metrics.

The method inside

  1. Extract decisions, facts, and unresolved questions needed for BigQuery service-level signal design.
  2. Organize BigQuery diagnostic telemetry mapping around the reader's next decision or action rather than the source order.
  3. Draft BigQuery actionable alert definition with source traceability and no invented behavior.
  4. Run a completeness, consistency, audience, and actionability review before returning the artifact.

Deliverables

  • BigQuery service-level signal design assessment
  • BigQuery diagnostic telemetry mapping decision and action plan
  • BigQuery actionable alert definition verification checklist

Evidence requirements

  • User-visible symptoms and SLO impact
  • Timeline, telemetry, deploys, and dependency state
  • Current mitigations and operational constraints

Quality gates

  • Every material claim traces to supplied evidence or is labeled as a hypothesis.
  • The response follows the declared deliverable contract.
  • No execution, access, measurement, or verification is invented.
  • Secrets and personal data are redacted rather than repeated.
  • The user receives a concrete independent verification step.
  • The relevant failure modes in this domain were considered rather than only the reported symptom.
  • No listed anti-pattern was recommended as a solution.

Example task

Input

Apply the observability design specialist to our BigQuery system before the next production change. We can provide table partitioning, clustering, SQL, reservations, and scheduled jobs; the main concern is unpruned scans or high-cardinality shuffle turning a small result into large cost.

Expected output

Instrument bytes processed, stage timelines, slot use, shuffle, and spill at the same boundary as the user-visible objective. The dashboard must make unpruned scans or high-cardinality shuffle turning a small result into large cost distinguishable from ordinary load. Page only on symptoms that require action, retain causal dimensions within a bounded cardinality budget, and test every alert with a controlled failure.

Boundaries and compatibility

Ideal for

  • BigQuery service-level signal design: produce a decision or artifact grounded in supplied evidence.
  • BigQuery diagnostic telemetry mapping: produce a decision or artifact grounded in supplied evidence.
  • BigQuery actionable alert definition: produce a decision or artifact grounded in supplied evidence.

Out of scope

  • Replacing incident command authority
  • Calling a trigger the root cause without a causal chain

Agent compatibility

  • GitHub Copilot Agent Skills
  • Cursor Agent Skills
  • Claude Code Skills
  • OpenAI Codex Skills
  • JetBrains Junie Skills

Tool policy: Advisory by default. No tools are assumed. If the host provides tools, use read-only evidence gathering unless the user explicitly approves a scoped write or execution action.