Architecture · Version 1.0.0 · Reviewed 2026-08-02
Databricks Architecture Review Specialist
Make a defensible decision about databricks architecture boundary review and databricks failure-mode modeling with evidence, explicit trade-offs, and a verification plan.
4 method steps
4 documented failure modes
4 diagnostic checks
7 quality gates
Reviews architecture boundaries, operating assumptions, and failure behavior in Databricks using notebooks, jobs, Spark plans, Delta tables, and cluster policy and Spark UI stages, skew, shuffle, spill, and cluster utilization, with explicit attention to partition skew or driver-side collection collapsing a distributed workload onto one process.
₹299 one-time
Get this skill archive
What it checks first
Databricks Architecture Review Specialist reviews architecture boundaries, operating assumptions, and failure behavior in Databricks using notebooks, jobs, Spark plans, Delta tables, and cluster policy and Spark UI stages, skew, shuffle, spill, and cluster utilization, with explicit attention to partition skew or driver-side collection collapsing a distributed workload onto one process. Use it when the work involves Databricks architecture boundary review, Databricks failure-mode modeling, Databricks architecture decision record.
- The quality attribute that actually constrains the design: latency, consistency, availability, cost, or compliance.
- The critical path and the number of network hops on it.
- Where state lives and who owns it, since ownership ambiguity becomes a correctness problem.
- The failure behavior of every dependency: fail open, fail closed, or degrade.
Example task
Input
Apply the architecture review specialist to our Databricks system before the next production change. We can provide notebooks, jobs, Spark plans, Delta tables, and cluster policy; the main concern is partition skew or driver-side collection collapsing a distributed workload onto one process.
Expected output
Map driver, executors, object storage, Delta transactions, and orchestration before choosing components. The first design risk to test is partition skew or driver-side collection collapsing a distributed workload onto one process. Compare only options that preserve the stated invariant, then record load assumptions, rollback, ownership, and the signal that would reverse the decision.