H.10 Platform
A large repository is served by SCM hosting, review, search, indexes, ownership, workspace, mirrors, snapshots, and other capabilities. Treating that portfolio as one invisible “repo” makes reliability, capacity, support, and cost impossible to reason about.
Define The Service Before Its Objective
H.10.1 Repository Services inventories clients, capabilities, authorities, data, dependencies, owners, and normal, degraded, emergency, maintenance, and retirement modes. H.10.2 Service Reliability can then set availability, integrity, durability, recovery, and data-loss objectives for named workflows. A restore exercise, not the existence of a backup, is the evidence that source metadata, review state, ownership, or indexes can recover.1
H.10.3 Service Capacity segments cold and warm workflows, teams, languages, regions, change classes, partial workspaces, peak demand, backlog, growth, and failure state. Repository size alone cannot reveal whether transfer, indexing, query, review, concurrency, storage, or a downstream service is saturated.
Operate Demand And Economics
H.10.4 Platform Support routes incidents, product defects, documentation gaps, onboarding, feature intake, migration help, and domain questions differently. Repeated questions, workarounds, and exceptions feed product or policy repair; enabling work transfers capability instead of making a central queue permanent.
H.10.5 Economics compares build, buy, and hybrid services through engineering, infrastructure, storage, network, security, compliance, support, integration, incidents, migration, switching, and exit. The same option may move cost from developer wait to platform toil or from control to vendor dependency, so comparisons need workflow and cohort evidence.
H.10.6 Repository Metrics closes the loop with time to first change, feedback and review wait, recovery, repeated support, toil, and developer experience by cohort. These outcomes test whether service and policy changes improve repository work without collapsing people or workflows into one health score. Telemetry purpose, access, retention, aggregation, and privacy are part of that measurement contract.
Operate repository capabilities as named services. Reliability protects their state and workflows, capacity explains demand and bottlenecks, support closes the product loop, economics exposes where cost and risk move, and repository metrics test whether those decisions improve real workflows.
Footnotes
-
The Site Reliability Workbook — Ch. 8: On-Call — explicit service responsibility, incident response, and recovery readiness ↩
Sections in this chapter · 6
Map and operate the SCM, review, search, indexing, ownership, workspace, and related services that support repository workflows.
Protect repository-service availability and integrity through objectives, backup, restore, disaster recovery, and degraded modes.
Plan repository-service capacity with cohort-specific cold, warm, peak, growth, and failure-state evidence.
Operate developer support, intake, triage, escalation, documentation, capability transfer, and product feedback loops.
Compare build, buy, and hybrid repository platforms through total cost, control, switching, support, and recovery obligations.
Evaluate repository effectiveness through workflow outcomes, cohort experience, toil, recovery, and privacy-aware telemetry.