H.7.8 Main Recovery

Reading progress2 missing
Article in preparation — showing conspect notes

Main recovery restores shared source after an integration failure; repository-service and product-delivery failures have different owners.

  • Choose a build-cop, platform on-call, component owner, or incident commander model from incident frequency and scale.
  • Classify product, flaky-evidence, infrastructure, capacity, policy, and unknown failure before retrying.
  • Define authority to pause affected admission, request revert or fast fix, route ownership, and communicate scope.
  • Preserve bounded escalation and handoff; coordination must not become permanent domain maintenance.
  • Measure detection, classification, ownership, recovery, repeated incidents, and blocked developer time.