- All four failover selectors (three workers + the verdict job) and the paired env/cache expressions now exclude dependabot[bot]: under failover, dependency-supplied code keeps queueing for the hosted pool instead of executing on the persistent VM. A delayed Dependabot PR during an outage is an acceptable cost; dependency code on the privileged host is not. - Runbook (both languages): records the shipped failover bounds (coverage 8, snapshots 12, sized for six instances) and documents that the verdict job follows the selector too — operators previously had no explanation for a verdict queued after all workers passed. - Local static gate green: 32 passed, 0 failed (translation pairing 519 pairs consistent).
4.8 KiB
Agent Note: CI failover runbook — hosted pools → in-house pool
Status: implemented
English | 中文
Problem
The three required Linux worker jobs in CI (node 24 / static, node 24 / coverage, node 24 / snapshots and artifacts) and the required verdict job that aggregates them (all checks passed) run on the hosted enterprise 32-core pools. When those pools degrade — jobs queue indefinitely, the enterprise labels vanish, or GitHub-side capacity fails — every open pull request becomes unmergeable, and the ordinary recovery of merging a fix is itself deadlocked behind the very required checks that cannot run. An outage therefore needs a switch a repository admin can throw without merging anything.
Decision
Each of the three required Linux worker jobs — and the all checks passed verdict job, which would otherwise stay queued on the failed pool even after every worker passed — resolves its runner pool through the DSH_CI_FAILOVER repository variable. Unset (normal), they run on the hosted enterprise pools. Set to selfhosted by a repository admin, all four retarget onto the in-house self-hosted vm-backup pool, coverage and snapshot concurrency drop to shared-VM bounds, and the hosted-path pnpm cache restores are skipped. The switch is admin-only repository state, not a merge, so it works while every check is red. The in-house pool's readiness is continuously re-proven by the serial / linux (self-hosted standby) lane, which runs the complete unsharded aggregate on every master push.
What the in-house pool is
vm-backup: one 64-core VM, six always-on systemd-managed runner instances. Check the latest serial / linux (self-hosted standby) run before switching: a green standby is verified-yesterday capacity.
Switch (repo admin, ~1 minute, no merge)
- Repository Settings → Secrets and variables → Actions → Variables → New repository variable: name
DSH_CI_FAILOVER, valueselfhosted. - Retrigger the required jobs so they re-resolve their pool. Jobs already queued for the hosted labels do not retarget and cannot be re-run in place, so for the documented indefinite-queue outage, cancel the stuck run and re-run all jobs, or push a new commit; "Re-run failed jobs" only helps once a job has actually failed rather than queued.
- That is the entire switch. Under failover the workflow also, automatically: drops
DSH_COVERAGE_MAX_WORKERSto 8 andDSH_SNAPSHOT_MAX_CONCURRENCYto 12 (sized for six always-on instances: worst case 6 × 8 = 48 coverage workers on the 64-core VM) (shared-VM contention bounds), and skips the hosted-path pnpm cache restores (the VM's persistent store serves warm installs).
Capacity during failover
Six always-on instances absorb normal PR traffic (the pool's steady-state load is one serial standby job per master push, so failover capacity is effectively the full pool). If queues still build, register additional instances with an org registration token (org Settings → Actions → Runners → New runner) — cloning an existing runner directory and running config.sh takes about a minute per instance.
Switch back
Delete the DSH_CI_FAILOVER variable (or set it to anything other than selfhosted). New runs resolve back to the hosted enterprise pools. Remove any extra instances that were registered during the incident.
Trust boundary
The variable is repository-admin-only state: a pull request can neither set it nor read a different value into effect, and the expressions live in the base branch's workflow definition. This failover path therefore adds no PR-editable route to the self-hosted pool. Runner-side enforcement — an org-level runner group restricting these runners to the master-ref workflow — is tracked separately and composes with this mechanism.
Alternatives considered
Merge a workflow change to switch pools. Rejected because the outage that motivates the switch is exactly the state in which no PR can merge: the required checks are the ones failing. A repository variable is admin-controlled state that takes effect on re-run without a merge.
Keep the self-hosted pool always in the required path. Rejected because it trades hosted-pool availability for the in-house VM's, moving a single point of failure rather than adding a fallback. The variable keeps the hosted pools primary and the self-hosted pool a proven, one-action standby.
Consequences
Recovering from a hosted-pool outage is a single admin variable plus a re-run, with no merge on the critical path. The cost is a second runner topology to keep working: the standby lane exercises it on every master push so the failover target never goes stale, and the concurrency and cache-restore branches in ci.yml carry a selfhosted leg that must stay in step with the hosted leg.