Skip to content

Failover and Failback Workflows

Failover and failback are operator-driven workflows executed through the X-Ray panel in the Monitor tab. Each workflow requires Failover permission.

Pre-flight checks run at the start of Planned Failover and Planned Failback, when Pre-failover Validation is enabled for the setup. If Pre-failover Validation is disabled, the checklist panel does not render at all and the operator proceeds directly through the Planned workflow without health-gate enforcement. In Emergency modes the Live Damage Report is shown instead of the checklist, regardless of the toggle. None of its rows can block the workflow, but the step has a gate of its own: Next is disabled unless at least one link is affected.

CheckPass ConditionPlanned SeverityEmergency Severity
Target cluster reachableBootstrap probe reports upFail (cannot override)Fail (cannot override)
Source cluster reachableBootstrap probe reports upWarnWarn
All link directions activeAll directions in an operational stateBlock (overridable)Warn
Replication lag within RPONo active SLA breachesBlock (overridable)Warn
Consumer group offset syncSync enabled and last sync < 60 seconds agoWarnWarn
No partial-failover remnantsNo Failed direction alongside a completed direction (Failed-Over or Promoted) on a FailedOver linkWarnWarn
Consumer groups covered by replicationAll source CGs are in the replication filterWarnWarn

Severity key:

  • Fail: hard block. Cannot proceed. Fix the underlying issue.
  • Block: soft block. Operator can override with an explicit acknowledgement when no hard fails exist.
  • Warn: informational. Operator may proceed.

The guided steps on the right side of X-Ray render only after you select a link direction in the sidebar. Nothing is selected when you first open the view, so the panel reads “Pick a link direction to inspect” until you choose one, and no operation buttons are available before then.

The sidebar takes one direction at a time. Every operation you start from X-Ray therefore applies to that single link, and the confirm dialog names the resolved scope, which reads as Partial for one link.

In Emergency modes, the Assess Damage step replaces the pre-flight checklist with a Live Damage Report:

RowContent
Cluster reachabilityOne row per cluster, colour-coded by status: green “reachable”, red “down {duration}” or “unreachable” when the downtime isn’t known, amber “degraded” (still serving traffic), or grey “unknown”
Link healthOne row per link, colour-coded by severity, showing the link name and its worst (max) lag
Active SLA BreachesListed below the link rows as the breach level (Critical or Warning), the link name, and the source and destination cluster aliases

A cluster reads “unknown” whenever its health can’t be established, which covers an unprobeable cluster (a Kerberos or OAuth probe failure, for example), a missing health entry, and a stale monitoring feed.

All links appear in the report, whether or not they are affected. The Next button is enabled only when at least one link is affected, and its tooltip names the count. A link counts as affected when its source cluster is down, when either direction reports Failed, or when the link itself is Failed or Degraded, which is what the button’s own tooltip means by “a link is unhealthy”.

Planned Failover executes a controlled, pre-checked cutover to the secondary cluster.

Use Planned Failover when the primary cluster is healthy and you want a zero-data-loss cutover, for example, during scheduled maintenance or a controlled migration.

  1. Pre-Flight Checks: checks run automatically. Review results. Resolve or override any block-level items. Hard fails must be fixed before proceeding.

  2. Shutdown Producers & Consumers: stop all producer and consumer clients on the source cluster. When done, click Confirm Shutdown to record the attestation.

  3. Reverse and Start: click the Reverse and Start button to promote mirror topics to regular topics on the destination and reverse the link direction. A confirmation modal appears before the operation executes.

  4. Restart Producers & Consumers on Destination: restart all clients, now pointed at the new primary. Click Confirm Restart to record the attestation.

  5. Verify Traffic: confirm that traffic is flowing on the new primary. Click Verify Traffic to finalise the failover.

Status transitions: Replicating / Degraded / Paused -> Failover-In-Progress -> Failed-Over