Failover and Failback Workflows
Failover and failback are operator-driven workflows executed through the X-Ray panel in the Monitor tab. Each workflow requires Failover permission.
Pre-Flight Checks
Section titled “Pre-Flight Checks”Pre-flight checks run at the start of Planned Failover and Planned Failback, when Pre-failover Validation is enabled for the setup. If Pre-failover Validation is disabled, the checklist panel does not render at all and the operator proceeds directly through the Planned workflow without health-gate enforcement. In Emergency modes the Live Damage Report is shown instead of the checklist, regardless of the toggle. None of its rows can block the workflow, but the step has a gate of its own: Next is disabled unless at least one link is affected.
| Check | Pass Condition | Planned Severity | Emergency Severity |
|---|---|---|---|
| Target cluster reachable | Bootstrap probe reports up | Fail (cannot override) | Fail (cannot override) |
| Source cluster reachable | Bootstrap probe reports up | Warn | Warn |
| All link directions active | All directions in an operational state | Block (overridable) | Warn |
| Replication lag within RPO | No active SLA breaches | Block (overridable) | Warn |
| Consumer group offset sync | Sync enabled and last sync < 60 seconds ago | Warn | Warn |
| No partial-failover remnants | No Failed direction alongside a completed direction (Failed-Over or Promoted) on a FailedOver link | Warn | Warn |
| Consumer groups covered by replication | All source CGs are in the replication filter | Warn | Warn |
Severity key:
- Fail: hard block. Cannot proceed. Fix the underlying issue.
- Block: soft block. Operator can override with an explicit acknowledgement when no hard fails exist.
- Warn: informational. Operator may proceed.
What a Failover or Failback Applies To
Section titled “What a Failover or Failback Applies To”The guided steps on the right side of X-Ray render only after you select a link direction in the sidebar. Nothing is selected when you first open the view, so the panel reads “Pick a link direction to inspect” until you choose one, and no operation buttons are available before then.
The sidebar takes one direction at a time. Every operation you start from X-Ray therefore applies to that single link, and the confirm dialog names the resolved scope, which reads as Partial for one link.
Live Damage Report
Section titled “Live Damage Report”In Emergency modes, the Assess Damage step replaces the pre-flight checklist with a Live Damage Report:
| Row | Content |
|---|---|
| Cluster reachability | One row per cluster, colour-coded by status: green “reachable”, red “down {duration}” or “unreachable” when the downtime isn’t known, amber “degraded” (still serving traffic), or grey “unknown” |
| Link health | One row per link, colour-coded by severity, showing the link name and its worst (max) lag |
| Active SLA Breaches | Listed below the link rows as the breach level (Critical or Warning), the link name, and the source and destination cluster aliases |
A cluster reads “unknown” whenever its health can’t be established, which covers an unprobeable cluster (a Kerberos or OAuth probe failure, for example), a missing health entry, and a stale monitoring feed.
All links appear in the report, whether or not they are affected. The Next button is enabled only when at least one link is affected, and its tooltip names the count. A link counts as affected when its source cluster is down, when either direction reports Failed, or when the link itself is Failed or Degraded, which is what the button’s own tooltip means by “a link is unhealthy”.
Planned Failover executes a controlled, pre-checked cutover to the secondary cluster.
When to Use
Section titled “When to Use”Use Planned Failover when the primary cluster is healthy and you want a zero-data-loss cutover, for example, during scheduled maintenance or a controlled migration.
-
Pre-Flight Checks: checks run automatically. Review results. Resolve or override any block-level items. Hard fails must be fixed before proceeding.
-
Shutdown Producers & Consumers: stop all producer and consumer clients on the source cluster. When done, click Confirm Shutdown to record the attestation.
-
Reverse and Start: click the Reverse and Start button to promote mirror topics to regular topics on the destination and reverse the link direction. A confirmation modal appears before the operation executes.
-
Restart Producers & Consumers on Destination: restart all clients, now pointed at the new primary. Click Confirm Restart to record the attestation.
-
Verify Traffic: confirm that traffic is flowing on the new primary. Click Verify Traffic to finalise the failover.
Status transitions: Replicating / Degraded / Paused -> Failover-In-Progress -> Failed-Over
Emergency Failover executes an immediate cutover without a pre-check gate.
When to Use
Section titled “When to Use”Use Emergency Failover when the primary cluster is unavailable and waiting for pre-checks to pass is not possible.
-
Assess Damage: review the Live Damage Report: cluster reachability, per-link health and lag, and active SLA breaches. Nothing in the report blocks you, but Next is disabled until at least one link is affected, and the links you select in the sidebar are the ones the failover applies to.
-
Shutdown Producers & Consumers: stop all clients on the source (if reachable). Click Confirm Shutdown.
-
Failover Mirrors (Immediate): mirror topics are force-converted to regular topics on the destination. This is a hard cutover. The suite does not wait for lag to drain.
-
Restart Producers & Consumers on Destination: restart clients on the new primary. Click Confirm Restart.
-
Verify Traffic: confirm traffic is flowing. Click Verify Traffic to finalise.
Status transitions: Replicating / Degraded / Paused / Failed -> Failover-In-Progress -> Failed-Over
Planned Failback returns traffic to the original primary after a Planned Failover.
When to Use
Section titled “When to Use”Use Planned Failback when the original primary has been restored and you want to return to the normal operating configuration with zero data loss.
Schema Failback Mode
Section titled “Schema Failback Mode”If schema replication is enabled, select a schema failback mode before initiating:
| Mode | Behaviour | When to Use |
|---|---|---|
| Skip | No schema orchestration. Operator handles schema replication manually. | When you want full manual control. |
| Attempt | Patches the existing schema exporter without clearing the destination schema registry. | Default. Try this first. |
| Force | Deletes scoped subjects on the destination registry first, then patches the exporter. | When Attempt fails due to import-mode collisions from outage-era schemas. |
-
Pre-Flight Checks: checks run against the reverse direction (secondary -> primary). Resolve or override block-level items.
-
Shutdown Producers & Consumers: stop all clients currently using the failover cluster. Click Confirm Shutdown.
-
Reverse and Start: re-establishes replication back to the original primary. The link direction is reversed and mirroring resumes.
-
Restart Producers & Consumers on Original Primary: redirect clients back to the original primary. Click Confirm Restart.
-
Verify Traffic: confirm traffic is flowing on the original primary. Click Verify Traffic to finalise the failback.
Status transitions: Failed-Over -> Failback-In-Progress -> Replicating
Emergency Failback returns traffic to the original primary after an Emergency Failover.
When to Use
Section titled “When to Use”Use Emergency Failback when the original primary has recovered and an emergency failover was previously executed. Because mirror topics were force-promoted during the emergency failover, the original primary must be reset before replication can resume.
-
Assess Damage: review the Live Damage Report for the reverse direction: cluster reachability, link health and lag, and active SLA breaches. Nothing in the report blocks you, but Next is disabled until at least one link is affected.
-
Shutdown Producers & Consumers: stop all clients on the current primary (the failover cluster). Click Confirm Shutdown.
-
Truncate and Restore: the suite resets the promoted mirror topics on the original primary so they can re-import data from the current primary. This step prepares the original primary for replication.
-
Reverse and Start: flips the link direction and resumes mirroring from the current primary back to the original primary.
-
Restart Producers & Consumers on Original Primary: redirect clients back. Click Confirm Restart.
-
Verify Traffic: confirm traffic is flowing on the original primary. Click Verify Traffic to finalise.
Status transitions: Failed-Over -> Failback-In-Progress -> Replicating