Skip to content

Monitor Tab

The Monitor tab provides live health data for all clusters, links, and schema exporters in a DR setup.

The status header runs across the top of the tab and shows:

ElementDescription
Aggregate health iconsHealth indicator per cluster and link
Last checked timestampTime of the most recent health poll
Auto-refresh toggleEnables 5-second polling. Enabled by default.
Manual refresh buttonTriggers an immediate health poll
X-Ray toggleSwitches the view to the X-Ray guided operations panel
SLA breach badgeCount of active SLA breaches. Shown when at least one breach is active.

Monitor tab status header with health icons and X-Ray toggle

If DataReplicatorService stops writing health snapshots for longer than the staleness threshold, the status header shows a warning banner:

Monitoring feed down, cluster health is stale since {time}. Showing last known state.

The threshold defaults to 20 seconds, which is three to four times the service’s 5 second probe interval, and can be changed with the DR_HEALTH_SNAPSHOT_STALE_MS environment variable.

While the feed is stale, every cluster’s health shows as unknown rather than its last-known state, and the header’s freshness indicator changes from Updated: {time} to Stale: {elapsed}. This clears automatically on the next snapshot once the feed resumes.

Click a cluster’s health chip in the status header to open the Brokers modal. It opens with a summary line carrying the broker count, and, for anything other than Confluent Cloud, the current controller as “Controller: b{id}”. When no controller ID is reported, the summary shows a “No controller” warning only if the probe positively reported that no controller is present; if it simply couldn’t determine one, the line is left out entirely rather than warning. Below that is one row per broker, with a search box and pagination once there are more than eight. Each row shows:

ElementDescription
Broker {id}The broker identifier, with a Controller badge on the current controller. The badge and the summary’s controller line are both omitted for Confluent Cloud, where the reported controller rotates between probes
Availability zoneAn AZ badge, when the platform reports one
LatencyConnection latency in milliseconds, when reported
StatusReachable, Unreachable, or Unknown

If per-broker data hasn’t been reported yet, the modal reads “Per-broker reachability not reported. {N} broker(s) detected.”, or “No brokers detected yet.” when the count is zero. If no health snapshot exists yet, it shows “Cluster health not available yet.”

An interactive canvas shows all clusters as nodes and all links as edges, colour-coded by health:

  • Green: healthy
  • Yellow / Orange: degraded or warning
  • Red: failed or critical

Click a cluster node to expand all link accordions related to that cluster. Click a link edge to select that specific link.

When active SLA breaches exist, a banner appears above the link accordions. It turns red when any breach is Critical and amber when all of them are Warnings. Each row reads as one line, for example:

RPO Critical: Link link-1 forward lag exceeds target. Current: 2m 30s

RTO Warning: Cluster dr-east downtime exceeds target. Current: 5m

So each row carries the severity, the dimension (RPO or RTO), the link and direction or the cluster, and the current value. It does not repeat the configured threshold; for lag against threshold, use the SLA view.

Time-based lag is formatted for readability (for example 45s, 2m 30s, 1h 30m, 1d 12h). Downtime uses a coarser format with a single unit (for example 45s, 5m, 36.0h), so a multi-day outage reads in hours rather than days. Message-based lag values show as a locale-formatted count (for example 12,345 messages).

Each link has a collapsible accordion showing live data for both directions.

Each direction (forward / reverse) shows its current link status:

StatusMeaning
ActiveMirror topics are syncing normally
PausedMirror topic sync is paused
DegradedSync is impaired
PendingLink is initialising
DormantLink exists but is not yet active
PromotingMirror topics are being promoted (planned failover in progress)
PromotedMirror topics have been promoted to regular topics
Failed-OverEmergency failover completed; mirrors converted with potential data loss
StoppedLink has been stopped
FailedLink has encountered an error
  • Max lag (ms): worst single-topic replication lag in milliseconds
  • Max lag (messages): worst single-topic replication lag in message count
  • Last sync timestamp
  • Enable / disable toggle (requires Manage Links permission)

Per-exporter status for the link, consistent with the Schemas Tab.

The following actions appear inside each link accordion when conditions are met and the user has Failover permission:

ActionDescription
PromotePromotes mirror topics to regular topics (planned failover path)
FailoverExecutes an emergency per-link failover
Truncate & RestoreResets promoted mirrors on the original primary for re-import (emergency failback path)
Reverse & StartReverses the link direction and resumes replication (failback path)

For orchestrated failover and failback across all links, use X-Ray Guided Operations.