Skip to main content

Command Palette

Search for a command to run...

OEM 24ai and Data Guard: The Metrics That Catch a Broker Failure Before Your RPO Does

Updated
•8 min read•View as Markdown
R
Transforming Reactive Monitoring into AI-driven Multi Cloud Observability.

Oracle observability post #9 — the last two posts covered the OEM job system and HA/DR architecture for the management tier itself. This one turns to something a lot of shops get half-right: monitoring Data Guard through OEM instead of just running dgmgrl by hand when someone asks "are we in sync?"

Here's the problem I run into constantly. A customer has Data Guard configured correctly — broker in place, Fast-Start Failover enabled, protection mode set deliberately. Then I ask how they'd know if apply lag started climbing at 2 a.m. on a Saturday, and the answer is some version of "someone would notice during the morning check." That's not monitoring. That's hoping the person who notices, notices in time.

Why DGMGRL Isn't Enough

dgmgrl and SHOW CONFIGURATION VERBOSE are the right tools when you're actively troubleshooting. They're the wrong tools for finding out something's wrong in the first place, because they require a human to run them. Data Guard failure modes are exactly the kind that develop quietly: a network blip causes a redo transport gap, the gap doesn't resolve because the archive isn't reachable, and six hours later your RPO is blown and nobody got paged because nobody was looking.

OEM's value here isn't that it knows something dgmgrl doesn't — it's the same broker data underneath. The value is that OEM polls it continuously, evaluates it against thresholds you set once, and routes a notification the moment something crosses the line. That's the difference between "we would have caught it" and "we caught it."

Getting Data Guard Under OEM's Management

If your primary and standby are already OEM targets — which they should be if you followed target discovery in this series — Data Guard monitoring isn't a separate onboarding step so much as a configuration step on targets you already have.

Console path: Targets → Databases → (select the primary) → Availability → Data Guard Administration. From here OEM reads the broker configuration directly and gives you the same primary/standby topology view that SHOW CONFIGURATION gives you at the command line — protection mode, Fast-Start Failover status, and the current role of each member.

If Data Guard hasn't been added to the broker configuration yet, or the standby isn't showing up correctly, verify the broker is actually managing the pair before you troubleshoot OEM's view:

Note: EM CLI syntax below is representative. Exact parameter names vary by OEM version and job type. Always validate against your environment: emcli help <verb> and Oracle's official EM CLI reference.

dgmgrl sys/password@primary_db
DGMGRL> show configuration verbose;
DGMGRL> show database verbose 'standby_db';

Confirm Transport Lag and Apply Lag are both reporting real values (not blank, not stuck at zero with the managed recovery process down) before you move to OEM-side threshold configuration. A broker that isn't tracking lag correctly will feed OEM the same bad data.

The Metrics That Actually Matter

OEM's Data Guard Performance metric group gives you more numbers than you need. Three are worth alerting on for almost every configuration; the rest are useful for diagnosis after something already fired.

Apply Lag. How far behind the primary the standby's applied redo actually is, in seconds. This is your RPO exposure in real time. Set the warning threshold to whatever your documented RPO tolerance is, and the critical threshold to double that — if you promised the business a 5-minute RPO, warning at 5 minutes and critical at 10 gives operations a window to react before you've actually breached the SLA.

Transport Lag. How far behind the primary the standby is in terms of redo received, not yet applied. A transport lag that's climbing while apply lag stays flat usually means a network or archiving problem, not an apply-side bottleneck — useful to alert on separately so your team isn't guessing which layer broke.

Log Transport / Gap Status. Whether there's an actual gap in the redo stream between primary and standby. This should page immediately, not wait for a lag threshold to be crossed, because a gap means the automatic resolution process has failed and someone needs to intervene before the archived logs it needs get purged.

Note: EM CLI syntax below is representative. Exact parameter names vary by OEM version and job type. Always validate against your environment: emcli help <verb> and Oracle's official EM CLI reference.

emcli get_metric_data -target_name="standby_db" -target_type="oracle_database" -metric_name="dataguard_performance" -metric_column="applylag"
emcli set_metric_threshold \
  -target_name="standby_db" -target_type="oracle_database" \
  -metric_name="dataguard_performance" -metric_column="applylag" \
  -warning_threshold="300" -critical_threshold="600"

Don't Skip Fast-Start Failover Health

If you're running Fast-Start Failover, monitoring the databases isn't enough — you also need to know the Observer is alive. The Observer is the thing that actually initiates failover when the primary goes dark, and it runs as a lightweight process independent of both databases. If it dies quietly on whatever host it's running on, you've silently disabled your automatic failover and nothing about the primary or standby will tell you that.

Set up a host-level target and process monitoring on wherever your Observer runs, and treat an Observer outage as a critical incident — not a warning. This is the single most common Data Guard monitoring gap I find in customer environments: broker configured correctly, FSFO enabled, and an Observer that's been dead for three weeks because it was running on a host that got patched and rebooted without anyone restarting it.

Protection Mode Changes Are Events Too

Most shops set a protection mode once at build time — Maximum Availability, usually — and never think about it again. But protection mode can change silently. If a network issue forces a synchronous standby to fall behind past the NetTimeout window, Data Guard can automatically downgrade from synchronous to asynchronous redo transport to keep the primary from stalling. That's the broker doing exactly what it's supposed to do, and it's also a silent reduction in your data-loss guarantee that nobody told the business about.

Alert on protection mode changes as their own event, separate from lag thresholds. A downgrade from Maximum Availability to a degraded transport mode should generate a notification even if lag stays within tolerance, because the guarantee you're actually operating under just changed, not just the number.

Don't Forget Switchover Readiness

Monitoring lag and gaps tells you Data Guard is healthy today. It doesn't tell you the standby is actually ready to take over if you need it to. Periodically validate switchover readiness rather than assuming continuous replication equals continuous failover-readiness — a standby can be perfectly in sync on data and still fail a switchover because of a stale password file, a missing standby redo log, or a tnsnames entry nobody updated after the last host migration.

Console path: Targets → Databases → (select the primary) → Availability → Data Guard Administration → Verify Configuration, which runs the broker's built-in configuration checks without actually performing a role transition. Schedule this as a recurring OEM job (see this series' post on the job system) rather than relying on someone remembering to click it before your next DR test.

Wiring Incident Rules Correctly

This series covered incident rules and notification routing back in post #3, and Data Guard is where that setup actually earns its keep. Don't route Data Guard alerts into the same general database incident rule set that catches tablespace warnings and long-running-session alerts — a gap or an Observer failure needs to reach a different, faster escalation path than "someone will look at it during business hours."

Build a dedicated rule set scoped to your Data Guard target metrics (apply lag critical, transport lag critical, gap status, Observer down) and route it to whatever channel actually gets an immediate human response — paging, not just email. Test it by manually inducing a gap in a non-production standby and confirming the notification actually arrives and reaches the right person before you trust it in production.

What Good Looks Like

A properly monitored Data Guard configuration means nobody has to remember to check it. Apply lag and transport lag thresholds are set to your documented RPO, not to whatever OEM shipped with by default. Gap status and Observer health page immediately rather than waiting for a lag threshold to trip. The alert routes to on-call, not to a shared inbox someone checks in the morning. And when you do get paged, the HA Console gives you the same topology and lag data you'd have pulled from dgmgrl anyway — so the page and the diagnosis use the same numbers instead of forcing you to reconcile two different views of reality.

The Bottom Line

Data Guard without monitoring is an insurance policy you haven't actually verified is active. The broker will keep quietly doing its job right up until the moment something breaks the automatic recovery path, and if nobody's watching, you find out during the outage instead of before it. The configuration work here is a few hours per database once you've built the monitoring template — it's not the hard part. The hard part is admitting that "the DBA checks it manually" was never actually a monitoring strategy.

Next up: OEM 24ai self-monitoring — what happens when the repository that's tracking everyone else's health problems develops its own.

Oracle EM 24ai & OCI Observability

Part 8 of 8

A comprehensive series covering Oracle Enterprise Manager 24ai and OCI Observability. Explore AI-powered monitoring, diagnostics, and observability features to manage and optimize your Oracle Cloud Infrastructure environments.

Start from the beginning

Your Enterprise Manager 24ai Is Installed. Your Oracle Stack Is Still Flying Blind!

Actionable Steps to Close the Observability Gap

More from this blog

E

Enterprise Management & OCI Observability | rajeshravi.com

9 posts

Regular deep dives on Oracle Enterprise Manager, OCI Observability & Management, and multicloud monitoring — written by an expert practitioner who has implemented these stacks for hundreds of enterprise customers. Expect version-specific guidance, real-world architecture patterns, and honest takes on what actually works in production.