The Legacy Obituaries

Case q20 · howto

Our Azure standby is in sync and the on-premises primary is healthy. For the cutover, should we do a Data Guard switchover or a failover, and what do we check first?

Expected: Switchover, because the primary is healthy: zero data loss and the old primary stays as a standby for switch-back. Pre-flight with the Broker (SHOW CONFIGURATION success; VALIDATE DATABASE ready for switchover; zero transport/apply lag; Flashback Database on both), test client redirection first; failover only if the primary is gone; keep any observer off the primary host.

Added after the frozen run to test the Knowledge Base on how-to prose. Not counted in the frozen totals.

No tools

The model alone, answering from memory.
larger model · 1 turn · 20 s · 16.5k in / 1.9k out

✓ verdict ✗ grounded ✗ cited

Grader: Correct verdict; 12c, port 1521, V$ views unsupported by passages

Answer

Do a switchover, not a failover. The primary is healthy and the standby is in sync, which is exactly the case switchover is for. It's a planned role swap with zero data loss. The old on-prem primary becomes a standby that stays in sync, so you can switch back if the cutover goes wrong. Failover is for when the primary is lost. It can lose data, and the old primary has to be reinstated (Flashback) or rebuilt before you have a fallback again. Using failover here would give up your rollback path for no benefit.

Check these first (T-24h, then again right before the switchover):

  1. Broker readiness. In DGMGRL, run SHOW CONFIGURATION and look for SUCCESS with no warnings. Then run VALIDATE DATABASE VERBOSE <azure_stby> and look for Ready for Switchover: Yes. It also checks standby redo logs, the apply process and log file mismatches.
  2. Lag. Transport lag and apply lag should both be at or near 0. Check with SHOW DATABASE <azure_stby> or V$DATAGUARD_STATS. Confirm there are no archive gaps with V$ARCHIVE_GAP.
  3. Standby redo logs. Both sides need SRLs that are the same size as the online redo logs, with at least one more group per thread than the online logs have. The on-prem side needs them because it becomes the standby afterward.
  4. Dry run without the broker. ALTER DATABASE SWITCHOVER TO <azure_stby> VERIFY; (12c and later) reports blockers without changing any roles.
  5. Flashback Database on both sides. You don't strictly need it for a switchover, but it's your safety net if you later have to fail over or reinstate a database.
  6. Redo transport in the new direction. Check that Azure-to-on-prem redo will work: the LogXptMode properties, Oracle Net and TNS entries in both directions, NSG, firewall and ExpressRoute/VPN rules for port 1521 (or whatever port you use), and latency if you plan to run SYNC.
  7. Keys and passwords. The TDE wallet or keystore must be open and current on Azure (Key Vault or OKV if you use one). The password file must match on both sides.
  8. Application cutover. Clients should connect through role-based services (srvctl or DBMS_SERVICE with -role PRIMARY). Check DNS and SCAN changes, connection-string timeouts and retries, and FAN/TAC/AC if you use them. Also plan for jobs, schedulers and DB links that still point at on-prem.
  9. Operational hygiene. Take a fresh backup, pause heavy batch and RMAN jobs, check that no long-running transactions or conflicting ALTER SYSTEM work are in flight, and decide who can declare a go/no-go.
  10. Rollback plan. Write down the exact SWITCHOVER TO <onprem> switch-back steps and the time limit for using them. Test the whole sequence in non-prod first.

Execute: DGMGRL> SWITCHOVER TO <azure_stby>;. Then check SHOW CONFIGURATION, open services on the new primary, and confirm that redo is flowing to the on-prem standby before you declare done. Keep on-prem as the standby until you've passed your stability window.

References:

Keyword search

BM25 over the same 446 documents: top 8 passages pasted in, one pass, references not followed.
larger model · 1 turn · 8 s · 19.3k in / 0.7k out

✓ verdict ✓ grounded ✓ cited

Grader: Switchover and broker pre-flight covered; omits observer placement

8 keyword-search passages
  1. pattern_data-guard-switchover-vs-failover-at-cutover score 42.909
  2. glossary_switchover-vs-failover score 39.301
  3. pattern_rac-vs-data-guard-on-azure score 32.885
  4. feature_data-guard score 31.83
  5. mapping_data-guard--azure-sql-mi score 28.801
  6. pattern_choose-migration-method score 26.673
  7. source_ua-blog-dg-switchover-vs-failover score 24.29
  8. pattern_rman-moves-to-oracle-on-azure score 22.48
Answer

Do a switchover, not a failover. Your primary is healthy and reachable, so this is a planned role transition. A switchover loses no data and can be reversed.

Why a switchover [1][2]

  • The on-premises primary stops accepting transactions and ships its final redo. The Azure standby applies that redo and becomes the primary.
  • No data is lost because the old primary takes part in the hand-off.
  • The old on-premises primary stays on as a valid standby of the new Azure primary. If the cloud database misbehaves in the first few days, you can switch back without losing data. This is why Data Guard has the best rollback story of the migration methods. It's also what Zero Downtime Migration's physical-online mode does.
  • A failover is for when the primary has been lost. The standby is promoted with whatever redo it already has, so you could lose data: none in synchronized synchronous mode, seconds in Maximum Performance. The old primary also drops out of the configuration until you reinstate it with Flashback Database or rebuild it. There's no reason to accept that when your primary is healthy.

Check first, through the Broker [1]

  1. SHOW CONFIGURATION reports SUCCESS.
  2. VALIDATE DATABASE <standby> says the standby is ready for switchover.
  3. Transport lag and apply lag are both zero.
  4. Flashback Database is on for both databases.
  5. If the standby isn't ready, the usual causes are an archive gap, missing standby redo logs or apply lag. Fix these before the cutover, not during it.
  6. Test that clients get redirected to the new primary. The passage is cut off at this point, so it doesn't give the details of how to test this.

Source: Uptime Architect, Data Guard Switchover vs Failover: Which Role Transition, and When: https://uptimearchitect.com/blog/oracle-data-guard-switchover-vs-failover/ [7]

Sanity Context · GROQ

The agent with the GROQ endpoint: schema, joins, Portable Text.
smaller model · 12 turns · 55 s · 179.9k in / 4.6k out

✓ verdict ✓ grounded ✓ cited

Grader: Covers every expected element, including keeping the observer off the primary

11 tool calls
  1. ToolSearch on · 0 chars back
    select:mcp__sanity-groq__initial_context,mcp__sanity-groq__groq_query,mcp__sanity-groq__array_field_reader,mcp__sanity-groq__schema_explorer
  2. initial_context on groq · 4,126 chars back
  3. groq_query on groq · 662 chars back
    *[_type=="oracleFeature" && (name match "Data Guard*" || slug.current match "*data-guard*")]{_id, name, "slug": slug.current, category, introducedIn, deprecatedIn, desupportedIn, oracleReplacement}
  4. groq_query on groq · 1,337 chars back
    *[_type=="glossary" && (term match "*switchover*" || term match "*failover*" || term match "*Data Guard*")]{_id, term, definition, alsoKnownAs}
  5. groq_query on groq · 1,599 chars back
    *[_type=="mapping" && oracleFeature._ref=="feature_data-guard"]{_id, fidelity, effort, rationale, appliesToOracleVersions, "azureTarget": azureTarget->{name, "slug": slug.current}, "sources": sources[]->{title, url}}
  6. groq_query on groq · 452 chars back
    *[_type=="caveat" && mapping->oracleFeature._ref=="feature_data-guard"]{_id, severity, statement, appliesTo, firsthand, "evidence": evidence->{title, url}, "mappingId": mapping->_id}
  7. groq_query on groq · 747 chars back
    *[_type=="dispute" && (claimA->mapping->oracleFeature._ref=="feature_data-guard" || claimB->mapping->oracleFeature._ref=="feature_data-guard")]{_id, title, whatDisagrees, resolution, "claimA": claimA->{statement, "evidence": evidence->{title,url,publisher}}, "claimB": claimB->{statement, "evidence": evidence->{title,url,publisher}}, "resolvedBy": resolvedBy->{title,url}}
  8. groq_query on groq · 446 chars back
    *[_type=="caveat" && mapping._ref=="mapping_data-guard--azure-sql-mi"]{_id, severity, statement, appliesTo, firsthand, "evidence": evidence->{title, url, publisher}}
  9. groq_query on groq · 2,165 chars back
    *[_type=="pattern" && references("mapping_data-guard--azure-sql-mi")]{_id, title, problem, whenNotTo, "sources": sources[]->{title,url}}
  10. groq_query on groq · 5,475 chars back
    *[_type=="pattern" && (title match "*cutover*" || title match "*switchover*" || title match "*Data Guard*" || problem match "*cutover*" || problem match "*switchover*")]{_id, title, problem, whenNotTo, "sources": sources[]->{title,url}}
  11. array_field_reader on groq · 5,510 chars back
    {
     "mode": "range",
     "documentId": "pattern_data-guard-switchover-vs-failover-at-cutover",
     "field": "approach",
     "range": {
      "startIndex": 0,
      "endIndex": 30
     }
    }
Answer

Verdict: Do a switchover, not a failover — your stated conditions (standby in sync, primary healthy) are exactly the trigger for a switchover. Reserve failover strictly for a scenario where the primary is actually lost. Cut over with a Data Guard switchover, and reserve failover for a lost primary

Why:

  • A switchover is planned: the healthy primary ships its final redo, roles reverse with zero data loss, and the old primary is left as a valid standby — reversible, and it's the standard cutover move for a Data Guard–based migration. A failover is a response to a lost primary: the standby is promoted with whatever redo it has (loss depends on protection mode — none in synced Maximum Availability/synchronous mode, seconds in Maximum Performance), and the old primary drops out of the configuration until reinstated via Flashback Database or rebuilt. [Switchover vs failover (Data Guard) — glossary]
  • Calling a failover against a healthy source throws away the zero-loss hand-off and built-in rollback and can lose the last seconds of redo; attempting a switchover against an already-dead source cannot complete. Data Guard Switchover vs Failover: Which Role Transition, and When

What to check first (pre-flight, before touching anything):

  1. Confirm the source is actually healthy — if yes, this stays a switchover.
  2. Through the Data Guard Broker: SHOW CONFIGURATION reports success; VALIDATE DATABASE on the standby confirms it is ready for switchover, with transport and apply lag at zero, and Flashback Database enabled on both primary and standby (needed as the safety net even though the plan is a switchover).
  3. Test client redirection to the Azure primary ahead of time — a stalled cutover on firewall ports or a decommissioned host is a network/readiness failure, not a Data Guard failure.
  4. If Fast-Start Failover/observer is configured, keep the observer off the primary host (place it on a separate small VM, per the Azure VM reference architecture pattern), since an observer on the primary dies with the failure it's meant to detect.
  5. Rehearse the switchover at least once beforehand — "a standby that has never been switched over is a hope, not a plan."

Then: quiesce/drain the application, run the switchover through the Broker, verify the new Azure primary shows role PRIMARY and open mode READ WRITE, confirm redo is now flowing back to the on-prem standby, and keep the old primary as a standby for an agreed fallback window before decommissioning. Cut over with a Data Guard switchover, and reserve failover for a lost primary · Architectures for Oracle database on Azure Virtual Machines

Scope check: This pattern applies when the Azure target is a physical standby (Oracle on Azure VM or Oracle Database@Azure). It does not apply for cross-endian or cross-version moves, Standard Edition 2 sources, Autonomous Database targets, or Azure SQL/PostgreSQL targets — those cut over via GoldenGate drain-and-reconnect or a Data Pump outage instead, with no switch-back option. Cut over with a Data Guard switchover, and reserve failover for a lost primary

Gaps: No caveat or dispute documents in the dataset are attached specifically to the Data Guard cutover pattern or its mapping — the guidance above is drawn from the pattern document only. The one mapping document tied to the Data Guard feature targets Azure SQL Managed Instance (failover groups, partial fidelity) and is about ongoing DR parity, not this migration-cutover question, so it isn't cited above as directly on-point.

GROQ + Knowledge Base

The agent with both endpoints.
smaller model · 12 turns · 52 s · 250.3k in / 3.9k out

✓ verdict ✓ grounded ✓ cited

Grader: Switchover and full pre-flight correct; omits observer placement

11 tool calls
  1. ToolSearch on · 0 chars back
    select:mcp__sanity-groq__initial_context,mcp__sanity-kb__initial_context,mcp__sanity-groq__groq_query,mcp__sanity-groq__array_field_reader,mcp__sanity-kb__knowledge_base_read,mcp__sanity-groq__schema_explorer
  2. initial_context on groq · 4,126 chars back
  3. initial_context on kb · 7,679 chars back
  4. knowledge_base_read on kb · 5,894 chars back
    {
     "knowledgeBase": "kb6MdvV3ZA3Q",
     "paths": [
      "ha_and_cutover"
     ]
    }
  5. groq_query on groq · 682 chars back
    *[_type=="oracleFeature" && (name match "*Data Guard*" || slug.current match "*data-guard*")]{_id, name, slug, category, introducedIn, deprecatedIn, desupportedIn, oracleReplacement}
  6. groq_query on groq · 1,973 chars back
    *[_type=="mapping" && oracleFeature._ref=="feature_data-guard"]{_id, fidelity, effort, rationale, appliesToOracleVersions, "azureTarget": azureTarget->{name}, "sources": sources[]->{title,url,publisher}, "caveats": *[_type=="caveat" && references(^._id)]{severity, statement, appliesTo, firsthand, "evidence": evidence->{title,url,publisher}}}
  7. groq_query on groq · 1,588 chars back
    *[_type=="glossary" && (term match "*switchover*" || term match "*failover*")]{_id, term, definition, alsoKnownAs, "sources": sources[]->{title,url,publisher}}
  8. groq_query on groq · 813 chars back
    *[_type=="source" && (title match "*RAC*Data Guard*" || title match "*switchover*" || title match "*Data Guard*")]{_id,title,url,publisher}
  9. groq_query on groq · 11,032 chars back
    *[_type=="pattern" && (title match "*Data Guard*" || title match "*cutover*" || title match "*RAC*")]{_id,title,problem,"sources":sources[]->{title,url,publisher}}
  10. groq_query on groq · 4,209 chars back
    *[_type=="dispute"]{_id,title,"claimA_text":claimA->statement,"claimB_text":claimB->statement}
  11. groq_query on groq · 433 chars back
    *[_type=="mapping" && references("feature_data-guard") && azureTarget->name match "*Oracle*"]{_id, fidelity, "azureTarget": azureTarget->name, appliesToOracleVersions, rationale}
Answer

Verdict: Switchover, not failover. Since your on-premises primary is healthy and the Azure standby is in sync, this is a planned role transition — the definition of a switchover, not a failover, which is reserved for a lost/unreachable primary. Cut over with a Data Guard switchover, and reserve failover for a lost primary; Architectures for Oracle database on Azure Virtual Machines

Why it matters: A switchover ships the primary's final redo, reverses roles with zero data loss, and leaves the old (on-prem) primary as a valid standby — reversible if problems surface. A failover promotes the standby with whatever redo it already has (zero loss only in synchronous Maximum Availability/Protection mode, seconds of loss in Maximum Performance), and the old primary is excluded from the configuration until reinstated via Flashback Database or a rebuild. Calling a failover against a healthy source throws away the zero-loss hand-off and the built-in rollback for no reason. Data Guard Switchover vs Failover: Which Role Transition, and When

What to check first — pre-flight checklist via Data Guard Broker, before opening the cutover window:

  1. SHOW CONFIGURATION reports success
  2. VALIDATE DATABASE on the Azure standby shows ready for switchover
  3. Transport and apply lag are at zero
  4. Flashback Database is enabled on both databases (this is also your fallback path if something forces an actual failover later)
  5. Client redirection to the Azure primary has been tested

Common pre-flight failures to rule out: archive gap, missing standby redo logs, apply lag, and firewalls blocking client redirection — fix these before the window, not during it. Cut over with a Data Guard switchover, and reserve failover for a lost primary

Cutover sequence once checks pass: quiesce/drain the app → run the switchover through Broker → verify the new (Azure) primary shows role PRIMARY and open mode READ WRITE → confirm redo is now flowing back to the old on-prem instance (reverse redo flow, proving it's a valid standby) → keep the old primary as standby for an agreed fallback window before decommissioning. Rehearse at least one full switchover before the real event — an untested standby is "a hope, not a plan." Cut over with a Data Guard switchover, and reserve failover for a lost primary

Scope caveats:

Disputes: None of the dataset's 5 recorded disputes touch Data Guard, switchover/failover, or RAC — no conflicting claims to present here.

← q19All cases