Architect highly available Oracle Database on Amazon EVS and FSx for ONTAP¶
Summary¶
An AWS Architecture Blog reference architecture for running Oracle Database at high availability on Amazon EVS (Elastic VMware Service) backed by Amazon FSx for NetApp ONTAP storage, with cross-region disaster recovery via NetApp SnapMirror. The design targets enterprises with existing VMware Cloud Foundation (VCF) investments who want to migrate Oracle to AWS without rearchitecting applications or retraining operations teams — EVS runs VCF 9.1 natively on EC2 bare-metal inside the customer's VPC, so vSphere/NSX/vSAN operational tooling is preserved while gaining sub-millisecond storage latency, snapshot backup, independent storage scaling, and cross-region DR. The post's durable system-design content is the two-tier storage split (local vSAN for transient I/O + FSx ONTAP NFS datastores for persistent/replicated Oracle data), the host-level NFS datastore vs in-guest dNFS decision (avoiding an NSX Edge choke point), the SnapMirror async cross-region DR topology with pre-provisioned standby VMs, and the Oracle DR-licensing constraint that shapes how replicated volumes are mounted.
Key takeaways¶
-
EVS = VCF 9.1 natively on EC2 bare metal inside the customer VPC (self-deployed model). With VCF 9.x, Amazon EVS provisions the bare-metal infrastructure and VLAN subnets; the customer then deploys VCF using Broadcom's VCF Installer. VCF 9.x supports evaluation mode (validate before applying license keys), and the Solutions-for-Amazon-EVS GitHub repo provides CloudFormation and Terraform templates to automate the phased deployment. An EVS environment supports 4–32 hosts per cluster; instance types can be mixed across clusters within the same SDDC. (Source: sources/2026-10-02-aws-architect-highly-available-oracle-database-on-amazon-evs-and-fsx-ontap)
-
Two-tier storage split: vSAN for transient, FSx ONTAP for persistent + replicated. vSAN (local NVMe on each ESXi host) handles low-latency, non-replicated workloads — VM boot disks, OS swap, Oracle temp tablespace (single-digit-ms latency at no extra cost). FSx for NetApp ONTAP handles Oracle data files and redo logs that require snapshot backup and cross-region replication. ESXi hosts mount FSx ONTAP volumes as NFS datastores; Oracle VMs access standard VMDKs on those datastores. This is a clean storage-media-tiering decision made on the replication/durability axis, not just the speed axis.
-
Host-level NFS datastore beats in-guest dNFS — to avoid an NSX Edge choke point. With in-guest NFS (Oracle dNFS), all NFS traffic routes through the NSX overlay and an NSX Edge node before reaching FSx ONTAP — a single Edge chokepoint. With NFS datastores, each ESXi host talks NFS directly to FSx ONTAP using its own network bandwidth — no Edge bottleneck, no extra latency hop, bandwidth distributed across the cluster. Each host has its own independent NFS path. (Source: sources/2026-10-02-aws-architect-highly-available-oracle-database-on-amazon-evs-and-fsx-ontap)
-
SnapMirror async cross-region DR with pre-provisioned standby VMs. SnapMirror replicates FSx ONTAP volumes (data, logs, binaries) to the DR region; replication frequency determines RPO. Pre-provisioning standby Oracle VMs in the DR cluster reduces RTO. Replicating binary volumes means Oracle installation is not required during recovery. This is the async-cross-region-replication
-
pilot-light shape applied to a VMware/Oracle stack. (Source: sources/2026-10-02-aws-architect-highly-available-oracle-database-on-amazon-evs-and-fsx-ontap)
-
Oracle DR-licensing constraint shapes the mount topology. To avoid Oracle double licensing, keep DR replicated volumes as data-protection (DP) volumes that are NOT mounted as NFS datastores on DR hosts until a failover is declared. Only then break the SnapMirror relationship, mount the datastore, and power on the VM. Pre-mounting the SnapMirror volume as a datastore — even with no VM powered on — means Oracle binaries are accessible on those hosts, which Oracle may consider an "installation" requiring licenses across the entire DR cluster. This is a licensing constraint directly dictating a technical DR runbook. (Source: sources/2026-10-02-aws-architect-highly-available-oracle-database-on-amazon-evs-and-fsx-ontap)
-
FSx ONTAP read/write throughput asymmetry is a sizing trap. On a 6 GB/s filesystem, read throughput can reach 6 GB/s but write throughput is limited to ~1 GB/s — and this write ceiling applies per high-availability (HA) pair. To scale write throughput beyond one HA pair, deploy a file system with more than one HA pair and manually distribute Oracle volumes across the additional aggregates (placement is not automatic). Size throughput capacity on the write workload. (Source: sources/2026-10-02-aws-architect-highly-available-oracle-database-on-amazon-evs-and-fsx-ontap)
-
100% SSD, no tiering, 20% headroom for Oracle volumes. Keep the FSx ONTAP tiering policy set to
nonefor all Oracle volumes (no capacity-pool tiering). If the SSD tier fills to capacity, FSx ONTAP blocks writes — provision a minimum 20% free headroom and monitor SSD utilization. (Source: sources/2026-10-02-aws-architect-highly-available-oracle-database-on-amazon-evs-and-fsx-ontap) -
One-way BGP via VPC Route Server; static beyond the VPC. The NSX Tier-0 gateway peers through BGP with Amazon VPC Route Server — but it is one-way BGP: Route Server listens for routes advertised by NSX and writes them into the VPC route table, but does not advertise VPC routes back to NSX. Beyond the VPC, routes stay static at Transit Gateway. Security group rules are not enforced on VLAN subnet interfaces — use network ACLs and the NSX distributed firewall for traffic control. (Source: sources/2026-10-02-aws-architect-highly-available-oracle-database-on-amazon-evs-and-fsx-ontap)
Architecture — eight data-flow components¶
The reference architecture spans two AWS Regions with this numbered flow (from on-premises through hybrid connectivity into the production EVS environment and cross-region DR):
- Hybrid link — on-prem data center → AWS via AWS Direct Connect for dedicated, low-latency bandwidth.
- Transit routing — AWS Transit Gateway routes between the production VPC, DR region, and on-prem.
- Workload landing — traffic reaches the EVS Database Cluster on
i7i.metal-24xlEC2 bare-metal instances with VCF 9.1. - Live migration — VMware HCX (Hybrid Cloud Extension) migrates Oracle VMs from on-prem VMware to EVS with near-zero downtime using Replication Assisted vMotion or bulk migration.
- NFS data path — each ESXi host mounts FSx ONTAP as an NFS datastore; Oracle VMs access volumes as VMDKs with sub-ms latency and up to 80,000 IOPS; each host has an independent NFS path.
- Cross-region DR — SnapMirror asynchronously replicates FSx ONTAP volumes (data, logs, binaries) to the DR region at a configurable RPO.
- Dynamic routing — NSX Tier-0 gateway peers via one-way BGP with VPC Route Server (listens, writes to VPC route table; does not advertise VPC routes back).
- DR failover — on failover, Transit Gateway routes to the DR region where pre-provisioned standby Oracle VMs mount the SnapMirror replica volumes.
Instance sizing (operational numbers)¶
ESXi host instance types supported by EVS (choice = per-core compute
speed i7i vs per-host density i4i):
| Instance | vCPU / cores | RAM | Notes |
|---|---|---|---|
| i7i.metal-24xl (recommended default) | 96 / 48 | 768 GiB | 5th Gen Intel Xeon; 3rd Gen Nitro SSDs for vSAN; ~10% price-performance over i4i.metal |
| i7i.metal-48xl | 192 / 96 | 1,536 GiB | double capacity/host; up to 100 Gbps net, 60 Gbps EBS; vertical scale for large SGAs; same per-core profile |
| i4i.metal | 128 / 1,024 | 1,024 GiB | higher density; fewer/larger hosts to reduce VMware per-host licensing |
- All three support VCF 9.1 with ESXi 9.1 (build
9.1.0.0100.25433460). VCF 5.2.2 (ESXi 8.0U3g) remains available but is on a path to end of support; new deployments should default to 9.x. - FSx ONTAP: sub-ms latency, multiple GB/s throughput, up to 80,000 IOPS per file system.
- Summary claims: i7i.metal-24xl delivers up to 23% better compute and 50% lower IO latency; "recovery in seconds" via SnapCenter regardless of DB size.
- VM sizing: match vCPU to Oracle
CPU_COUNT; size memory for SGA+PGA+OS (SGA typically 75–85% of VM memory); place VM swap + Oracle temp tablespace on vSAN (not NFS); NUMA-aware placement for VMs exceeding single-socket core count.
Migration options (on-prem VMware → EVS)¶
| # | Method | Best for | Downtime |
|---|---|---|---|
| 1 | VMware HCX Live Migration | Existing VMware on-prem | Near-zero (vMotion) or planned bulk |
| 2 | SnapMirror ONTAP-to-ONTAP |
On-prem Oracle on NetApp ONTAP | Minutes (final sync switchover) |
| 3 | Oracle PDB Relocation | PDB/CDB multitenant | Brief (final switchover only) |
| 4 | RMAN Backup/Restore | Non-ONTAP on-prem (universal) | Hours (backup + restore + apply) |
Security layers¶
NSX-T Tier-1 gateways isolate DB/App/perimeter segments; NSX Distributed Firewall restricts Oracle listener (TCP 1521) to authorized app segments (micro-segmentation prevents lateral movement); FortiGate (or equivalent) inspection VPC for N-S filtering; FSx ONTAP volumes encrypted with AWS KMS; VPC encryption for NFS + IPsec for SnapMirror cross-region; Oracle TDE for DB-level encryption; zero-trust access (Banyan/Zscaler) for VMware admin consoles; network ACLs (not security groups) on VLAN interfaces.
Caveats¶
- Reference architecture, not a production retrospective. No named customer, no measured production numbers from a real deployment. The 23% compute / 50% IO-latency / "recovery in seconds" figures are AWS marketing claims, not an incident/scale writeup.
- Oracle licensing guidance is repeatedly hedged with "request an AWS Optimization and Licensing Assessment (AWS OLA)" — the post is not authoritative legal/licensing advice.
- Step-by-step deployment (provisioning, Oracle install, SnapMirror setup) is deferred to a companion post ("Deploy Oracle Database step by step on Amazon EVS with FSx for ONTAP").
Source¶
- Original: https://aws.amazon.com/blogs/architecture/architect-highly-available-oracle-database-on-amazon-evs-and-fsx-for-ontap/
- Raw markdown:
raw/aws/2026-10-02-architect-highly-available-oracle-database-on-amazon-evs-and-1ab9dab7.md
Related¶
- systems/amazon-evs — the VCF-on-bare-metal substrate this Oracle architecture runs on.
- systems/amazon-fsx-for-netapp-ontap — the SnapMirror/NFS-datastore storage tier.
- systems/oracle-database — the migrated workload; DR-licensing constraint shapes the mount topology.
- patterns/async-replication-for-cross-region — SnapMirror as the cross-region async mechanism.
- patterns/pilot-light-deployment — pre-provisioned standby DR cluster with replicated-but-unmounted volumes.
- concepts/rpo-rto — SnapMirror frequency sets RPO; standby VMs reduce RTO.
- concepts/storage-media-tiering — vSAN (transient) vs FSx ONTAP (persistent/replicated) split.