Introduction
The client is a globally distributed logistics organization that relies on continuous availability of its infrastructure and business applications to support freight operations, customer services, and regional offices. Its production environment runs on Nutanix hyperconverged infrastructure — the AHV hypervisor with the Nutanix Distributed Storage Fabric — at on-premises sites in Buenos Aires, Hong Kong, and Miami.
- The engagement was driven by the need to improve business continuity and recoverability across the client's critical production environments without the cost and operational overhead of maintaining secondary physical data centers. To establish disaster-recovery capability for each site without standing up a second physical data center, the client engaged Nutanix Xpert Services to design and deploy Nutanix Clusters on AWS as the DR target platform.
Problem Statement (Challenges)
Prior to the engagement, each of the client's three production sites operated on a single physical Nutanix cluster with no secondary site, which motivated the disaster recovery initiative. A single disaster recovery method could not be proposed given the existing workloads: the on-premises estate — spanning SQL databases, file services, Citrix, DMZ/WAF, and backup systems — was varied enough that it needed a mixed protection approach, such as separate protection domains for SQL versus files and category-based Leap policies, rather than one blanket DR policy.
- Without a DR target, the client faced the risk of extended service disruption to freight operations, customer services, and regional offices in the event of a site-level failure, and would otherwise have had to bear the capital investment and operational overhead of building and maintaining dedicated secondary physical disaster recovery data centers. The existing inter-site bandwidth was also not sufficient as-is: additional inter-site bandwidth had to be procured specifically to hit the 1-hour RPO / 4-hour RTO target being planned for the DR replication traffic.
- A VPN appliance had to be stood up before the disaster recovery connectivity work could even start: a Fortinet appliance was a customer-side prerequisite, and the connectivity design could not begin until it existed, adding a procurement and deployment constraint to the timeline.
Solution Proposed
Nutanix Xpert Services proposed a relocate-style migration in which the existing Nutanix/AHV stack — including Nutanix Files, Citrix infrastructure, DMZ/WAF services, and other production and test virtual machines — is extended into AWS bare-metal instances through Nutanix's native AWS integration, rather than re-architected onto native AWS compute or container services. Three independent Nutanix Clusters on AWS were deployed as regional DR targets: a Singapore cluster for the Hong Kong site, a Frankfurt cluster for the Buenos Aires site, and an N. Virginia cluster for the Miami site, onboarded region by region using AWS's Assess, Mobilize, and Migrate & Modernize methodology aligned to the AWS Migration Acceleration Program (MAP).
- For each cluster, the team established a multi-account landing zone and per-region VPC, configured site-to-site IPSec VPN and Route 53 Resolver connectivity back to the corresponding on-premises site, and deployed Nutanix Clusters on AWS via the Nutanix Clusters Console with Prism Central registration and Cluster Redundancy Factor 2 (RF2) configuration.
- To support disaster recovery, the team configured protection domains and Leap replication policies to protect the different workload categories, with compute delivered through AWS bare-metal EC2 instances (i3.metal / i3en.metal) running the AHV hypervisor and Nutanix CVMs on the Nutanix Distributed Storage Fabric. Infrastructure health was monitored through Prism Central and Prism Element alerting, Amazon GuardDuty and AWS Security Hub were integrated for centralized threat detection, and As-Built documentation was produced for each of the three regional deployments.
Services Used
AWS bare-metal EC2 instances (i3.metal / i3en.metal) hosted the AHV hypervisor and Nutanix CVMs, with AWS VPC providing per-region private and public subnets and site-to-site IPSec VPN for connectivity back to each on-premises site. Amazon Route 53 Resolver forwarded DNS queries to on-premises DNS infrastructure, and the Nutanix Clusters Console together with Prism Central and Prism Element supported cluster provisioning and ongoing management.
- These were supported by AWS Secrets Manager for securely storing and managing credentials required for provisioning and operational activities, Amazon GuardDuty and AWS Security Hub for centralized threat detection and security visibility, AWS Config for configuration and compliance monitoring, and AWS Cost Explorer alongside resource tagging for cost governance and allocation by cluster and AWS Region.
Benefits
The engagement delivered three regional DR clusters, each configured and validated with Cluster Redundancy Factor 2 (RF2): three bare-metal nodes in Singapore, five in Frankfurt, and five in N. Virginia, each expandable up to the platform limit of sixteen nodes per cluster without architectural redesign. Tiered recovery objectives were established and validated through DR failover testing — an RPO of ≤ 1 hour and RTO of ≤ 4 hours for Tier 0 SQL/database volume groups, and an RPO of ≤ 4 hours and RTO of ≤ 8 hours for Tier 1 file and application services — with protection-domain replication validated across all three clusters (Singapore: 24 SQL and 24 file-service protection domains; Frankfurt: 27 SQL and 21 file-service protection domains; N. Virginia: 23 SQL and 25 file-service protection domains).
- Beyond the measurable outcomes, the solution eliminated the need for the client to build and maintain dedicated secondary physical DR data centers, reducing capital investment and operational overhead while preserving the existing Nutanix operating model. The consistent operational model across on-premises and AWS environments allowed a single knowledge-transfer session to bring the operations team up to speed rather than retraining staff on a new platform, and the As-Built documentation provided for each regional deployment gives the team a clear reference for administration, operational procedures, and disaster recovery workflows going forward.
