Recovering a resiliency group using replication-based recovery
Recover is an activity initiated by a user when the source data center is down due to a natural calamity or other disaster, and the virtual machines need to be restored at the target data center to provide business continuity. The user starts the virtual machines at the recovery data center with the available data. Since it is an unplanned event, the data available at the recovery data center may not be up to date. You need to evaluate the tolerable limit of data loss, and accordingly take the necessary action - start the virtual machines with the available data, or first use any other available data backup mechanism to get the latest copy of data, and thereafter start the virtual machines. The recover operation brings up the virtual machines at the target data center using the last available data.
Perform the resync operation after successful completion of recover operation.
If you are recovering to vCloud Director data center, without adding Hyper-V Server or vCenter Server, then recover operation from cloud to production (on-premises) data center is not supported.
It is recommended to stop or disable NetworkManager on RHEL hosts having multiple NICs. This is required if the recovery data center is AWS, Azure cloud.
For VMware virtual machines, ensure that the network mapping of all the required port groups, or subnets across the data centers is complete.
For Hyper-V virtual machines, ensure that the network mapping of all the required virtual switches across the data centers is complete.
See Creating network pairs between source and target data centers.
If the source data center is in AWS, then ensure that the network mapping of all the required subnets between the source and target data center is complete.
There should not be any resources on Azure having the same name or substring of name as that of the virtual machine display name or FQHN name on on-premises data center.
If the source data center is in cloud, then you need to set the SAN policy to either OnlineAll or OfflineShared based on whether you have shared or non-shared disks. For more details refer to Microsoft Documentation.
If the status of the virtual machine on the source data center is not correctly displayed, then you need to refresh the cloud discovery or the virtualization server discovery.
For the replication technology HPE 3PAR Remote Copy, ensure that for VMware virtual machines the config.vpxd.filter.hostRescanFilter value is set to false.
When you upgrade from an earlier version to version 10.4 or later, then after performing the recover operation with replicated data, a risk is raised. This risk is regarding the changes in NIC configuration when you recover to any cloud data center. Suppress this risk while the resiliency group is online on cloud. Migrate back to the on-premises data center and then edit the resiliency group to fix the NIC configuration.
To perform recover operation on virtual machines
- Navigate
Assets (navigation pane) > Resiliency Group(s) tab
- Double-click the resiliency group to view the details page. Click Recover.
- Do the following:
Select the target data center.
If there is an outage on the source data center, select the Confirm outage of assets check box.
During the recover operation, if Resiliency Platform detects a probability of data loss, you have the option to abort the recover operation to avoid any data loss. Select the check box if you want to abort the operation in such a situation.
Click Continue for warnings.
- Click Submit.
To perform recover operation on resiliency group configured with CDP enabled
- Navigate
Assets (navigation pane) > Resiliency Group(s) tab
- Double-click the resiliency group to view the details page. Click Recover.
- To select the recovery point for CDP, do the following:
- On the Select Recovery points panel, you can select the date, time, and range to list the recovery points with the latest data against the assets.
Select the target data center.
If there is an outage on the source data center, select the Confirm outage of assets check box.
Do not select the Abort recover if these subtasks fail check box and click Next.
Select the recovery points. To choose from a specific time range, select the date and the time range. Then select the hours or minutes since the start time. Click Search.
- ClickSubmit.
If the recover operation fails, check to know the reason and fix it. You can then launch the operation. The operation restarts the recover workflow, it skips the steps that were successfully completed and retries those that had failed.
Do not restart the workflow service while any workflow is in running state, otherwise the operation may not work as expected.
For more information on troubleshooting specific scenarios,