Recovering Resiliency Manager
Following are the conditions in which Resiliency Manager may go offline, along with its impact, and steps to recover from such scenario. The table also describes the mitigation steps for each of the condition mentioned.
Table: Recovering Resiliency Manager
Scenario | Impact | Steps to recover | Mitigation steps |
|---|---|---|---|
Single Resiliency Manager deployed in cloud and Resiliency Manager goes offline due to public cloud outage. |
| Wait till the cloud provider fixes the cloud outage. There is no action required from Resiliency Platform side. Business is resumed as usual after Resiliency Manager is brought online. | You can deploy multiple ( a minimum of three) Resiliency Managers in a data center to achieve resiliency of Resiliency Manager. If your cloud data center is in AWS, It is recommended that you deploy multiple Resiliency Managers in different Availability Zones to achieve maximum resiliency of Resiliency Manager. |
One Resiliency Manager deployed per site and any one of the Resiliency Managers goes offline. | Resiliency of Resiliency Managers gets impacted, no impact on product functionality | Bring the affected Resiliency Manager online. Use this recovered Resiliency Manager to control the assets. If Resiliency Manager is in irrecoverable state then perform the Leave domain operation from the other Resiliency Manager in the domain. Deploy a new Resiliency Manager on the impacted site. | Ensure appropriate redundancy of Resiliency Manager against the Hypervisor, storage, and network level failure. You can deploy multiple ( a minimum of three) Resiliency Managers in a data center to achieve resiliency of Resiliency Manager. |
Rare scenario: For target Resiliency Manager, services did not start after power on. | Resiliency Manager (say RM2) has issues because of "/var/opt" primary RM (say RM1) services did not start after power on. The primary RM (RM1) database was moved into rebuild mode but the rebuild failed with an exception because Resiliency Manager (RM2) was not in good responsive state. | For recovery, either first the Resiliency Manager (RM2) needs to be corrected and if not possible then power down (RM2) to get into outage scenario. Perform manual restart of services on primary RM (RM1). | NA |
See Planning a resiliency domain for efficiency and fault tolerance.