Chaos Engineering for Multi-Cloud Resilience
Manuscript Title
Chaos Engineering for Multi-Cloud Resilience
A Hypothesis-Driven Method for Validating Provider Failover, Partition Tolerance, and Recovery Objectives
Siri Chandana Chava
Sr. Cloud Architect
Independent Researcher, Tennessee, USA
Abstract—Resilience in a multi-cloud system is not something you can establish by looking at an architecture diagram. A diagram can assert that a workload survives the loss of a region or a whole provider, but only a controlled experiment can show whether that assertion holds when the failure actually arrives. This paper presents a disciplined, hypothesis-driven method for validating multi-cloud resilience through chaos engineering. It states the experimental protocol plainly, defines the three failure modes that make multi-cloud systems different from single-provider ones, and shows how faults are specified as bounded, version-controlled artifacts. To make the method concrete it walks through an illustrative provider-failover exercise. The exercise is a worked example rather than a report of measurements from a particular production system: it describes the shape of such an exercise and the class of finding it tends to surface, namely that a system can meet its availability objective while quietly breaching its recovery-point objective because data replicated across providers had not caught up at the moment of failover. The paper closes with the operational and ethical conditions the practice depends on. Its argument is simple: resilience that has not been exercised should be treated as unverified, and unverified resilience is an open risk.
Index Terms—chaos engineering, multi-cloud, resilience, fault injection, failover, partition tolerance, recovery point objective, disaster recovery.