search

LEMON BLOG

Multicloud Looks Efficient on Paper — Until Disaster Recovery Is Put to the Test

Multicloud has become an attractive strategy for organisations that want flexibility, redundancy and less dependence on a single provider. On paper, spreading workloads across multiple cloud environments appears to reduce concentration risk and give businesses more options when something goes wrong.

The reality becomes much more complicated once an actual outage begins.

When services are distributed across different cloud platforms, recovery is no longer simply about restoring a virtual machine or switching to a backup environment. Teams have to identify where the failure started, understand which systems are affected, coordinate across providers, move data between environments and restore services quickly enough to remain within regulatory and business tolerances.

For Malaysian financial institutions, that clock is particularly unforgiving.

In 2024, Bank Negara Malaysia imposed penalties exceeding RM5 million on CIMB Bank and Maybank in connection with breaches involving service availability and operational resilience. Maybank experienced repeated disruptions affecting its mobile banking services and MAE application, while CIMB faced penalties associated with shortcomings in incident response and recovery.

The message was clear: resilience is not simply an infrastructure goal. For regulated businesses, it is an operational and regulatory obligation.

Downtime Is No Longer Just an IT Problem

BNM's Risk Management in Technology requirements set strict expectations around service availability.

For critical user-facing services, cumulative unplanned downtime must remain within tightly controlled limits, while individual incidents are also expected to stay below specified maximum tolerances.

From an IT perspective, that changes the way disaster recovery needs to be viewed.

An organisation cannot simply say that systems eventually came back online. The more important questions are:

A recovery process that takes several hours may still look technically successful to an infrastructure team, but from a regulatory perspective, those hours can consume a significant portion of the organisation's annual availability allowance.

That makes recovery speed measurable in business risk rather than just technical performance.

Multicloud Can Create a False Sense of Resilience

One of the most common assumptions behind multicloud architecture is that using more than one provider automatically makes the environment more resilient.

It can—but only if the environments are designed to work together during failure.

Simply placing workloads in AWS, Microsoft Azure and another cloud provider does not automatically create effective disaster recovery.

In fact, poorly planned multicloud environments can make recovery harder.

Each provider may have its own:

Under normal operating conditions, these differences are manageable.

During an outage, they become friction.

If an application suddenly becomes unavailable, teams may first need to determine whether the problem originates from the application itself, a cloud provider, a network path, a security policy, an identity service or one of the integrations between environments.

That investigation consumes something every regulated institution has very little of during an outage: time.

Failures Often Happen Between Clouds, Not Inside Them

Individual cloud platforms typically provide mature disaster recovery capabilities.

The more difficult part is what happens when recovery depends on moving data or services between providers.

Replication traffic between cloud environments may need to cross external networks. Backup data may have to move from one provider to another. Applications may rely on dependencies hosted somewhere else entirely.

The weakest point may therefore not be AWS or Azure themselves.

It may be the connection between them.

When recovery traffic depends heavily on public internet routing, organisations have less control over factors such as latency, congestion, routing changes and available throughput.

That uncertainty matters when large volumes of data need to be replicated or restored quickly.

A recovery plan can be perfectly designed at the application layer and still miss its recovery objective because the underlying network cannot move the data quickly enough.

Recovery Time Starts Before the Restore Button Is Pressed

Organisations also tend to think of recovery time as the period between starting a restore and bringing the application online.

In reality, the clock starts much earlier.

Before recovery can begin, teams may need to:

Every step consumes part of the recovery window.

A technically fast restore process therefore means little if it takes an hour just to understand where the problem originated.

Ransomware Makes Multicloud Recovery Even More Complicated

The situation becomes even more difficult when an outage is caused by a cyberattack rather than a simple infrastructure failure.

With ransomware, restoring services quickly is only one objective.

The organisation must also avoid restoring compromised data or reconnecting infected systems to production.

Security teams may need to determine when the attack began, which systems were affected and whether replicated copies have already been contaminated.

In a fragmented multicloud environment, those answers may exist across several monitoring systems and logging platforms.

The organisation could end up correlating records from multiple cloud providers, security products and networking systems before it can confidently begin recovery.

Again, the problem is not necessarily a lack of data.

It is the difficulty of bringing all that data together quickly enough to make a safe decision.

Multicloud Resilience Can Become Expensive Very Quickly

The financial cost of poorly integrated multicloud architecture is another issue.

Businesses often duplicate infrastructure across providers in the name of resilience.

They may maintain secondary environments, duplicate monitoring tools and employ teams with expertise across several cloud platforms.

Individually, each decision may appear reasonable.

Collectively, the organisation can end up paying for:

Multicloud can therefore reduce dependency on a single provider while simultaneously creating a much larger operational footprint.

The key question is whether all of that additional complexity actually improves recovery.

If the organisation still struggles to fail over quickly during an incident, it may be paying for redundancy without receiving meaningful resilience in return.

The Network Layer Is Often the Missing Part of Disaster Recovery

One approach to improving cross-cloud recovery is to make the network connecting those environments more predictable.

This is where solutions such as Time Cloud Xchange Hub attempt to simplify multicloud connectivity.

Instead of relying primarily on public internet routes between cloud environments, the platform provides private Layer 2 and Layer 3 connectivity into a shared multicloud fabric.

Separate connections can then extend from that fabric toward providers such as AWS and Microsoft Azure.

The idea is relatively straightforward.

Cloud platforms continue managing their native replication and backup processes, while the network carrying that traffic becomes more controlled.

Rather than sending recovery data through unpredictable public routes, traffic can move across a dedicated cloud-to-cloud path.

For disaster recovery, that distinction can be significant.

A Predictable Recovery Path Can Save Valuable Time

The advantage of private connectivity is not simply higher performance.

Predictability can be just as important.

If an organisation knows how recovery traffic will travel, it becomes easier to estimate recovery times, test disaster recovery plans and identify network bottlenecks before an actual incident occurs.

The organisation can design around known characteristics rather than hoping sufficient internet bandwidth will be available when disaster strikes.

This becomes particularly important when recovery involves large datasets.

Moving several terabytes between environments may be manageable during routine replication, but emergency restoration places very different demands on the network.

A deterministic path reduces one of the major uncertainties in that process.

Cloud-to-Cloud Recovery Can Reduce Dependence on Physical DR Sites

Traditional disaster recovery strategies often involve maintaining a separate physical recovery site.

A business might operate its primary systems in one data centre and maintain another facility purely for disaster recovery.

That model can provide strong resilience, but it is expensive.

The organisation may need to maintain duplicate servers, storage, network equipment, licences, security systems and physical facilities that spend much of their life waiting for something to fail.

Cloud-based disaster recovery changes the equation.

Instead of maintaining an entire secondary physical environment, organisations can potentially recover workloads into another cloud platform.

When connectivity between those cloud environments is properly engineered, the requirement for dedicated physical DR infrastructure can be reduced.

That does not mean every organisation should abandon traditional disaster recovery sites, but it creates another architectural option.

A Unified Multicloud Fabric Can Simplify Operations

One of the larger benefits of consolidating cloud connectivity under a single networking layer is visibility.

Without such consolidation, an organisation may effectively operate several separate environments.

Network teams monitor one set of connections.

Cloud engineers monitor another.

Security teams maintain their own telemetry.

During an incident, the organisation must reconstruct the full picture from all of those sources.

A common connectivity layer can simplify that architecture.

Instead of thinking about AWS, Azure and another cloud as entirely separate islands, the business can treat them as parts of a larger controlled environment.

That can make troubleshooting easier because network traffic follows more clearly defined paths.

It can also simplify capacity planning, monitoring and change management.

Auditability Matters Almost as Much as Recovery Speed

For regulated organisations, restoring service is only part of the challenge.

After the outage, someone will almost certainly ask what happened.

Management will want answers.

Risk teams will need evidence.

Auditors may need documentation.

Regulators may expect a clear timeline.

A business must therefore be able to reconstruct the incident accurately.

That means knowing when the disruption began, when it was detected, which systems were affected, when recovery started, how traffic moved and when services were fully restored.

If the organisation has to reconcile several independent cloud logs after every incident, producing that timeline becomes difficult.

A more centralised network path can make the investigation easier.

Instead of piecing together several unrelated views of the same incident, teams have a more consistent record of how recovery traffic moved across environments.

For Banks, Every Minute Eventually Needs an Explanation

This is particularly important for financial institutions.

When a banking service becomes unavailable, the issue does not stop with customer inconvenience.

There may be transaction failures, settlement implications, delayed payments and operational impacts across other institutions.

Afterwards, the bank may need to explain exactly how long each service was unavailable and how the recovery process unfolded.

If multiple teams are using different logs with slightly different timestamps or definitions of downtime, producing a definitive answer can become surprisingly difficult.

A strong disaster recovery architecture therefore needs to support two objectives simultaneously

The second objective is easy to overlook until regulators or auditors start asking questions.

Why Sovereign Cloud Is Becoming More Important

Another major consideration in Malaysia's evolving cloud landscape is data sovereignty.

A sovereign cloud model is designed to keep data, processing systems and security controls within a defined jurisdiction while ensuring that the relevant operators remain subject to that jurisdiction's laws.

This has become increasingly important for regulated industries.

Directors and senior management are no longer expected to understand cloud adoption only at a high level.

They increasingly need answers to practical governance questions:

These are no longer purely technical questions.

They are governance questions.

Sovereignty and Resilience Are Closely Connected

Data sovereignty is often discussed primarily in terms of privacy and regulatory compliance, but it also has implications for disaster recovery.

An organisation cannot simply move regulated data anywhere it wants during an emergency.

The recovery environment still needs to satisfy legal, security and governance requirements.

That means the fastest available cloud may not always be an acceptable recovery destination.

A properly designed sovereign multicloud architecture can reduce that uncertainty by ensuring recovery options have already been evaluated against regulatory requirements.

Rather than deciding where data is allowed to move in the middle of an outage, the organisation makes those decisions during architecture planning.

That is far safer.

Visibility Becomes More Valuable as Cloud Complexity Grows

The more cloud platforms an organisation adopts, the harder it becomes to maintain a single operational view.

Each provider naturally reports its own metrics.

Each security platform generates its own alerts.

Each environment has its own logs and configuration information.

Individually, these systems can provide excellent visibility.

Collectively, they can still leave teams struggling to understand what is happening across the entire organisation.

A unified multicloud architecture attempts to reduce that fragmentation.

By bringing connectivity under a common control layer, organisations can create a clearer picture of how workloads communicate and where data moves.

That visibility is valuable during normal operations.

During an incident, it can become critical.

Cloud Resilience Should Be Tested Before the Outage

One of the most dangerous assumptions in disaster recovery is believing that a documented plan automatically equals a working plan.

It does not.

Recovery needs to be tested.

An organisation should know:

The first time an organisation discovers that its multicloud recovery process does not work should not be during a production outage.

Recovery Metrics Need to Reflect Business Reality

Traditional disaster recovery planning often focuses heavily on Recovery Time Objective (RTO) and Recovery Point Objective (RPO).

These remain important, but organisations should also consider the complete operational timeline.

For example, a system may have an RTO of 60 minutes.

Technically, the infrastructure team may be able to restore the application within 45 minutes.

But if it takes 40 minutes to detect the outage and another 30 minutes to authorise failover, the actual customer impact is already far beyond the stated objective.

Recovery planning therefore needs to measure the complete journey:

Optimising only the restore step gives an incomplete picture.

Multicloud Is Not the Same as Resilience

This is perhaps the most important point.

Having multiple clouds does not automatically mean an organisation has high availability.

Redundancy does not automatically create resilience.

You can duplicate infrastructure across three providers and still experience a prolonged outage if the organisation cannot coordinate those environments effectively.

True resilience requires the pieces to work together.

The organisation needs reliable connectivity, clear ownership, tested recovery procedures, consistent monitoring, accurate logging and enough automation to reduce human delays during emergencies.

Multicloud provides the building blocks.

Architecture determines whether those blocks actually form a resilient system.

The Real Test Begins When Everything Goes Wrong

When systems are operating normally, almost any cloud architecture can look impressive.

Dashboards are green.

Replication jobs are running.

Backups are completing.

Performance metrics look healthy.

The real test begins when something fails unexpectedly.

A critical cloud region goes offline.

A ransomware alert appears.

A network dependency stops responding.

A major application becomes unavailable during peak banking hours.

That is when the theoretical advantages of multicloud architecture are tested against operational reality.

The question is no longer how many cloud providers the organisation uses.

It becomes much simpler:

How quickly can the business restore service?

Final Thoughts

Multicloud can provide enormous flexibility, but complexity grows alongside that flexibility.

The more environments an organisation operates, the more important connectivity, visibility and governance become. Without those foundations, multicloud can create exactly the kind of fragmented recovery process that businesses hoped to avoid when they adopted it.

For Malaysian financial institutions and other regulated organisations, this matters because downtime has regulatory consequences as well as technical ones.

The strongest disaster recovery strategies therefore need to do more than maintain backups. They need to provide predictable recovery paths, tested failover procedures, consolidated visibility and clear audit trails.

Private multicloud connectivity and sovereign cloud architectures can play an important role by reducing uncertainty between environments and creating clearer governance over where recovery traffic travels.

Ultimately, the success of a multicloud strategy should not be judged when everything is operating normally.

It should be judged when something breaks.

When the outage clock starts, organisations need to know exactly where the problem is, where their data is going, how quickly services can return and how they will prove what happened afterwards.

That is the difference between simply having multiple clouds and actually having cloud resilience.

Rakyat Digital Opens the Door to AI Tools and Digi...
Why Sans-Serifs Took Over — And What We Lost Along...

Related Posts

 

Comments 0

Loading latest comments...
Sunday, 16 August 2026

Captcha Image

LEMON VIDEO CHANNELS

Step into a world where web design & development, gaming & retro gaming, and guitar covers & shredding collide! Whether you're looking for expert web development insights, nostalgic arcade action, or electrifying guitar solos, this is the place for you. Now also featuring content on TikTok, we’re bringing creativity, music, and tech straight to your screen. Subscribe and join the ride—because the future is bold, fun, and full of possibilities!

My TikTok Video Collection