When a production line stops, the cost does not wait for the IT team to finish troubleshooting. The Siemens True Cost of Downtime 2024 report found that the average manufacturing facility loses USD 260,000 per hour (approximately RM1.22 million) when a critical line sits idle. For Malaysian manufacturers in electronics, F&B, automotive, or palm oil processing, even a fraction of that figure leaves a noticeable impact on quarterly profits and losses.
The problem is that most DR plans in manufacturing were written for IT systems like ERP, email, and file servers, but they do not account for the SCADA systems, MES platforms, and PLCs that actually run the production line. A plan that restores the database but leaves the assembly line offline recovers records of lost production, but not production.
Below, we cover what a complete manufacturing business continuity plan in Malaysia looks like, from cost drivers through architecture to multi-site rollout, for organisations exploring cloud disaster recovery solutions in Malaysia.
The Four Cost Drivers of Factory Downtime
- Production loss. This is the direct value of goods not produced. The cross-sector average is USD 260,000 per hour (approximately RM1.22 million), but continuous-process plants in chemicals or F&B run higher because restarts involve recalibration, purging, and quality revalidation.
- Scrap and rework. Interrupted mid-cycle processes produce waste. In food manufacturing, temperature-controlled batches that stop mid-run cannot be reintroduced. This also applies to precision engineering, where a partial machining cycle may scrap the workpiece.
- SLA penalties and customer churn. Just-in-time supply chains leave no margin for missed windows. An automotive parts supplier that misses a shipment can trigger cascading delays worth multiples of the original order value.
- Recovery overtime. Most facilities require 1.5 to 2 hours of overtime per hour of downtime to catch up. At 1.5x labour rates, recovery costs can match the original production loss.
DR for OT/IT Convergence: MES, SCADA, and ERP
Modern factories run two interconnected technology layers. IT handles ERP, databases, and business applications. OT (Operational Technology) handles MES, SCADA, PLCs, and industrial IoT sensors that control the physical production process.
Effective factory downtime recovery for converged environments requires coverage of both layers:
- ERP and MES replication. Both need synchronised backup. Restoring ERP without MES leaves the business with financial records but no production scheduling. Restoring MES without ERP means producing against stale orders.
- SCADA configuration backup. SCADA and PLC configurations must be backed up with version control. A six-month-old configuration may not match the current production setup after recipe changes or line modifications.
- Network segmentation in recovery. OT networks stay segmented from IT during recovery, as during normal operations. In Europe, 80% of manufacturers operate critical OT systems with known vulnerabilities, making segmentation during recovery as important as during production.
- Tested failover, including OT. If the quarterly drill only covers ERP restoration, it validates half the recovery plan.
RTO and RPO Benchmarks for Production Lines
RTO (Recovery Time Objective) defines how quickly the system must be back. RPO (Recovery Point Objective) defines how much data loss is acceptable. Both must be set per system:
- MES: RTO 1 to 2 hours, RPO 15 to 30 minutes. Losing 30 minutes of MES data during a multi-stage process means the line restarts with incomplete batch records.
- SCADA and PLCs: RTO 30 minutes to 1 hour, and RPO is configuration-based. The priority is restoring the current configuration and re-establishing device communication.
- ERP: RTO 2 to 4 hours, RPO 1 hour. A 4-hour ERP outage is disruptive but survivable. Losing more than one hour of transactional data creates reconciliation problems.
- Email and file servers: RTO 4 to 8 hours, RPO 4 hours. Important for communication, not production-critical.
According to DataNumen’s 2024 analysis, 93% of businesses that experience prolonged data loss file for bankruptcy within one year. For a manufacturer, “prolonged” does not mean weeks. A production line that cannot recover its MES and SCADA configuration within the shift window has already lost the output.
DRaaS for Plants With Limited Connectivity
Cloud-based DRaaS (Disaster Recovery as a Service) is standard for IT workloads, but factory locations do not always have reliable high-bandwidth internet.
With that, the architecture that works uses a hybrid model:
- Local replication appliance. Captures snapshots of ERP, MES, and SCADA configurations at defined intervals (every 15 to 60 minutes). Replication runs at LAN speed with no internet dependency at the point of capture.
- Asynchronous cloud sync. The appliance syncs to the cloud during off-peak hours, transmitting only changed data blocks. If connectivity drops, the appliance queues the sync and continues local capture.
- Cloud failover for IT. In a site-level disaster, ERP and MES workloads spin up on cloud infrastructure while the physical site recovers.
- Local failover for OT. SCADA and PLCs require real-time communication with physical equipment and cannot run in the cloud. The local appliance restores configurations to standby hardware on-site.
This architecture ensures that disaster recovery for manufacturing in Malaysia does not depend on connectivity that many factory locations cannot reliably provide.
Phasing a DR Rollout Across Multiple Sites
Deploying DR across all sites simultaneously is expensive and disruptive. Here’s how a three-phase approach for multi-site manufacturing business continuity planning in Malaysia might go:
- Phase 1 (Months 1 to 3): Begins with the pilot site with the highest production value or greatest downtime risk. Deploy local appliance, configure cloud sync, document RTO/RPO per system tier, and run the first failover drill. The output is a validated runbook that can be replicated.
- Phase 2 (Months 4 to 8): Roll the architecture to two or three high-priority sites, adapting for each site’s OT stack.
- Phase 3 (Months 9 to 12+): Complete rollout to remaining facilities. Establish quarterly DR drills rotating across sites, with each site tested at least twice per year.
Manufacturing was the most targeted sector for ransomware globally in 2025, with attacks rising 56%. In Malaysia, ransomware surged 153% in 2024. All sites need a proper DR plan to prevent instances like ransomware encrypting the production database and the backup server in the same attack.
Getting Factory DR Right From the First Site
Factory disaster recovery sits at the intersection of IT and OT, cloud and local infrastructure, and continuity economics. Once the runbook is validated at the pilot site, scaling across additional sites becomes an operational rollout.
If you are a Malaysian manufacturer evaluating disaster recovery for production environments, dealing with converged OT/IT systems, or looking to phase a DR rollout across multiple factory sites, the conversation starts with mapping the systems that actually run the production line.
This is where Net Onboard is here to help. Our AmplifyContinuity pillar provides the Disaster Recovery as a Service and business continuity framework built for environments where downtime is measured in lost production.
Talk to the Net Onboard team today about cloud disaster recovery solutions in Malaysia and get a DR plan scoped for your production-critical systems.
