SIOS SANless clusters

SIOS SANless clusters High-availability Machine Learning monitoring

  • Home
  • Products
    • SIOS DataKeeper for Windows
    • SIOS Protection Suite for Linux
  • News and Events
  • Clustering Simplified
  • Success Stories
  • Contact Us
  • English
  • 中文 (中国)
  • 中文 (台灣)
  • 한국어
  • Bahasa Indonesia
  • ไทย

Surviving the Friday Night Crash: From Scrappy Bare Metal to Seamless Data Replication

September 27, 2026 by Jason Aw Leave a Comment

Surviving the Friday Night Crash From Scrappy Bare Metal to Seamless Data Replication

Surviving the Friday Night Crash: From Scrappy Bare Metal to Seamless Data Replication

It was my sophomore year of high school. It was late on a Friday night, and like most teenagers, I was zoned out in front of the TV. Then, the glow of the screen was interrupted by my phone lighting up. A buzz. Then another. Then a relentless, vibrating chorus that could only mean one thing: Discord was blowing up, the support tickets were flooding in, and the servers were down. Again.

When Server Downtime Became Part of the Business

At that point, catastrophic downtime had practically become a feature of my daily life. Game hosting is the Wild West of infrastructure. Between targeted DDoS attacks from rival servers, inexperienced customers accidentally nuking their own configurations, and the inevitable hardware failures of budget bare-metal rigs, chaos was just part of the business model. While my classmates were worrying about biology tests, I was frantically SSH-ing into servers at 2 AM, praying a hard reboot would bring the nodes back online.

Eventually, that constant chaos started taking a heavy toll. Personally, the sheer burnout of being a one-man, 24/7 incident response team was exhausting. But from a business perspective, the bleeding was even worse. Every minute of downtime meant paying out SLA credits from a razor-thin margin. It meant hemorrhaging organic growth, losing prospective leads, and dealing with justifiably frustrated gamers. I was learning the hard way that downtime isn’t just a technical glitch; it is a massive financial liability.

Searching for a Better High Availability Strategy

I knew something had to change. I spent hours reading up on enterprise architecture, but as a solo teenager running a bootstrapped operation, my high availability strategy was essentially a collection of duct-tape bash scripts. I found ways to make rudimentary failovers work, but I could never get a reliable, affordable way to ensure true data replication.

I eventually handed over the keys, selling the hosting business to a competitor just before trading my server racks for college lecture halls. But as I progressed through my degree, the operational scars of those frantic outages never faded. I couldn’t stop thinking about the infrastructure, the downtime, and the massive tech giants that effortlessly survived the exact failures that used to cripple my network. I worked relentlessly to understand that gap, which ultimately drove me to land my internship here at SIOS.

Discovering SIOS DataKeeper and Seamless Data Replication

Walking into this engineering environment felt like stepping into an alternate reality where my teenage infrastructure nightmares had already been solved. When I first saw SIOS DataKeeper in action, it was a genuine moment of awe. It was exactly what I had been searching for all those years ago.

Watching edits on one machine move to another seamlessly, and then just work exactly as intended during a simulated failover, was incredible. Seeing how effortlessly it integrates with everything else to make downtime a thing of the past proved to me that the perfect infrastructure I used to dream about actually exists.

From Scrappy Infrastructure to Mission-Critical High Availability

Today, my role bridges the gap between my scrappy roots and elite industry standards. I am directly on the front lines in tech support, ensuring our software and our customers’ architectures run flawlessly. Every single day, I get to help businesses keep their mission-critical data online.

There is an incredibly profound, full-circle satisfaction I feel in that. When I help a customer smooth out their high availability setup, I think back to that solo teenager scrambling in the dark on a Friday night. It is an absolutely amazing feeling to know I am no longer just wishing for a safety net. I am actively helping other businesses make rock-solid, panic-free infrastructure a reality.

Author: Brit Weinstein, Customer Experience Engineer Intern at SIOS

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: data replication, High Availability

Eliminating Single Points of Failure

June 14, 2026 by Jason Aw Leave a Comment

Eliminating Single Points of Failure

Eliminating Single Points of Failure

In the world of enterprise IT, the phrase “Single Point of Failure” (SPOF) is enough to keep any system administrator awake at night. A SPOF is any component in your infrastructure—be it a server, a network switch, or a storage array—that, if it fails, brings the entire system down with it. As businesses increasingly demand 99.99% (or higher) uptime, identifying and eliminating these vulnerabilities is no longer optional; it’s a critical requirement.

If you are looking to bulletproof your infrastructure, combining High Availability (HA) with data replication provides a robust, enterprise-grade solution to eliminate SPOFs and ensure continuous operations.

The Power of Clustering to Eliminate SPOFs

At the heart of high availability is the clustering concept. A cluster is a group of independent servers (nodes) configured to work together to provide highly reliable services. These services could be anything from a custom application to a file share.

In a typical HA cluster, one node actively hosts the services while one or more nodes remain on standby. Cluster management software, such as SIOS LifeKeeper, continuously monitors the health of the active node to ensure it can properly host the services.

If a critical failure is detected on the primary node, the cluster software automatically orchestrates a failover. It shifts the application services, IP addresses, storage, and dependencies to a healthy standby node. By automating this process, the individual server ceases to be a single point of failure, ensuring service continuity with minimal interruption.

Eliminating the SAN Single Point of Failure

Traditional clustering typically depends on a Storage Area Network (SAN) to provide shared access to data across all nodes. However, this design presents a critical vulnerability: the SAN becomes a Single Point of Failure. If the shared storage array experiences downtime, the entire cluster is rendered inoperative, even if the individual nodes remain functional.

To eliminate the shared storage SPOF, administrators utilize data replication to create a “SANless” cluster. Instead of a SAN, each node relies on its own local attached storage. Software like SIOS DataKeeper sits at the operating system level and performs continuous, block-level replication from the active node’s storage to the standby node’s storage.

Because the data is continuously replicated and mirrored in real-time, the standby node is always ready to take over with the latest data on its local storage.

Multiple Communication Paths and Quorum/Witness Solutions

For a cluster to operate safely, the nodes must be in constant communication to verify each other’s status. They do this by exchanging “heartbeats”—small, frequent data packets that indicate a node is alive and healthy.

If a standby node stops receiving heartbeats, it might assume the primary node is dead and attempt to bring the application online. If the primary node is actually still running, you end up with two nodes trying to write data simultaneously—a scenario known as “split-brain.“ To avoid this, you should always configure a quorum or witness solution to your cluster, which acts as a tiebreaker to determine which node should safely own the active workload.

Furthermore, to prevent network infrastructure from becoming a SPOF, a resilient cluster architecture requires multiple communication paths. By ensuring there are multiple distinct ways for nodes to communicate, you ensure that a single faulty network switch or severed cable doesn’t break the cluster’s logic.

Systematically Find & Eliminate SPOFs with SIOS

Building a truly highly available environment means looking at your architecture through the lens of worst-case scenarios. By combining the intelligent application monitoring of SIOS LifeKeeper with the robust, SANless replication of SIOS DataKeeper, you can systematically find and eliminate Single Points of Failure.

Author: Trey Isaac, Sr. Product Support Engineer at SIOS

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: data replication, High Availability

Reliable Data Replication with SIOS DataKeeper: Why Communication (and Ports) Matter

July 22, 2025 by Jason Aw Leave a Comment

Reliable Data Replication with SIOS DataKeeper Why Communication (and Ports) Matter

Reliable Data Replication with SIOS DataKeeper: Why Communication (and Ports) Matter

In nearly every aspect of IT, communication is key, and when it comes to data replication, it’s critical. For DataKeeper, ensuring that data stays synchronized between nodes in a high availability cluster starts with ensuring the systems can talk to each other over the network.

Whether you’re replicating data across regions or data centers, your first task is to enable secure, reliable communication between all participating nodes. At the heart of this communication lies the TCP/IP protocol. DataKeeper uses a predefined set of TCP ports to establish and maintain replication.

What Are TCP Ports, and Why Are They Important for Data Replication?

A TCP port is a numeric identifier that serves as an endpoint used by network protocols to route traffic to specific applications running on a system. Think of them as addresses for an apartment building. You can send a message to the apartment building, but it probably won’t make it out of the lobby. With the address, you can ensure your message gets to the resident. If the desired address (port) is blocked, the data never gets to where it needs to go.

In the context of DataKeeper, these ports serve as the designated pathways through which nodes exchange critical replication data. Without open and correctly routed ports, the nodes won’t be able to communicate, causing replication to fail or stall.

Which Ports Does SIOS DataKeeper Use?

To establish replication and maintain communication between nodes, DataKeeper requires the following TCP ports to be open:

  • 137, 138, 139, 445 – These are Windows networking ports used for file and printer sharing (NetBIOS and SMB).
  • 9999 – This is the default port used by the DataKeeper service for control and status updates.
  • 10000–10025 – These ports are used for the actual replication traffic. Each port in this range corresponds to a drive letter:

10000 = Volume A

…

10025 = Volume Z

If you’re replicating volume F, for example, you’ll need to ensure that port 10005 is open between nodes.

What to Check When Data Replication Isn’t Working

If replication isn’t starting or is repeatedly disconnecting, consider the following:

  1. Firewall Configuration
    1. Check that Windows Firewall is not blocking any required ports. You can create an inbound rule to allow traffic on the needed ports:
      1. Open Windows Defender Firewall with Advanced Security
      2. Go to Inbound Rules > New Rule
      3. Choose Port, select TCP, and specify:

137, 138, 139, 445, 9999, 10000-10025

  1. Allow the connection and apply the rule to all profiles (Domain, Private, Public).
  1. Network Security Groups / Cloud Firewalls

If your nodes are hosted in cloud environments like AWS, Azure, or GCP, make sure the security groups or NSGs also allow the above ports between the relevant IP addresses.

  1. Ping and Connectivity Tests
    1. Use ping or Test-NetConnection in PowerShell to verify network reachability.
    2. Use telnet or Test-NetConnection -Port to check if specific ports are open.

Best Practices for a Smooth SIOS DataKeeper Deployment

Beyond enabling TCP traffic, there are a couple of other networking best practices that can improve your DataKeeper experience. To ensure reliable replication with DataKeeper, start by verifying that all nodes can resolve each other’s hostnames consistently. This can either be through DNS or static entries in the hosts file. Name resolution issues can be a common source of silent failures and should be addressed early. Additionally, think about configuring a dedicated network interface for replication traffic whenever possible. Separating replication from production traffic not only improves performance and reduces latency but can also enhance security and reliability by isolating data transfer from user and application activity.

Ensure Port Connectivity for Reliable SIOS DataKeeper ReplicationIn Summary

For DataKeeper to perform reliably, network communication must be unrestricted across the defined set of TCP ports. Understanding and configuring these ports, especially the volume-specific replication ports, is essential for avoiding downtime and ensuring your high-availability setup delivers on its promise.

Taking a few minutes to audit your firewall rules and confirm connectivity can save you hours of troubleshooting when replication suddenly stalls. As with all things in IT, clear communication, both between people and between systems, makes all the difference.

Want to take the next step? Consider how high-availability strategies, such as clustering, can support safer, disruption-free patching in your environment. Request a demo today to see how SIOS can help you protect critical workloads, minimize downtime, and ensure seamless patching.

Author: Tristan Allen, Associate Customer Experience Software Engineer at SIOS Technology Corp.

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: data replication

How does Data Replication between Nodes Work?

June 19, 2022 by Jason Aw Leave a Comment

How does Data Replication between Nodes Work?

How does Data Replication between Nodes Work?

In the traditional datacenter scenario, data is commonly stored on a storage area network (SAN). The cloud environment doesn’t typically support shared storage.

SIOS DataKeeper presents ‘shared’ storage using replication technology to create a copy of the currently active data. It creates a NetRAID device that works as a RAID1 device (data mirrored across devices).

Data changes are replicated from the Mirror Source (disk device on the active node – Node A in the diagram below) to the Mirror Target (disk device on the standby node – Node B in the diagram below).

In order to guarantee consistency of data across both devices, only the active node has write access to the replicated device (/datakeeper mount point in the example below). Access to the replicated device (the /datakeeper mount point) is not allowed while it is a Mirror Target (i.e., on the standby node).

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: data replication

Data Replication

December 13, 2021 by Jason Aw Leave a Comment

Data Replication

 

 

Data Replication

Real-Time Data Replication for High Availability

What is Data Replication

Data replication is the process by which data residing on a physical/virtual server(s) or cloud instance (primary instance) is continuously replicated or copied to a secondary server(s) or cloud instance (standby instance). Organizations replicate data to support high availability, backup, and/or disaster recovery.  Depending on the location of the secondary instance, data is either synchronously or asynchronously replicated. How the data is replicated impacts Recovery Time Objectives (RTOs) and Recovery Point Objectives (RPO).

For example, if you need to recover from a system failure, your standby instance should be on your local area network (LAN). For critical database applications, you can then replicate data synchronously from the primary instance across the LAN to the secondary instance. This makes your standby instance “hot” and in sync with your active instance, so it is ready to take over immediately in the event of a failure. This is referred to as high availability (HA).

In the event of a disaster, you want to be sure that your secondary instance is not co-located with your primary instance. This means you want your secondary instance in a geographic site away from the primary instance or in a cloud instance connected via a WAN. To avoid negatively impacting throughput performance, data replication on a WAN is asynchronous. This means that updates to standby instances will lag updates made to the active instance, resulting in a delay during the recovery process.

Why Replicate Data to the Cloud?

There are five reasons why you want to replicate your data to the cloud.

  1. As we discussed above, cloud replication keeps your data offsite and away from the company’s site. While a major disaster, such as a fire, flood, storm, etc., can devastate your primary instance, your secondary instance is safe in the cloud and can be used to recover the data and applications impacted by the disaster.
  2. Cloud replication is less expensive than replicating data to your own data center. You can eliminate the costs associated with maintaining a secondary data center, including the hardware, maintenance, and support costs.
  3. For smaller businesses, replicating data to the cloud can be more secure especially if you do not have security expertise on staff. Both the physical and network security provided by cloud providers is unmatched.
  4. Replicating data to the cloud provides on-demand scalability. As your business grows or contracts, you do not need to invest in additional hardware to support your secondary instance or have that hardware sit idle if business slows down. You also have no long-term contracts.
  5. When replicating data to the cloud, you have many geographic choices, including having a cloud instance in the next city, across the country, or in another country as your business dictates.

Why Replicate Data Between Cloud Instances?

While cloud providers take every precaution to ensure 100 percent up-time, it is possible for individual cloud servers to fail as a result of physical damage to the hardware and software glitches – all the same reasons why on-premises hardware would fail. For this reason, organizations that run their mission-critical applications in the cloud should replicate their cloud data to support high availability and disaster recovery. You can replicate data between availability zones in a single region, between regions in the cloud, between different cloud platforms, to on-premise systems, or any hybrid combination.

SIOS Real-Time Data Replication for High Availability and Disaster Recovery

SIOS Datakeeper™ uses efficient, block-level, data replication to keep your primary and secondary instances synchronized. If a failover happens, the secondary instance(s) continues to operate, providing users with access to the most recent data. With SIOS solutions, RPO is always zero and RTO is dependent on the application but typically 30 seconds to a few minutes.

SIOS products uniquely protect any Windows- or Linux-based application operating in physical, virtual, cloud or hybrid cloud environments and in any combination of site or disaster recovery scenarios, enabling high availability and disaster recovery for applications such as SAP and databases, including Oracle, HANA, MaxDB, SQL Server, DB2, and many others. The “out-of-the-box” simplicity, configuration flexibility, reliability, performance, and cost-effectiveness of SIOS products set them apart from other clustering software.

In a Windows environment, SIOS DataKeeper Cluster Edition seamlessly integrates with and extends Windows Server Failover Clustering (WSFC) by providing a performance-optimized, host-based data replication mechanism. While WSFC manages the software cluster, SIOS performs the data replication to enable disaster protection and ensure zero data loss in cases where shared storage clusters are impossible or impractical, such as in cloud, virtual, and high-performance storage environments.

In a Linux environment, SIOS LifeKeeper and SIOS DataKeeper provide a tightly integrated combination of high availability failover clustering, continuous application monitoring, data replication, and configurable recovery policies, protecting your business-critical applications from downtime and disasters.

———————————————————————————————————————————

Here is a real-world example of how one leading manufacturing company uses SIOS to create a high availability solution in the cloud using real-time data replication.

How to Achieve HA in a Cloud Environment with Real-Time Data Replication

Bonfiglioli is a leading Italian design, manufacturing, and distribution company, specializing in industrial automation, mobile machinery, and wind energy products and employing over 3,600 employees in locations around the globe. To run its business, the company relies on various mission-critical applications, including its SAP ERP system. The company’s IT infrastructure includes an on-premises VMware data center and a remote data center for business continuity and disaster protection. Since most of their applications run in a Windows environment, Bonfiglioli used guest-level Windows Server failover clustering in their VMware environment to provide high availability and disaster protection.

The company’s IT team implemented a program to move part of its IT operations into the Microsoft Azure cloud and to leverage Azure as their disaster recovery site. An important requirement of the company’s migration plan was to ensure the cloud architecture could provide better high availability protection than before and ensure Bonfiglioli could continue to meet its strict Service Level Agreements (SLAs).

In its on-premises environment, the company uses VMware clustering, which allows Windows Server Failover Clustering (WSFC) to manage failover to a secondary server in the event of an infrastructure failure. However, it was a challenge to provide this type of protection in the cloud because using guest-clustering with shared-bus disks is not a viable cloud solution. Creating a cluster in VMware using Raw Device Mapping and shared-bus disks (RDM) is challenging and creates limitations for backing up the virtual machines.

The Solution

After evaluating several solutions, Bonfiglioli chose SIOS DataKeeper as their cloud high availability and disaster recovery solution upon learning that SIOS DataKeeper is the only certified high availability clustering solution for SAP in a public cloud. In addition, Bonfiglioli’s management consulting partner, BGP, had experience with SIOS DataKeeper and knew that it is easy to install, transparent to the operating system, and a proven, highly effective solution.

With SIOS, the IT team fashioned a cluster environment without RDM. They created a two-node cluster in VMware and added SIOS DataKeeper Cluster Edition to synchronize storage via real-time data replication in each cluster instance. In an on-premises environment, synchronized storage appears to WSFC as a single shared storage disk.

SIOS DataKeeper also provides high availability protection for the company’s SAP instance and eliminates single point of failure. Using SIOS DataKeeper, the IT team replicated an SSD-tiered disk partition in the company’s on-premises data center using real-time data replication. This allows Bonfiglioli to restore their virtual machines to Microsoft Azure in the event of a disaster.

The Results

Daniele Bovina, Systems Architect at Bonfiglioli, comments about the results, “SIOS DataKeeper gave us an easy way to move our business-critical SAP system to the Microsoft Azure cloud while meeting our stringent SLAs for availability, disaster recovery, and performance.”

—————————————————————————————————————————–

For more information about SIOS Clustering Solutions, contact us or request a free trial.

References

  • https://storageservers.wordpress.com/2018/02/12/difference-between-backup-and-replication-2/
  • http://www.bbc.co.uk/newsbeat/article/16838342/could-the-digital-cloud-used-for-storage-ever-crash

Reproduced from SIOS

Filed Under: Clustering Simplified Tagged With: data replication, High Availability

  • 1
  • 2
  • 3
  • 4
  • Next Page »

Recent Posts

  • What Is High Availability (HA)?
  • Surviving the Friday Night Crash: From Scrappy Bare Metal to Seamless Data Replication
  • Grounded: What Missing Percona Live Amsterdam Taught Me About HA
  • SIOS LifeKeeper vs. Red Hat High Availability Add-On:
  • The State of Application Resilience: 2026 SIOS High Availability Survey

Most Popular Posts

Maximise replication performance for Linux Clustering with Fusion-io
Failover Clustering with VMware High Availability
create A 2-Node MySQL Cluster Without Shared Storage
create A 2-Node MySQL Cluster Without Shared Storage
SAP for High Availability Solutions For Linux
Bandwidth To Support Real-Time Replication
The Availability Equation – High Availability Solutions.jpg
Choosing Platforms To Replicate Data - Host-Based Or Storage-Based?
Guide To Connect To An iSCSI Target Using Open-iSCSI Initiator Software
Best Practices to Eliminate SPoF In Cluster Architecture
Step-By-Step How To Configure A Linux Failover Cluster In Microsoft Azure IaaS Without Shared Storage azure sanless
Take Action Before SQL Server 20082008 R2 Support Expires
How To Cluster MaxDB On Windows In The Cloud

Join Our Mailing List

Copyright © 2026 · Enterprise Pro Theme on Genesis Framework · WordPress · Log in