SIOS SANless clusters

SIOS SANless clusters High-availability Machine Learning monitoring

  • Home
  • Products
    • SIOS DataKeeper for Windows
    • SIOS Protection Suite for Linux
  • News and Events
  • Clustering Simplified
  • Success Stories
  • Contact Us
  • English
  • 中文 (中国)
  • 中文 (台灣)
  • 한국어
  • Bahasa Indonesia
  • ไทย

Taking Over a SIOS LifeKeeper for Linux Cluster

June 21, 2026 by Jason Aw Leave a Comment

Taking Over a SIOS LifeKeeper for Linux Cluster

Imagine you’re standing outside of your minivan, baby in hand, reaching for the baby’s diaper bag when a large black van with a red stripe pulls up beside you.  The van door slowly opens, revealing a motley crew of individuals.  An imposing figure exits the vehicle, his mohawk stretches straight to the sky, and a strong aura of no-nonsense looms as large as the gold chains around his neck.  A savvy veteran exits and tells you that you are now an integral part of a mission-critical operation.

Imagine you’re seated front row, center left, at an exclusive, all-access rehearsal of your favorite band.  You listen as they rip through your favorite hits, one after another.  You strum your air guitar to the classics.  In perfect rhythm, you alternately bang out the drum riffs on your seat and legs.  Eventually,y you find two straws from a discarded Big Gulp and use them to kill the drum solo. A lifetime in the making, the wait is over, and you are finally a part of the sold-out crowd, but suddenly the drummer runs off stage, and a frantic manager is pointing at you to get on stage.

Carlos the Great Dane is the odds-on favorite to win best in show for Big Boys Kennel Club.  You’ve only heard about Carlos through coworkers who love big breed dogs, but today you’re seeing a whole new side of Carlos.  In fact, you are seeing every side of Carlos as his trainer just dropped him off at your house while he waits outside for the tow truck to haul away his work van and the rental car to bring a replacement.  “It will only be one hour,” he assures you.  But keeping a million-dollar best in show out of danger is no easy task.

Let’s be honest, chances are the original A-Team or any real version of the same will not pull up beside your minivan and whisk you away to be a part of a top-secret, mission-critical, save-the-nation type of event.  Likewise, your best Jimi Hendrix or Sheila E impersonation, with or without the big gulp straws, will probably not move you from the crowd to the spotlight of a sold-out show.  And while you may see Carlos or another best in show favorite or winner, unless you are the owner, trainer, or judge, you will likely not be spending any unsupervised time with them.  While these scenarios are unlikely to happen, it is possible that you might be asked to take over a LifeKeeper for Linux cluster that comes with similar risks, responsibilities, and criticality.

What to Do When You Inherit a LifeKeeper for Linux Cluster

Many businesses report that downtime costs upwards of $300,000 per hour, to the millions.  Companies seeking to avoid disasters and downtime will frequently deploy robust architectures with redundancy at multiple layers.

In addition to these redundancies, many companies deploy High Availability (HA) software like SIOS LifeKeeper for Linux to add an essential layer of monitoring and recovery capabilities to their business-critical infrastructure, applications, and databases.  LifeKeeper for Linux provides resource monitoring and recovery for infrastructure, applications, services, and databases, ensuring that business continuity is maintained and downtime is minimized or avoided.  While it doesn’t come in a black van manned by operatives, it is mission-critical.

So, what do you do if you are suddenly responsible for being the lead administrator or soloist for a LifeKeeper for Linux (LK-L) environment that brings in more revenue than a couple of dozen versions of Carlos?

8 Steps for Taking Over a SIOS LifeKeeper for Linux Cluster

Eight critical steps for taking over a SIOS LifeKeeper for Linux Cluster include:

1. Locate and review existing runbooks

Find any existing runbooks.  These are often documents stored in a document repository created by the previous administrator.  A detailed runbook often provides insight into the cluster’s configuration and architecture.  These details will be helpful for administration and future operations.

2. Locate your LifeKeeper for Linux product version

Understanding your product version is an important part of taking ownership of the cluster.  SIOS releases frequent product updates that offer more feature-rich content, security updates, and improvements.  When you take over an existing product cluster, you’ll need to know what version you are on so that you can assess several factors:

  1. Where is your product within the product and support lifecycle?
  2. Are you on the latest version of the product?
  3. What new features or fixes have been added to the product since your version?
  4. Where to find version-specific documentation?

You can find the product version via the UI.  If your runbook indicates LifeKeeper versions 9.8.x or newer, you can use https://<servername>:5110 (or https://<server_IP>:5110) to launch the LifeKeeper Web Management Console (LKWMC).  Once logged in, select properties:

LifeKeeper Web Management Console properties menu

After the properties page loads, locate your version information in the section beneath the product name:

LifeKeeper for Linux version information in the Web Management Console

If you’re running an older version of LifeKeeper for Linux, enable X11 forwarding and launch the Java UI via the command /opt/LifeKeeper/bin/lkGUIapp from a SSH client session.

Command to launch the LifeKeeper Java UI

Your product version can also be found via the command line as follows:

# rpm -qi steeleye-lk

Command to check the LifeKeeper for Linux product version

Once you’ve launched the Java UI and logged in, you’ll be able to navigate to help. Once you obtain your product version, check your product lifecycle and version-specific information via docs.us.sios.com

3. Review your technical support agreement

The Technical Support agreement (TSA) outlines the support that SIOS provides for the SIOS products.  The TSA is helpful in identifying critical information regarding maintenance, upgrades, product support, and product fixes.  The TSA also provides valuable information regarding SIOS’s 24/7 support offering and the contact information for access.  Understanding the TSA goes a long way towards ensuring confidence that you are not on an island, but supported by the SIOS team.  Understanding the TSA also helps you avoid surprises during ongoing deployment and maintenance by identifying what is and is not covered.

4. Obtain SIOS Administrator Training

If your transition is immediate, there may not be a previous administrator available to provide you or your new team members with training.  Don’t panic.  SIOS has convenient online training available.  This training provides a comprehensive overview of the LifeKeeper for Linux product and the roles and actions required for an administrator.  If your Account Representative was documented in the runbook, reach out to them directly for more information.  Otherwise, contact sales@us.sios.com or support@us.sios.com for assistance.

5. Create a demo or test cluster

Armed with your administrator training, deploy a test environment where you can practice and hone your skills and understanding without directly endangering your company’s data or applications.  Creating a demo or test cluster helps you and your future team understand the product basics and get familiar with the UI.

In addition, if your team inherited a runbook, building your own cluster via the runbook helps you validate and update these books for the future. If possible, do your best to mimic the applications and data protected by the production cluster. Ideally, the test cluster your team builds should be as similar to the production systems as possible.  This helps your team understand dependencies, behaviors, and operations in a safe environment before executing commands on production.  Be sure to run through several key exercises, such as:

  1. Manual switchovers
  2. Server failovers
  3. Application recovery
  4. Maintenance operations

6. Schedule a cluster health check

A cluster health check validates the entire SIOS LifeKeeper for Linux environment.  Think of it as a multipoint inspection for your HA cluster.  A team of SIOS experts will conduct a detailed review and validation of the system logs, system settings, run books, LifeKeeper operation, and other documentation to ensure the LifeKeeper environment, including application recovery kits, is configured and operating in an optimized fashion.  The health check report will provide you with recommendations for correcting and/or improving operation, de-risking potential issues, and increasing your awareness of the products.

7. Leverage SIOS support and professional services

The A-Team was a team, not just a single individual on a critical mission.  Your favorite band is more than just the lead singer, drummer, or guitarist.  It is a group of like-minded professionals seeking to accomplish great things, make great music, avoid disappointing fans, and enjoy the rewards of sold-out shows.  Carlos’ success includes handlers, groomers, coaches and trainers, walkers, veterinarians, and a bevy of experts and agents.  Their success is a team effort, and so is yours.  Leverage your SIOS support team at support@us.sios.com or via the support portal support.us.sios.com to gain access to invaluable insights and information.

The SIOS Support Portal contains hundreds of helpful knowledge-based articles (KBAs), access to the latest software, and is ready to help engineers guide you towards success. Reach out to the SIOS support team to ensure that you have a login to the support portal and can manage your cluster effectively.  The support team can also help you get access to an array of SIOS Professional Services offerings, including more advanced training, additional health checks and validations, assistance for new cluster installations, or standby engineering services for your first or any future maintenance or go-live windows.

8. Stay connected with SIOS

A secret to success when you inherit a new cluster is to stay in touch with SIOS.  Establish a frequent check-in with your Account Representatives.  This touchpoint enables you to keep up with new options and opportunities, stay ahead of license renewals, and understand how to add new clusters to expand your protection of additional applications and services.

Stay in touch with your SIOS support team via the newsletter and email blasts. The email blasts provide updates on any new features, releases, or critical updates that may impact your software.  Open cases via the support email inbox or Support Portal whenever you need clarity on an RCA or have an issue that just needs a second set of trained eyes.

Taking over a SIOS LifeKeeper for Linux cluster doesn’t have to feel overwhelming. Request a demo to see how SIOS can help you protect critical applications, reduce downtime risk, and manage high availability with confidence.

Author: Cassius Rhue, VP, Customer Experience, SIOS Technology Corp.

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: Linux

Eliminating Single Points of Failure

June 14, 2026 by Jason Aw Leave a Comment

Eliminating Single Points of Failure

Eliminating Single Points of Failure

In the world of enterprise IT, the phrase “Single Point of Failure” (SPOF) is enough to keep any system administrator awake at night. A SPOF is any component in your infrastructure—be it a server, a network switch, or a storage array—that, if it fails, brings the entire system down with it. As businesses increasingly demand 99.99% (or higher) uptime, identifying and eliminating these vulnerabilities is no longer optional; it’s a critical requirement.

If you are looking to bulletproof your infrastructure, combining High Availability (HA) with data replication provides a robust, enterprise-grade solution to eliminate SPOFs and ensure continuous operations.

The Power of Clustering to Eliminate SPOFs

At the heart of high availability is the clustering concept. A cluster is a group of independent servers (nodes) configured to work together to provide highly reliable services. These services could be anything from a custom application to a file share.

In a typical HA cluster, one node actively hosts the services while one or more nodes remain on standby. Cluster management software, such as SIOS LifeKeeper, continuously monitors the health of the active node to ensure it can properly host the services.

If a critical failure is detected on the primary node, the cluster software automatically orchestrates a failover. It shifts the application services, IP addresses, storage, and dependencies to a healthy standby node. By automating this process, the individual server ceases to be a single point of failure, ensuring service continuity with minimal interruption.

Eliminating the SAN Single Point of Failure

Traditional clustering typically depends on a Storage Area Network (SAN) to provide shared access to data across all nodes. However, this design presents a critical vulnerability: the SAN becomes a Single Point of Failure. If the shared storage array experiences downtime, the entire cluster is rendered inoperative, even if the individual nodes remain functional.

To eliminate the shared storage SPOF, administrators utilize data replication to create a “SANless” cluster. Instead of a SAN, each node relies on its own local attached storage. Software like SIOS DataKeeper sits at the operating system level and performs continuous, block-level replication from the active node’s storage to the standby node’s storage.

Because the data is continuously replicated and mirrored in real-time, the standby node is always ready to take over with the latest data on its local storage.

Multiple Communication Paths and Quorum/Witness Solutions

For a cluster to operate safely, the nodes must be in constant communication to verify each other’s status. They do this by exchanging “heartbeats”—small, frequent data packets that indicate a node is alive and healthy.

If a standby node stops receiving heartbeats, it might assume the primary node is dead and attempt to bring the application online. If the primary node is actually still running, you end up with two nodes trying to write data simultaneously—a scenario known as “split-brain.“ To avoid this, you should always configure a quorum or witness solution to your cluster, which acts as a tiebreaker to determine which node should safely own the active workload.

Furthermore, to prevent network infrastructure from becoming a SPOF, a resilient cluster architecture requires multiple communication paths. By ensuring there are multiple distinct ways for nodes to communicate, you ensure that a single faulty network switch or severed cable doesn’t break the cluster’s logic.

Systematically Find & Eliminate SPOFs with SIOS

Building a truly highly available environment means looking at your architecture through the lens of worst-case scenarios. By combining the intelligent application monitoring of SIOS LifeKeeper with the robust, SANless replication of SIOS DataKeeper, you can systematically find and eliminate Single Points of Failure.

Author: Trey Isaac, Sr. Product Support Engineer at SIOS

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: data replication, High Availability

3 Challenges of Maintaining High Availability with a Legacy Infrastructure

June 9, 2026 by Jason Aw Leave a Comment

3 Challenges of Maintaining High Availability with a Legacy Infrastructure

3 Challenges of Maintaining High Availability with a Legacy Infrastructure

High availability (HA) is critical for organizations that rely on continuous access to applications, services, and data. Whether supporting customer-facing platforms or internal business operations, downtime can quickly lead to financial loss, productivity issues, and reputational damage. While many companies continue to use legacy infrastructure due to cost, compatibility, or business requirements, maintaining high availability in older environments becomes increasingly difficult over time. Legacy IT systems often introduce technical limitations and operational risks that modern platforms are designed to avoid.

One of the most common issues with legacy infrastructure is the growing incompatibility between software packages, libraries, and system components. Older technologies are often built on tightly coupled dependencies that were designed years ago, before decoupling was really put into practice. Over time, these systems become difficult to update because the software has drastically changed or isn’t maintained anymore.

Here are some examples of what issues you can run into with older infrastructure:

  • Updating one package or library can unintentionally break another component that relies on an older version. I’ve run into this myself before, where one dependency leads to another and another, and hours pass as you’re recompiling a dozen packages!
  • Lack of documentation on how the services interact can make it difficult to upgrade.
  • Lastly, modern monitoring, security, or automation tools may not integrate cleanly with outdated systems. In HA environments, even small compatibility issues can trigger major disruptions.

Plan for infrastructure modernization as part of your high availability strategy

Maintaining high availability with older infrastructure presents both technical and operational challenges. Package incompatibilities, limited vendor support, and declining internal expertise can all threaten system stability and increase the risk of downtime.

While legacy systems may continue to serve important business functions, organizations should proactively plan for infrastructure modernization, improve documentation practices, and invest in knowledge transfer before critical expertise is lost. A strong HA strategy is about ensuring long-term reliability, security, and operational resilience for the future.

Author: Cassy Hendricks-Sinke, Senior System Engineer, IT Operations, SIOS

Reproduced with permission from SIOS

 

Filed Under: Clustering Simplified Tagged With: High Availability

LifeKeeper Generic Applications for High Availability and Disaster Recovery

June 4, 2026 by Jason Aw Leave a Comment

LifeKeeper Generic Applications for High Availability and Disaster Recovery

LifeKeeper Generic Applications for High Availability and Disaster Recovery

Keys to Success for Protecting Business-Critical Applications

High Availability and Disaster Recovery have to cover a broad range of use cases. There are as many use cases as there are organizations, far exceeding the capabilities of any single High Availability and Disaster Recovery solution to provide out-of-the-box support for every scenario. While many common applications have a wide array of High Availability and Disaster Recovery solutions available, more specific use cases limit the selection available to protect business-critical applications.

Of course, LifeKeeper cannot cover every use case out of the box. LifeKeeper, however, provides a versatile and flexible framework that can be adapted to a wide range of use cases to remedy this limitation. While powerful, this framework can appear complex to an outsider. This blog is here to help give a leg up when starting to conceptualize a Generic Application Recovery Kit for your specific use case.

Related Blogs and Background Reading Recommendations

Within this blog, there is also the assumed familiarity with the LifeKeeper Resource Hierarchy framework and LifeKeeper Clustering in general. For background on these topics, the blogs listed below provide fantastic context. Additionally, this blog builds upon a previous blog regarding one means to close the gap between possible use cases and supported protection mechanisms via the use of the “Quick Service Protection Application Recovery Kit” (QSP ARK) within LifeKeeper, linked below.

  • Linux Clustering / Windows Clustering (Writing credits to Ms. Hoagland, Vice President of SIOS Global Sales and Marketing, and the SIOS Marketing Team)
  • Application Intelligence in Relation to High Availability (Writing credits to Ms. Hendricks-Sinke, Senior Software Engineer at SIOS)
  • Resource actions and background on The Generic Application Recovery Kit (Writing credits to Mr. Birmingham, Senior Technical Evangelist)
  • Choosing Between GenApp and QSP: Tailoring High Availability for Your Critical Applications (Writing credits to Ms. Hendricks-Sinke, Senior Software Engineer at SIOS).

This blog, however, will explore the options available when the QSP ARK cannot meet the demands for High Availability and Disaster Recovery for a particular application or use case.

Conceptualizing Applications and Defining the Approach

Asking the Smallest Question About Application Health

System administration and software engineering are both fields with lots of nuance. There can be so many different elements behind a question that a simple, straightforward answer can be difficult to obtain. Conversational, this can be easily navigated. In code, complex answers are difficult to accommodate. Asking the “smallest” question is the practice of targeting an inquiry to the smallest element possible, while ensuring that the answer has clearly defined criteria.

“Is the application running?” This is a “big” question; it may require a verbose answer. Yes, the application is running, but it is not responding. Yes, the application is running, but it is running on that other system – not the one you’re talking about. The criteria of the answer are ambiguous, and the answer is nuanced – a level of detail that developers would rather not have to handle.

“Is the application’s process running, and is the application actively responding to queries?”

Though longer to say, it is a smaller question. It clearly defines the conditions under which the answer is yes or no. While this change is an improvement, it is not yet the “smallest” question. The previous falls victim to the same pitfall of asking “Are both X and Y true?” A yes or no answer cannot give the level of detail to determine the truth of X and Y independently. The smallest question requires specificity; it must provide full insight into the status of the smallest element of the greater whole. “Is the application’s process running on the desired system?” That’s a small question – in this case, this is the smallest question. Keep in mind, there might be multiple “smallest” questions – in this example, “is the application responding to queries” would also qualify.

While questions can be broken down almost indefinitely, there is a limit. Asking “the smallest question?, comes with the implication of “Asking the smallest question that still provides useful/actionable information”. Asking “Am I on the train to Philadelphia?” is sufficient; going further to ask “Am I on the train to Philadelphia, and which direction is Philadelphia?”  provides more information – but it is not actionable. I cannot change the direction of the train. I know from the answer to “Am I on the train to Philadelphia?” if I need to call into work to inform my boss that I will be late.

Though clear in this example, this is less obvious when developing for a generic application. Throughout the process of protecting the generic application, one must still keep perspective on the bigger picture. This, like anything else, is a skill – with practice and collaboration comes the ability to determine when a question is the smallest question, and when further nuance stops providing additional useful information.

Broad questions that have been broken into smaller, specific, and targeted inquiries about individual elements are the basis on which Generic Application Recovery Kits are built. Each “big question” can be answered through a composite of the answers provided for each of the elements implicated within.

Once the questions are broken into their smallest elements, the information that needs to be relayed becomes significantly clearer. Knowing the information needed, the remaining work in developing a Generic Application Recovery Kit is all a matter of how to get the information that is needed from the information that is provided. One must work with the information that is given.

Working with Application APIs and LifeKeeper APIs

Often, applications provide Graphical User Interfaces (GUIs) to display information or show changes that occurred to the application. While fantastic for human-driven use, this is less useful when the administration is being done by an application. GUIs are for people to use, and applications (foregoing a massive amount of programming effort and unnecessary complexity) are not equipped to interface with the GUI of another application as a human would. For the purposes of LifeKeeper and a Generic Application Resource, the exchange of information between the Generic Application Recovery Kit’s action scripts and the application being protected must be done by an Application Programming Interface, or an “API”.

LifeKeeper provides its own API for interacting with the LifeKeeper, the hierarchy, and the resources within the hierarchy. In the case of LifeKeeper’s API, the command-line utilities contained within the product are the easiest to use in a Generic Application. As a general recommendation, only the command line utilities outlined in LifeKeeper Product Documentation (Linux Commands Documentation / Windows Commands Documentation) should be used. Even with this recommendation, these commands should be used with care and attention to detail to ensure that unintended actions are not taken.

Of course, LifeKeeper is not the only factor in the Generic Application. The application being protected will also need an API presented so action scripts can leverage the application’s API to achieve the desired outcomes. Developing a Generic Application Recovery Kit does require knowledge of the protected application’s API and the use of that API within the action scripts that compose the Generic Application Recovery Kit.

Using Return Codes and Output Streams in Recovery Scripts

Whether it is the API for LifeKeeper or the protected application, information will be primarily output in two ways:

  • Return Codes
  • Output Streams (sometimes called “STDOUT/STDERR output” or just “terminal output”)

How Return Codes Help Determine Success or Failure

Return codes, in the broadest sense, provide a quick way to see if a utility succeeded or failed. Typically (In the context of a shell environment), a return code of 0 indicates success, while a nonzero return code indicates failure.

Depending on the application, the exact value of the return code may give more insight into the error encountered. Often, the outcome of actions performed via an application’s API can be surmised simply by checking the return code.

In more nuanced cases, it may be that the return code is simply used to tell the program which course of action to take following a call to the application’s API. Return codes are especially useful when dealing with utilities that concern the state of some underlying element.

How Output Streams Provide More Detailed Application Information

Output Streams, while more complex to make use of in a program, are sometimes necessary for information exchange or to verify outcomes. If running a utility to get the system’s hostname, the return code alone will not indicate what that hostname is, unless the utility was successful in retrieving the hostname. In some cases, an API utility may return a successful return code if the requested information was obtained, but that information has to be evaluated for validity based on the circumstances.

Whether using return codes or output streams, developing a Generic Application requires the use of the information at hand. When thinking of ways to achieve resource actions (outlined in the next section) or determine information about an application or LifeKeeper resource, try to think in terms of return codes and output streams, not GUI interfaces. It can be helpful to imagine trying to convey information over the phone. This is to say, information is best communicated, actions are best defined, and scenarios are best handled when the inputs and outputs of utilities are reported exactly as they are to be provided as input or are reported as output.

Building a Foundation for Generic Application Protection

This section kept the strategies very conceptual. These strategies lay the foundation for thinking about an application through the responses provided to questions and actions issued upon that application.  Going forward, the approach will become more specific to LifeKeeper and the process of creating a Generic Application Recovery Kit. In the meantime, these strategies develop like any other skill, through practice. In technical communication, writing procedures, or any capacity in which you find yourself, practicing these conceptualization strategies will benefit not only in the short term but over time as well.

Need help protecting a business-critical application that does not fit a standard high availability model? SIOS can help you evaluate your environment and determine the right LifeKeeper approach. Request a demo today.

Author: Philip Merry, L3 Support Engineer at SIOS Technology Corp.

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: disaster recovery, High Availability

SIOS Enterprise Support Guide: What Your Plan Covers

May 30, 2026 by Jason Aw Leave a Comment

SIOS Enterprise Support Guide What Your Plan Covers

SIOS Enterprise Support Guide: What Your Plan Covers

What’s Included in Your SIOS Enterprise Support Plan?

Here are some quick tips for what is covered and not covered with Enterprise level support, and where to go for additional information based on three common scenarios.

24/7 Support for Critical System Downtime

Scenario 1: System Down After Hours
Joan’s team: It’s 7 pm EST on Sunday. The routine switchover between SIOS LifeKeeper cluster nodes should have been simple.  But something unexpected happened, resulting in the switchover failure. Despite all of the team’s efforts to resolve the issue, the cluster remains down.  Joan needs help, but she is not sure that her SIOS Technical support plan covers weekends or how long it will take to get a support person on the phone.

Customers who have purchased (or renewed) their Enterprise level support prior to an incident have access to receive support 24 hours a day, 7 days a week.  This support includes weekends and holidays to address Critical Issues.  Critical Issues mean down production systems or applications, where Customer data cannot be accessed using SIOS Programs.  For all Priority 1 (critical) issues, where normal operation results in the loss of access to your production data, SIOS provides a 2-hour response time.

If Joan has valid Enterprise support, she will be able to reach out to the SIOS support team, and her after-hours issue will be covered.

Installation and Configuration Support

Scenario 2: Installation Assistance Needed
Scott’s team: It’s 4 pm EST on Thursday. The approvals have been completed for the new infrastructure project, including the required high availability configuration for critical applications and data. At the kickoff, the stakeholders moved the date for go-live. As a result, the team needs to get the systems installed and configured quickly to avoid service interruption.  Scott’s team knows how to configure the application and server, but they want to be doubly sure they install the HA solution correctly. They need help, but Scott’s not sure that their support plan covers help with installation errors.

Since Scott’s team is in the deployment phase, the new infrastructure project involves systems that have not been validated or successfully put into production.  If Scott’s team has valid SIOS Enterprise level support, he will have access to SIOS product documentation and installation pointers.  However, assistance with installation and configuration is not covered under Scott’s Enterprise support, but he can contact his SIOS sales representative to arrange a paid Professional Services installation engagement. This engagement will ensure that Scott’s team gets the assistance they need to properly install, configure, and validate their cluster. SIOS provides a wide range of professional services designed to help customers quickly and cost-effectively implement, manage, and maintain their HA environments.

Root Cause Analysis (RCA) After Failover

Scenario 3: Post-Failover RCA Support
Amol’s team: It’s 2 am EST on Tuesday.  An alert has been sent out to the entire application team at AjaxBjax Corp. The cluster protecting the company’s most critical application system is conducting a failover.  Amol checks the application dashboard and discovers that the failover was successful and all applications are functioning.  However, Amol knows that management will want some explanations and assurances.  Amol wants to make sure that all application services are up and functioning, but he isn’t sure that their support plan covers whatever this is.

Amol’s team is looking for an RCA and the confidence that their system is going to continue to be operational. Amol’s data is accessible, and his application is fully functional. His system is not a critical down production server, nor a P1 issue.  However, if AjaxBjax Corp has valid Enterprise support for their cluster, they will be able to reach out to the SIOS support team for guidance around the clock (US East), Monday through Friday, for RCA issues.  Amol’s 2 am call will be routed to one of the knowledgeable SIOS support centers, where the team will begin working with Amol.

Additional Questions About Contacting SIOS Support

Amol and Joan were able to contact support via the Support Hotline (US: 877.457.5113; International: +1.803.808.4270) with coverage included by their Enterprise Support.  Scott was able to receive the help he needed, not from the Support team, but through the purchase of services to assist with configuration and installation.  But what about other scenarios, where can Scott, Amol, Joan, and others find more about their support levels and support details?  Or whether their product has reached the maintenance or extended support phases?

When you need to find additional information about your support agreement, you can consult the SIOS Technical Support Agreement (TSA), which is included with each order.  The TSA is also conveniently located on our download site, and can be requested via an email to the SIOS Support Team at support@us.sios.com.  Additionally, product schedules and support tier information can be found online at the Product Lifecycle page.

Customers who already know what’s covered under their plan, but need help with a problem, answers to a general question, root cause analysis, the latest software, or pointers to more information can open a new case via the Support Portal website or via email to the Support Inbox at support@us.sios.com.  Once your case is created, the team will work to provide timely responses and resolution.

Author: Cassius Rhue VP, Customer Experience

Reproduced with permission from SIOS

Filed Under: Clustering Simplified Tagged With: Application availability, High Availability

  • « Previous Page
  • 1
  • 2
  • 3
  • 4
  • …
  • 117
  • Next Page »

Recent Posts

  • Patch Management
  • Observation and Calculation: Applying Experience to Better Business Decisions
  • Why High Availability and Disaster Recovery Are Now Business Priorities
  • Disaster Recovery Incident Response: The Discipline of Not Reacting Impulsively
  • High Availability and Disaster Recovery Everywhere: From General Concepts to Generating Solutions

Most Popular Posts

Maximise replication performance for Linux Clustering with Fusion-io
Failover Clustering with VMware High Availability
create A 2-Node MySQL Cluster Without Shared Storage
create A 2-Node MySQL Cluster Without Shared Storage
SAP for High Availability Solutions For Linux
Bandwidth To Support Real-Time Replication
The Availability Equation – High Availability Solutions.jpg
Choosing Platforms To Replicate Data - Host-Based Or Storage-Based?
Guide To Connect To An iSCSI Target Using Open-iSCSI Initiator Software
Best Practices to Eliminate SPoF In Cluster Architecture
Step-By-Step How To Configure A Linux Failover Cluster In Microsoft Azure IaaS Without Shared Storage azure sanless
Take Action Before SQL Server 20082008 R2 Support Expires
How To Cluster MaxDB On Windows In The Cloud

Join Our Mailing List

Copyright © 2026 · Enterprise Pro Theme on Genesis Framework · WordPress · Log in