Want to win a PS4? Go Premium and enter to win our High-Tech Treats giveaway. Enter to Win


Cluster Services will not start

Posted on 2009-07-02
Medium Priority
Last Modified: 2013-12-16
Specifically, we are trying to setup a two-node cluster to provide a highly available apache server. After reviewing the documentation, it appears that shared storage may not be necessary, though we would like to have the document root be on shared storage eventually.

We have followed the steps laid out in the howto:


However when we try to start the service in Luci, the httpd service fails to start. We get the following errors in /var/log/messages:

Jun 29 16:22:36 habox1 clurgmgrd: [10855]: <err> Stopping Service
apache:Apache_Test_Srvr > Failed
Jun 29 16:22:36 habox1 clurgmgrd[10855]: <notice> stop on apache
"Apache_Test_Srvr" returned 1 (generic error)
Jun 29 16:22:36 habox1 clurgmgrd[10855]: <crit> #13: Service
service:Web_Server failed to stop cleanly
Jun 29 16:26:29 habox1 clurgmgrd[10855]: <notice> Starting disabled
service service:Web_Server
Jun 29 16:26:29 habox1 clurgmgrd: [10855]: <err> Looking For IP
Addresses [apache:Apache_Test_Srvr] > Failed - No IP Addresses Found
Jun 29 16:26:29 habox1 clurgmgrd[10855]: <notice> start on apache
"Apache_Test_Srvr" returned 1 (generic error)
Jun 29 16:26:29 habox1 clurgmgrd[10855]: <warning> #68: Failed to start
service:Web_Server; return value: 1
Jun 29 16:26:29 habox1 clurgmgrd[10855]: <notice> Stopping service
Jun 29 16:26:35 habox1 clurgmgrd: [10855]: <err> Checking Existence Of
File /var/run/cluster/apache/apache:Apache_Test_Srvr.pid
[apache:Apache_Test_Srvr] > Failed - File Doesn't Exist
Jun 29 16:26:35 habox1 clurgmgrd: [10855]: <err> Stopping Service
apache:Apache_Test_Srvr > Failed
Jun 29 16:26:35 habox1 clurgmgrd[10855]: <notice> stop on apache
"Apache_Test_Srvr" returned 1 (generic error)
Jun 29 16:26:35 habox1 clurgmgrd[10855]: <crit> #12: RG
service:Web_Server failed to stop; intervention required
Jun 29 16:26:35 habox1 clurgmgrd[10855]: <notice> Service
service:Web_Server is failed
Jun 29 16:26:35 habox1 clurgmgrd[10855]: <crit> #13: Service
service:Web_Server failed to stop cleanly

Can you advise us as to what the problem may be? Let us know if you need
more information.

my cluster.conf file created in web GUI (luci)
<?xml version="1.0"?>
<cluster alias="app_server" config_version="16" name="app_server">
        <fence_daemon clean_start="0" post_fail_delay="0" 
                <clusternode name="habox2.nimh.nih.gov" nodeid="1" 
                                <method name="1"/>
                <clusternode name="habox1.nimh.nih.gov" nodeid="2" 
                                <method name="1"/>
        <cman expected_votes="1" two_node="1"/>
                        <apache config_file="conf/httpd.conf" 
name="Apache_Test_Srvr" server_root="/etc/httpd" shutdown_wait="0"/>
                        <ip address="" monitor_link="1"/>
                        <script file="/etc/rc.d/init.d/httpd" 
                        <fs device="/dev/hda2" force_fsck="0" 
force_unmount="0" fsid="36806" fstype="ext3" mountpoint="/HA" 
name="httpd_content" self_fence="1"/>
                <service autostart="1" exclusive="0" max_restarts="0" 
name="Web_Server" recovery="restart" restart_expire_time="0">
                        <apache ref="Apache_Test_Srvr">
                                <ip ref=""/>
                                <script ref="script_test"/>
                                <fs ref="httpd_content"/>

Open in new window

Question by:Justin_Edmands
Welcome to Experts Exchange

Add your voice to the tech community where 5M+ people just like you are talking about what matters.

  • Help others & share knowledge
  • Earn cash & points
  • Learn & ask questions
  • 2
  • 2

Author Comment

ID: 24768145
need some help!

Expert Comment

ID: 24770784
I think it might be easier to use hearbeat with DRBD: www.linux-ha.org and www.drbd.org. They work very nicely together. Basically you set up some information in their config files about each other, which IP address they will share, etc. Then at the end of it all, you have the heartbeat process start up, which then starts up the DRBD shared storage and places symlinks on the system to point to the shared storage. Heartbeat then handles the starting and stopping of Apache. It is pretty easy to do, and both projects are very well documented and there are many how-to articles on the web for doing exactly what you want to do.
LVL 80

Expert Comment

ID: 24774609
IMHO, it is better to load balance a web server rather than set it up in a fail over cluster.
You could use rsync to synchronize the document root data.

You could setup a cluster resource dealing with a specific IP.
This will deal with making the IP "available all the time"

One error I see is that you are not assigning an IP that will move with the web server.

You have to setup an IP that will move between/among the nodes.


Author Comment

ID: 24789982
already got DRBD to work and all. need to do RedHat Cluster Suite

Accepted Solution

JabbaDow earned 1500 total points
ID: 24791698
I have no experience with Red Hat Cluster Suite, but I guess the first thing would be to make sure that you have a virtual IP address (i.e. an address bound to a virtual interface like eth0:0), and make sure that that address is working on the active node of the cluster. When you failover to the other node, that address needs to follow the active node. Then set your Apache to listen on that address. So instead of Listen *:80, you need to have "Listen x.x.x.x:80" apache directive. Make sure that the networking comes up before apache does.

Featured Post

Free Tool: Site Down Detector

Helpful to verify reports of your own downtime, or to double check a downed website you are trying to access.

One of a set of tools we are providing to everyone as a way of saying thank you for being a part of the community.

Question has a verified solution.

If you are experiencing a similar issue, please ask a related question

SSH (Secure Shell) - Tips and Tricks As you all know SSH(Secure Shell) is a network protocol, which we use to access/transfer files securely between two networked devices. SSH was actually designed as a replacement for insecure protocols that sen…
Introduction This article explores the design of a cache system that can improve the performance of a web site or web application.  The assumption is that the web site has many more “read” operations than “write” operations (this is commonly the ca…
Learn how to find files with the shell using the find and locate commands. Use locate to find a needle in a haystack.: With locate, check if the file still exists.: Use find to get the actual location of the file.:
This demo shows you how to set up the containerized NetScaler CPX with NetScaler Management and Analytics System in a non-routable Mesos/Marathon environment for use with Micro-Services applications.
Suggested Courses

636 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question