Linux server keeps "crashing"

Posted on 2013-01-02
Last Modified: 2013-01-17
I have a Debian Linux server that has started becoming unavailable/unresponsive as of about 2 weeks ago. In my experience, this is usually caused by a high server load, caused by Apache, a poorly written PHP script, a corrupt database, or sometimes a disk I/O related issue. In this case, though, this doesn't appear to be the case. I installed various utilities to log and warn about high server loads. One example is the sysstat utility. According to these, server load was not the issue.

It was also not a network issue, since, for example, the system log stopped logging at the times the server went down. If it was just a network problem, the system log would have continued to log.

I also couldn't find anything useful in syslog.

Here's an example of what my server load average looked like during the last "crash" (the server went down just after 13:35 or 13:36, and was restarted at 15:14):

# sar -q -f /var/log/sysstat/sa02 -s 13:00:01
Linux 2.6.26-2-686      01/02/13        _i686_

13:05:01      runq-sz  plist-sz   ldavg-1   ldavg-5  ldavg-15
13:15:01            2       153      0.13      0.06      0.01
13:25:02            0       146      0.07      0.18      0.11
13:35:01            3       147      0.12      0.11      0.09
Average:            2       149      0.11      0.12      0.07

15:14:40          LINUX RESTART

15:15:01      runq-sz  plist-sz   ldavg-1   ldavg-5  ldavg-15
15:25:01            2       152      0.14      0.56      0.50
15:35:01            2       152      0.00      0.10      0.26
15:45:01            2       145      0.02      0.11      0.18
15:55:01            1       157      0.10      0.05      0.10
16:05:01            1       156      1.32      0.99      0.48
16:15:01            1       173      0.72      1.09      0.81
16:25:01            2       151      0.10      0.23      0.46
16:35:01            3       145      0.00      0.04      0.24
16:45:01            2       169      0.15      0.61      0.45
16:55:01            2       169      0.18      0.18      0.27
17:05:01            2       161      0.08      0.30      0.30
17:15:01            4       162      0.23      0.31      0.29
17:25:01            2       164      0.03      0.08      0.16
17:35:01            2       165      0.06      0.05      0.09
17:45:01            2       168      0.00      0.02      0.05
17:55:01            2       164      0.03      0.10      0.08
18:05:01            2       168      0.15      0.21      0.13
Average:            2       160      0.19      0.30      0.29

Open in new window

I'm wondering if someone might be able to help me identify the problem. I realise there could be many possibilities, but a couple of starting points would be good.

Many thanks!
Question by:Julian Matz
LVL 31

Accepted Solution

farzanj earned 175 total points
Comment Utility
It could also be a hardware issue itself or a filesystem level issue.  Pay attention on the % used of the related file system.   Any issues with hardware?
LVL 25

Assisted Solution

madunix earned 165 total points
Comment Utility
Check the following:
- look in /var/log any suspicious
- do you have free drive space..
- are all the file systems OK? fsck
- memory diagnostic.... could be a bad piece of RAM
- check apache config
- check mysql config it could memory setting bigger than your actual RAM (if you have mysql) if you run mysql
- Apache service starts with no errors??
- check Apache error log if contains hints
LVL 76

Assisted Solution

arnold earned 160 total points
Comment Utility
As others suggest, look in /var/log/messages for a kernel panic.
You need to collect info, memory use, vmstat, iostat, top, and sysstat.
Similar to the sar report.

You can use to poll data using snmp
The data collection should be every minute.
LVL 21

Author Comment

by:Julian Matz
Comment Utility
Nothing in logs, but the hardware, bar the hard drive, was replaced, and I haven't had any crashes since. Not sure was the motherboard replaced, actually. I was guessing it could have been the CPU, but I could be wrong; no way to know for sure now, but the main thing is that it's fixed. Thanks for your help/suggestions.

Featured Post

Maximize Your Threat Intelligence Reporting

Reporting is one of the most important and least talked about aspects of a world-class threat intelligence program. Here’s how to do it right.

Join & Write a Comment

Introduction We as admins face situation where we need to redirect websites to another. This may be required as a part of an upgrade keeping the old URL but website should be served from new URL. This document would brief you on different ways ca…
Why Shell Scripting? Shell scripting is a powerful method of accessing UNIX systems and it is very flexible. Shell scripts are required when we want to execute a sequence of commands in Unix flavored operating systems. “Shell” is the command line i…
Learn how to get help with Linux/Unix bash shell commands. Use help to read help documents for built in bash shell commands.: Use man to interface with the online reference manuals for shell commands.: Use man to search man pages for unknown command…
Get a first impression of how PRTG looks and learn how it works.   This video is a short introduction to PRTG, as an initial overview or as a quick start for new PRTG users.

763 members asked questions and received personalized solutions in the past 7 days.

Join the community of 500,000 technology professionals and ask your questions.

Join & Ask a Question

Need Help in Real-Time?

Connect with top rated Experts

13 Experts available now in Live!

Get 1:1 Help Now