求帮忙:centos当机分析

82 views
Skip to first unread message

Qf Yang

unread,
May 24, 2013, 3:49:27 AM5/24/13
to sh...@googlegroups.com
一台centos6,只跑了squid作反向代理,一年多运行平衡,近两天异常当机,只能让idc重启,查/var/log/message发现有如下消息。
请各位帮忙看下问题出在哪里?
多谢了!


May 24 14:49:45 centos2010 squid[2676]:   always_direct = 0
May 24 14:49:45 centos2010 squid[2676]:    never_direct = 0
May 24 14:49:45 centos2010 squid[2676]:        timedout = 0
May 24 14:50:35 centos2010 kernel: e1000e: eth1 NIC Link is Up 100 Mbps Full Duplex, Flow Control: RX/TX
May 24 14:50:35 centos2010 kernel: 0000:02:00.0: eth1: 10/100 speed: disabling TSO
May 24 14:50:57 centos2010 squid[2676]: Failed to select source for 'http://112.65.244.92/images/nav_on_left.gif'
May 24 14:50:57 centos2010 squid[2676]:   always_direct = 0
May 24 14:50:57 centos2010 squid[2676]:    never_direct = 0
May 24 14:50:57 centos2010 squid[2676]:        timedout = 0
May 24 14:56:37 centos2010 squid[2676]: TCP connection to tc4.bioon.com/80 failed
May 24 14:56:48 centos2010 kernel: irq 70: nobody cared (try booting with the "irqpoll" option)
May 24 14:56:48 centos2010 kernel: Pid: 0, comm: swapper Tainted: G        W  ----------------  2.6.32-71.29.1.el6.i686 #1
May 24 14:56:48 centos2010 kernel: Call Trace:
May 24 14:56:48 centos2010 kernel: [<c04af074>] ? __report_bad_irq+0x24/0x90
May 24 14:56:48 centos2010 kernel: [<c04af230>] ? note_interrupt+0x150/0x190
May 24 14:56:48 centos2010 kernel: [<c04b0a71>] ? move_native_irq+0x11/0x50
May 24 14:56:48 centos2010 kernel: [<c04af6dd>] ? handle_edge_irq+0xbd/0x130
May 24 14:56:48 centos2010 kernel: [<c040c192>] ? handle_irq+0x32/0x60
May 24 14:56:48 centos2010 kernel: [<c040b7c7>] ? do_IRQ+0x47/0xc0
May 24 14:56:48 centos2010 kernel: [<c040a0f0>] ? common_interrupt+0x30/0x38
May 24 14:56:48 centos2010 kernel: [<c04500d8>] ? print_tainted+0x8/0xb0
May 24 14:56:48 centos2010 kernel: [<c065420a>] ? acpi_idle_enter_bm+0x25e/0x28f
May 24 14:56:48 centos2010 kernel: [<c0740b22>] ? cpuidle_idle_call+0x72/0xf0
May 24 14:56:48 centos2010 kernel: [<c0408884>] ? cpu_idle+0x94/0xd0
May 24 14:56:48 centos2010 kernel: [<c0807e37>] ? start_secondary+0x209/0x24e
May 24 14:56:48 centos2010 kernel: handlers:
May 24 14:56:48 centos2010 kernel: [<f86bdb60>] (e1000_msix_other+0x0/0xb0 [e1000e])
May 24 14:56:48 centos2010 kernel: Disabling IRQ #70
May 24 14:56:58 centos2010 squid[2676]: TCP connection to tc4.bioon.com/80 failed
May 24 14:57:19 centos2010 squid[2676]: TCP connection to tc4.bioon.com/80 failed
May 24 14:57:35 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 14:57:40 centos2010 squid[2676]: TCP connection to tc4.bioon.com/80 failed
May 24 14:57:52 centos2010 squid[2676]: TCP connection to 222.73.104.106/80 failed
May 24 14:58:10 centos2010 squid[2676]: TCP connection to tc4.bioon.com/80 failed
May 24 14:58:13 centos2010 squid[2676]: TCP connection to 222.73.104.106/80 failed

Qf Yang

unread,
May 24, 2013, 3:56:45 AM5/24/13
to sh...@googlegroups.com
后面是这样的

May 24 15:00:41 centos2010 squid[2676]: TCP connection to 222.73.104.106/80 failed
May 24 15:01:02 centos2010 squid[2676]: TCP connection to 222.73.104.106/80 failed
May 24 15:01:02 centos2010 squid[2676]: Detected DEAD Parent: www
May 24 15:01:23 centos2010 squid[2676]: TCP connection to 222.73.104.106/80 failed
May 24 15:01:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:02:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:03:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:04:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:05:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:06:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:07:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:08:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:09:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:10:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:11:36 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:12:37 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:13:37 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:14:37 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:15:37 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:16:37 centos2010 kernel: possible SYN flooding on port 80. Sending cookies.
May 24 15:17:11 centos2010 init: tty (/dev/tty1) main process (2004) killed by TERM signal
May 24 15:17:11 centos2010 init: tty (/dev/tty2) main process (2011) killed by TERM signal
May 24 15:17:11 centos2010 init: tty (/dev/tty3) main process (2017) killed by TERM signal
May 24 15:17:11 centos2010 init: tty (/dev/tty4) main process (2022) killed by TERM signal
May 24 15:17:11 centos2010 init: tty (/dev/tty5) main process (2024) killed by TERM signal
May 24 15:17:11 centos2010 init: tty (/dev/tty6) main process (2026) killed by TERM signal
May 24 15:17:12 centos2010 acpid: exiting
May 24 15:17:12 centos2010 auditd[1734]: The audit daemon is exiting.
May 24 15:17:12 centos2010 kernel: type=1305 audit(1369379832.437:1346): audit_pid=0 old=1734 auid=4294967295 ses=4294967295 subj=system_u:system_r:auditd_t:s0 res=1
May 24 15:17:12 centos2010 NET[4180]: /etc/sysconfig/network-scripts/ifdown-post : updated /etc/resolv.conf
May 24 15:17:13 centos2010 kernel: IPv6 over IPv4 tunneling driver
May 24 15:17:13 centos2010 kernel: sit0: Disabled Privacy Extensions
May 24 15:17:13 centos2010 cpuspeed: Disabling ondemand cpu frequency scaling governor
May 24 15:17:13 centos2010 kernel: Kernel logging (proc) stopped.
May 24 15:17:13 centos2010 rsyslogd: [origin software="rsyslogd" swVersion="4.6.2" x-pid="1750" x-info="http://www.rsyslog.com"] exiting on signal 15.
May 24 15:18:52 centos2010 kernel: imklog 4.6.2, log source = /proc/kmsg started.
May 24 15:18:52 centos2010 rsyslogd: [origin software="rsyslogd" swVersion="4.6.2" x-pid="1721" x-info="http://www.rsyslog.com"] (re)start
May 24 15:18:52 centos2010 kernel: Initializing cgroup subsys cpuset
May 24 15:18:52 centos2010 kernel: Initializing cgroup subsys cpu
May 24 15:18:52 centos2010 kernel: Linux version 2.6.32-71.29.1.el6.i686 (mock...@c6b5.bsys.dev.centos.org) (gcc version 4.4.4 20100726 (Red Hat 4.4.4-13) (GCC) ) #1 SMP Mon Jun 27 18:07:00 BST 2011
May 24 15:18:52 centos2010 kernel: KERNEL supported cpus:
May 24 15:18:52 centos2010 kernel:  Intel GenuineIntel
May 24 15:18:52 centos2010 kernel:  AMD AuthenticAMD
May 24 15:18:52 centos2010 kernel:  NSC Geode by NSC



2013/5/24 Qf Yang <fen...@gmail.com>

Wizard

unread,
May 24, 2013, 4:00:36 AM5/24/13
to shlug



2013/5/24 Qf Yang <fen...@gmail.com>
感觉这个call trace不是造成当机的原因。 看代码,它就是把这个中断停掉了。 
 

--
-- You received this message because you are subscribed to the Google Groups Shanghai Linux User Group group. To post to this group, send email to sh...@googlegroups.com. To unsubscribe from this group, send email to shlug+un...@googlegroups.com. For more options, visit this group at https://groups.google.com/d/forum/shlug?hl=zh-CN
---
您收到此邮件是因为您订阅了 Google 网上论坛的“Shanghai Linux User Group”论坛。
要退订此论坛并停止接收此论坛的电子邮件,请发送电子邮件到 shlug+un...@googlegroups.com
要查看更多选项,请访问 https://groups.google.com/groups/opt_out。
 
 



--
Wizard

Qf Yang

unread,
May 24, 2013, 4:14:08 AM5/24/13
to sh...@googlegroups.com
请教一下,接下来怎么查找当机原因呢?
是否还有别的某个日志需要查呢


2013/5/24 Wizard <wizard...@gmail.com>

Wizard

unread,
May 24, 2013, 4:18:14 AM5/24/13
to shlug



2013/5/24 Qf Yang <fen...@gmail.com>
请教一下,接下来怎么查找当机原因呢?
是否还有别的某个日志需要查呢

 
哈,我也不是很懂。 我就是看了一下这个call trace的输出。

等大侠指导。
 

--
-- You received this message because you are subscribed to the Google Groups Shanghai Linux User Group group. To post to this group, send email to sh...@googlegroups.com. To unsubscribe from this group, send email to shlug+un...@googlegroups.com. For more options, visit this group at https://groups.google.com/d/forum/shlug?hl=zh-CN
---
您收到此邮件是因为您订阅了 Google 网上论坛的“Shanghai Linux User Group”论坛。
要退订此论坛并停止接收此论坛的电子邮件,请发送电子邮件到 shlug+un...@googlegroups.com
要查看更多选项,请访问 https://groups.google.com/groups/opt_out。
 
 



--
Wizard

尹川

unread,
May 24, 2013, 5:17:10 AM5/24/13
to sh...@googlegroups.com
查查是否有攻击。
开防火墙了吗。好多的syn


--
一片树林里分出两条路——
        而我选择了人迹更少的一条,
        从此决定了我一生的道路。

Wizard

unread,
May 24, 2013, 5:30:41 AM5/24/13
to shlug



在 2013年5月24日下午5:17,尹川 <yinchu...@gmail.com>写道:
查查是否有攻击。
开防火墙了吗。好多的syn


syn是这个当机的原因?




--
Wizard

尹川

unread,
May 24, 2013, 5:34:48 AM5/24/13
to sh...@googlegroups.com
之前遇到过大流量的攻击后端某台服务器导致varnish自动重启,日志里面也是一堆syn,然后kernel崩掉,也可能不是这个原因。建议查查网络流量。


On Friday, May 24, 2013, Wizard wrote:
--
-- You received this message because you are subscribed to the Google Groups Shanghai Linux User Group group. To post to this group, send email to sh...@googlegroups.com. To unsubscribe from this group, send email to shlug+un...@googlegroups.com. For more options, visit this group at https://groups.google.com/d/forum/shlug?hl=zh-CN
---
您收到此邮件是因为您订阅了 Google 网上论坛的“Shanghai Linux User Group”论坛。
要退订此论坛并停止接收此论坛的电子邮件,请发送电子邮件到 shlug+un...@googlegroups.com
要查看更多选项,请访问 https://groups.google.com/groups/opt_out。
 
 

Wizard

unread,
May 24, 2013, 5:56:51 AM5/24/13
to shlug
在 2013年5月24日下午5:34,尹川 <yinchu...@gmail.com>写道:
之前遇到过大流量的攻击后端某台服务器导致varnish自动重启,日志里面也是一堆syn,然后kernel崩掉,也可能不是这个原因。建议查查网络流量。


那这样的话,写个iptable啥的。

不过这样很难复现这个问题吧。 而且也不确定是不是一定是syn引起的。  


--
Wizard

尹川

unread,
May 24, 2013, 5:58:51 AM5/24/13
to sh...@googlegroups.com
查流量监控


On Friday, May 24, 2013, Wizard wrote:
--
-- You received this message because you are subscribed to the Google Groups Shanghai Linux User Group group. To post to this group, send email to sh...@googlegroups.com. To unsubscribe from this group, send email to shlug+un...@googlegroups.com. For more options, visit this group at https://groups.google.com/d/forum/shlug?hl=zh-CN
---
您收到此邮件是因为您订阅了 Google 网上论坛的“Shanghai Linux User Group”论坛。
要退订此论坛并停止接收此论坛的电子邮件,请发送电子邮件到 shlug+un...@googlegroups.com
要查看更多选项,请访问 https://groups.google.com/groups/opt_out。
 
 

Wizard

unread,
May 24, 2013, 6:01:41 AM5/24/13
to shlug



在 2013年5月24日下午5:58,尹川 <yinchu...@gmail.com>写道:
查流量监控



这个高级,不会。。。 

--
Wizard

Qf Yang

unread,
May 24, 2013, 6:15:14 AM5/24/13
to sh...@googlegroups.com
似乎不是,监测下来,没有发现syn很大
#netstat -n | awk '/^tcp/ {++S[$NF]} END {for(a in S) print a, S[a]}'
LAST_ACK 10
SYN_RECV 5
CLOSE_WAIT 1
ESTABLISHED 171
FIN_WAIT1 11
FIN_WAIT2 4
CLOSING 5
TIME_WAIT 321

当机时没有查,下等多次查看,都 SYN_RECV都很小,不超过过几十的。


--

Sherlock

unread,
May 24, 2013, 8:43:34 AM5/24/13
to Shanghai Linux User Group
Rule
No.1 这么大量的日志不要直接在 邮件正文里直接写
No.2 日志要给全

虽然有 syn attack 但 已经发cookie了,所以是其他原因的可能性更大。
你的crash point报的是IRQ的问题


2013/5/24 Qf Yang <fen...@gmail.com>



--
==========
      InitX
==========

Qf Yang

unread,
May 29, 2013, 1:30:52 AM5/29/13
to sh...@googlegroups.com
多谢提醒,下次会注意的。

dmesg 消息里有这样一条

    May 24 14:56:48 centos2010 kernel: handlers:
    May 24 14:56:48 centos2010 kernel: [<f86bdb60>] (e1000_msix_other+0x0/0xb0 [e1000e])
问题应该出在 网上驱动模块上,Intel Corporation 82574L Gigabit Network 网卡的 msi/msi-x驱动有bug,
参考intel官方文档, http://downloadmirror.intel.com/15817/eng/README.txt
    IntMode
    -------
    Valid Range: 0-2 (0=legacy, 1=MSI, 2=MSI-X)
    Default Value: 2 
通过 在 modprobe.d/ 写入自定义文件,内容如下
  alias eth0 e1000e
  alias eth1 e1000e
  options e1000e IntMode=0,0
屏蔽掉msi/msi-x.
今天上午做的,接下来监测一下效果再说。
还请有相关经验的朋友帮忙看看是否可行。

Qf Yang

unread,
May 29, 2013, 3:20:55 AM5/29/13
to sh...@googlegroups.com
再问:生产环境下,ACPI是否最好关掉?
昨天为了解决e1000e的msi/msi-x bug的问题,加 acpi=off 启动参数关掉acpi,不知是否有必要打开?

Sherlock

unread,
May 30, 2013, 6:08:14 PM5/30/13
to Shanghai Linux User Group
ACPI 关不关和生产环境无关,关键看你硬件和软件的支持程度。
我接触的很多生产系统开ACPI的很多


2013/5/29 Qf Yang <fen...@gmail.com>



--
==========
      InitX
==========
Reply all
Reply to author
Forward
0 new messages