Google Groups no longer supports new Usenet posts or subscriptions. Historical content remains viewable.
Dismiss

T5xxx servers short "freeze" behaviour

22 views
Skip to first unread message

Egrama

unread,
Jan 22, 2010, 5:39:23 AM1/22/10
to
Hi guys,

On a T5220 and a T5240 I noticed this strange behaviour: sometimes the
machine is freezing for a couple of secconds and then continues
working as if nothing happened. I noticed this because we are running
some realtime application and 2 secconds delays in processing
triggers alarms.
The machine CPU load is around 30% and also the memory.
I would say this is not a system related problem, but I noticed the
problem first hand when my terminal just hung and then the application
alerts came.
Has anybody experienced anything similar? I have no errors whatsoever
in the system logs.....
Any idea how to investigate this without a major performance impact ?

Thanks,
Emil

Richard B. Gilbert

unread,
Jan 22, 2010, 9:47:18 AM1/22/10
to
Egrama wrote:
> Hi guys,
>
> On a T5220 and a T5240 I noticed this strange behavior: sometimes the
> machine is freezing for a couple of seconds and then continues

> working as if nothing happened. I noticed this because we are running
> some realtime application and 2 seconds delays in processing

> triggers alarms.
> The machine CPU load is around 30% and also the memory.

Remember that the "CPU load" is an average over time. There is nothing
in that "30%" that precludes the CPU from being 100% busy for a few seconds.

ISTR something about "real time" priorities. I never had cause to use
them but: see
http://www.princeton.edu/~unix/Solaris/troubleshoot/schedule.html

or

Google!

Drazen Kacar

unread,
Jan 22, 2010, 1:20:22 PM1/22/10
to
Egrama wrote:
> Hi guys,
>
> On a T5220 and a T5240 I noticed this strange behaviour: sometimes the
> machine is freezing for a couple of secconds and then continues
> working as if nothing happened. I noticed this because we are running
> some realtime application and 2 secconds delays in processing
> triggers alarms.
> The machine CPU load is around 30% and also the memory.
> I would say this is not a system related problem, but I noticed the
> problem first hand when my terminal just hung and then the application
> alerts came.
> Has anybody experienced anything similar? I have no errors whatsoever
> in the system logs.....

I've seen something similar, but for a somewhat longer time period.
Another box announced the same IP address, so switch sent all network
packets there. It looked like the machine was hung, although it was
working perfectly fine.

You can confirm this by being logged in on the console via lights out
management. That will still work while network connections appear
unresponsive. Although your 2 seconds seem to short to catch this
behaviour reliably.

Is all your monitoring network based?

OTOH, perhaps your real time application spawned more threads than there
are CPUs, so some of them have to wait until another real-time thread goes
to sleep. The OS will not preempt RT class. How many CPUs do you have and
how many threads the application has?

--
.-. .-. Yes, I am an agent of Satan, but my duties are largely
(_ \ / _) ceremonial.
|
| da...@fly.srk.fer.hr

Chris

unread,
Jan 22, 2010, 5:01:30 PM1/22/10
to
On Jan 22, 1:20 pm, Drazen Kacar <d...@fly.srk.fer.hr> wrote:

> Egrama wrote:
> >  On a T5220 and a T5240 I noticed this strange behaviour: sometimes the
> >  machine is freezing for a couple of secconds and then continues
> >  working as if nothing happened. I noticed this because we are running
> >  some realtime application and 2 secconds delays in processing
> >  triggers alarms.
> >  The machine CPU load is around 30% and also the memory.
> >  I would say this is not a system related problem, but I noticed the
> >  problem first hand when my terminal just hung and then the application
> >  alerts came.
> >  Has anybody experienced anything similar? I have no errors whatsoever
> >  in the system logs.....

I have a call open to Sun support on that very issue. I've been
through several tech support people and haven't nailed it down yet.
However, it seems clear to me that it is a network interface problem.
There was a patch posted in mid December for a bug in the e1000g
driver that said the chipset would freeze up under certain
cirumstances. That patch didn't fix my problem. What I see is that
anything using the network interface is momentarily unreachable,
including a StorageTek 2510 iSCSI array direct attached to e1000g2. I
see freezeups in ssh terminal connections, I see alerts from mon which
is poking ping, http, https, drupal, etc, and I see alerts from Common
Array Manager claiming to have lost iSCSI connections. Meanwhile
serial connection to the console via ILOM is just dandy, and the
system seems perfectly responsive when looked at through that
connection. The load on the system is a miniscule fraction of what it
should be capable of. We haven't even ramped it up yet. It is supposed
to take over from an E250, but, at the moment, we are more comfortable
leaving a lot of our stuff on the E250, which, in principle, ought to
be 100 times slower or more.

Michael Laajanen

unread,
Jan 22, 2010, 5:01:44 PM1/22/10
to
Hi,

I don't know much about this exept from realtime apps :) but it sound
like something is falling to sleep or a garbage collection but on a OS
level I don't think its a garbage collction :) could it be a disk that
has powered down?

/michael

Drazen Kacar

unread,
Jan 23, 2010, 3:43:03 AM1/23/10
to
Chris wrote:

> I have a call open to Sun support on that very issue. I've been
> through several tech support people and haven't nailed it down yet.
> However, it seems clear to me that it is a network interface problem.
> There was a patch posted in mid December for a bug in the e1000g
> driver that said the chipset would freeze up under certain

Which server type? My T5240 has nxge on-board. Are there versions with
e1000g?

> cirumstances. That patch didn't fix my problem. What I see is that
> anything using the network interface is momentarily unreachable,
> including a StorageTek 2510 iSCSI array direct attached to e1000g2. I
> see freezeups in ssh terminal connections, I see alerts from mon which
> is poking ping, http, https, drupal, etc, and I see alerts from Common
> Array Manager claiming to have lost iSCSI connections. Meanwhile
> serial connection to the console via ILOM is just dandy, and the
> system seems perfectly responsive when looked at through that
> connection.

Hm. I had the same problem with T2000 attached to an 2540 array. It turned
out that 2540 had buggy firmware which advertised MAC address belonging to
the host where Common Array Manager was installed. ARP scans could not
detect it, though. Only e1000g0 (used to communicate with the array) was
affected. ALOM and other interfaces had no problems.

Support patched the array firmware and the problem went away.

Horst Scheuermann

unread,
Jan 25, 2010, 7:35:13 AM1/25/10
to

we had similar problems with X4500, the patch 141445-09 seams to help

--
11. Gebot: Wenn Du eine Fahrradklingel hörst, dreh Dich um, reiße
Mund, Nase und Augen auf, trete aber keinesfalls zur Seite.

Egrama

unread,
Jan 28, 2010, 10:27:04 AM1/28/10
to
On Jan 22, 4:47 pm, "Richard B. Gilbert" <rgilber...@comcast.net>
wrote:
> Egramawrote:

You are right abouut the CPU being an average, but I have seen many
overloaded Sun servers with cpu average close to 100% and load average
a few times the number of processors.
The ssh connection was a bit sluggish, but nothing like freezing.
Reading the posts below, I think that it might be network related. I
cannot reproduce the incident, so I cannot verify by using the serial
console instead of an ssh session.

Egrama

unread,
Jan 28, 2010, 10:29:34 AM1/28/10
to
On Jan 22, 8:20 pm, Drazen Kacar <d...@fly.srk.fer.hr> wrote:
> Egramawrote:
>      |        d...@fly.srk.fer.hr

There are a lot of virtual CPUs available - 256 to be exact! and few
real time threads - less than 10.
I think it could be network based.

Martha Starkey

unread,
Jan 29, 2010, 8:54:32 AM1/29/10
to
On 01/22/10 13:20, Drazen Kacar wrote:
> Egrama wrote:
>> Hi guys,
>>
>> On a T5220 and a T5240 I noticed this strange behaviour: sometimes the
>> machine is freezing for a couple of secconds and then continues
>> working as if nothing happened. I noticed this because we are running
>> some realtime application and 2 secconds delays in processing
>> triggers alarms.
>> The machine CPU load is around 30% and also the memory.
>> I would say this is not a system related problem, but I noticed the
>> problem first hand when my terminal just hung and then the application
>> alerts came.
>> Has anybody experienced anything similar? I have no errors whatsoever
>> in the system logs.....
>
> I've seen something similar, but for a somewhat longer time period.
> Another box announced the same IP address, so switch sent all network
> packets there. It looked like the machine was hung, although it was
> working perfectly fine.

That sounds like the "broadcom arp poisoning" issue with certain NIC
drivers acting in "Teamed mode". If you have Broadcom NICS on windows
teamed mode servers, you may need updated drivers:

http://blogs.sun.com/swas/entry/solaris_10_8_07_broadcom

http://blogs.sun.com/swas/entry/update_to_the_broadcom_pc

Here's Dell's updated driver page:

http://support.dell.com/support/topics/global.aspx/support/dsn/en/document?c=us&dl=false&l=en&s=gen&docid=49F4FB5AA612CFF6E040A68F5A28020D&doclang=en&cs

ChrisS

unread,
Feb 6, 2010, 11:02:47 AM2/6/10
to

Assuming they are both Enterprise servers and not the Netra T5220, the
latest firmware is 139439-08 (7.2.7.b). I've loaded this onto two of
my new SE T5220s without issues. Is diagnostics turned on Max in
ILOM? If, can you poweroff the host and turn it back on? Watch the
console carefully (start SP/console , via ILOM). I'm not sure if
diags get logged somewhere.

Someone suggested it might be NIC problems. Connect to the host's
console using the Service Processor (either NetMgr NIC or serial/
tip). If you still see "pausing" going on it "may not" be the NIC(s)
of the host (assuming you were remoted into the terminal before). Is
the application using NICs to process thru? Is there anything in /
var/adm/messages that would suggest a problem? If so, are you dumping
data via NFS or some other means?

Just food for thought. Good luck.


0 new messages