Google Groups no longer supports new Usenet posts or subscriptions. Historical content remains viewable.
Dismiss

Unix 5.0 in panic

0 views
Skip to first unread message

Erik Kurz

unread,
Sep 11, 2001, 8:52:14 AM9/11/01
to
Hi there,

unix stopped using following message:

unexpected trap in kernel mode:
...
...
PANIC: k_trap - kernel mode trap type 0x00000160
trying to dump 16287 pages to dumpdev hd (1/41) at block 0, 204 pages per
'.'

has anyone an idea an can help me ?
i'm a novice user!

thanks
erik


Karel Adams

unread,
Sep 11, 2001, 1:11:14 PM9/11/01
to

"Erik Kurz" <e...@beta.at> schreef in bericht news:9nl191$lhr$1...@murmel.gams.at...

And you will the answer that many novices got here:
Please do give as _much_ more information!!!

Some people will require exact version (output of uname -X)
and patchlevel (customquery listpatches | head -1)

For myself, I should at the very least like to know
-) if this is a newly installed machine that you
find no way to boot, or whether a server that's worked
for years without trouble and now suddenly dies on you
-) if you have been able to get it back to working
and if so, what you could read in /var/adm/syslog
and /var/adm/messages

Remember: the more you say, the better the answers you may get!

Grüße!
Karel


Erik Kurz

unread,
Sep 12, 2001, 4:36:41 AM9/12/01
to

"Karel Adams" <k_a...@glo.be> schrieb im Newsbeitrag
news:Myrn7.5749$35.5...@iguano.antw.online.be...
here is some info

+)the server is working for years without troubles
+)version: Release 3.2v5.0.0
KernelID 95/08/08

+)patchlevel: SCO:Tcl::7.3.2a oss 603a.TclX732a

+)syslog: no entry with the current date (file is to big to show info)
+)messages: no entry with the current date

+) i powered off the server and started it again, after that
i did an orderly shutdown (reboot)
at the moment it is working

greetings
erik

Karel Adams

unread,
Sep 12, 2001, 6:31:59 PM9/12/01
to

"Erik Kurz" <e...@beta.at> schreef in bericht news:9nn8mb$mrn$1...@murmel.gams.at...

>
> "Karel Adams" <k_a...@glo.be> schrieb im Newsbeitrag
> news:Myrn7.5749$35.5...@iguano.antw.online.be...
> >
> > "Erik Kurz" <e...@beta.at> schreef in bericht
> news:9nl191$lhr$1...@murmel.gams.at...
> > > Hi there,
> > > unix stopped using following message:
> > > unexpected trap in kernel mode:
> > > PANIC: k_trap - kernel mode trap type 0x00000160
> > > trying to dump 16287 pages to dumpdev hd (1/41) at block 0, 204 pages
> per
> > > '.'
> > > has anyone an idea an can help me ?
> > > i'm a novice user!
> >
> > And you will the answer that many novices got here:
> > Please do give as _much_ more information!!!
> >
> > Some people will require exact version (output of uname -X)
> > and patchlevel (customquery listpatches | head -1)
> >
> > For myself, I should at the very least like to know
> > -) if this is a newly installed machine that you
> > find no way to boot, or whether a server that's worked
> > for years without trouble and now suddenly dies on you
> > -) if you have been able to get it back to working
> > and if so, what you could read in /var/adm/syslog
> > and /var/adm/messages
> >
> > Remember: the more you say, the better the answers you may get!
> >
> > Grüße!
> > Karel
> >
> >
> here is some info
>
> +)the server is working for years without troubles

Well with any luck it ought to add many more!

> +)version: Release 3.2v5.0.0
> KernelID 95/08/08

That it rather old, not to say very old!

> +)patchlevel: SCO:Tcl::7.3.2a oss 603a.TclX732a

And this seams meagre. There must be a website
that tells you a minimum patch level but I am afraid
it will be a very long list for this old release.

> +)syslog: no entry with the current date (file is to big to show info)

Your pardon? ('''Bitte?''') At the very least 'tail' should yield some info.
Beyond that, vi has never failed me for the size of the file
I made it handle - I often used it on files of at least 10MB size!

> +)messages: no entry with the current date
>
> +) i powered off the server and started it again, after that
> i did an orderly shutdown (reboot)

This seems wise.

> at the moment it is working

Well that's the main thing, isn't it?

Finally: to my (limited) experience, this kind of trouble
is very hard to diagnosticize. If it doesn't come back
within a month of continuous operation, just forget it.
If it does, I think you ought to suspect hardware, and
start by swapping step for step, starting with memory
(easiest and cheapest), next CPU, next motherboard.

If a system begins to panic regularly, without it's
configuration (HW/SW) having been changed,
(check very thoroughly, here) then only
some dying hardware can be the cause.
Make sure to make _many_ backups!

Freundlich,
KA


Brian K. White

unread,
Sep 12, 2001, 11:18:45 PM9/12/01
to

"Karel Adams" <k_a...@glo.be> wrote in message
news:klRn7.5782$35.5...@iguano.antw.online.be...

while the memory might be the easiest and cheapest thing to replace, it is
among the least likely to change with age. The most likely peice of hardware
to look at for random problems after years of no problems, (barring the
obvious, cpu and powersupply fans) is the power supply.

if the power supply doesn't fix it, then all other components are about
equally likely, but I'd start with the motherboard (it has capacitors that
change value with age, as do some other components, but a dead cdrom drive
generally doesn't affect the rest of the system)

One thing I've seen, if the motherboard is old, try upping the voltage on
the cpu by 1 or 2 tenths of a volt. I have seen this work. I don't know
which of two most probable reasons it works. maybe a) the chip is "old" and
has changed inside, and now needs a higher voltage to work reliably. or
maybe b) the motherboard is "old" and when you set the jumpers for say, 2.9
volts, it's no longer actually delivering that, so you need to set it to
say, 3.1, to get 2.9. Either way, if this seems to fix the problem, do 2
things immediately. 1) install a new, oversized cpu fan and heat sink, and
use heatsink grease and remove any stickers or other foreign objects between
the cpu face and the heat sink. 2) tell the owner of the computer to be
ready to replace the server because this only buys you some time, it cannot
be considered a real fix.

Another thing I've seen more than once: scsi ribbon cables that touch metal,
where the insulation has worn away and the wire inside is just barely
visible, meaning it sometimes but not always gets shorted to the chassis
ground. You have to look close and careful for that, not just near sharp
corners either. Systems run mostly ok, but get a random corrupted file here
and there and messages in the log file and sometimes momentary server lockup
or even unexpected reboot. Basically, If you see symptoms that make you
think "hard drive is going south", look hard at the cable first.

--
Brian K. White -- br...@aljex.com -- http://www.aljex.com/bkw/
+++++[>+++[>+++++>+++++++<<-]<-]>>+.>.+++++.+++++++.-.[>+<---]>++.
filePro BBx Linux SCO Prosper/FACTS AutoCAD #callahans Satriani


Matthias Wiesner

unread,
Sep 13, 2001, 10:13:24 AM9/13/01
to
"Erik Kurz" <e...@beta.at> wrote in message news:<9nl191$lhr$1...@murmel.gams.at>...
Hi Erik,
We had the same problem (and really nobody could help, until ... we
recognized, that because of many processes running in the background
many mails were produced by "mmdf" - which we normally didn't use. So
we
- moved the directories "addr", "msg" and "q.local"
from /var/spool/mmdf/lock/home" to another location and
and made "7sur/spool/mail/root" empty.
It was amazing - the system was much faster and didn't crash anymore
(before that it happend sometime between 2 and 4 weeks).

regards, Matthias

Tom Parsons

unread,
Sep 13, 2001, 10:34:54 AM9/13/01
to sco...@xenitec.on.ca
Matthias Wiesner enscribed:

Absolutely and unequivocably the wrong solution.

IF your mmdf was taking up that amount of resources, then you had something
else badly wrong in your system. Fix the real cause and start by checking
for the existence of home direcories.

At one time, this could happen when someone thought they could backup
with tar and restored from the tape, only to find directories missing.

In the same vein, unknowledgeable types would find an empty /usr/sys
directory and remove it.
--
==========================================================================
Tom Parsons t...@tegan.com
==========================================================================

Karel Adams

unread,
Sep 13, 2001, 2:35:01 PM9/13/01
to

"Brian K. White" <br...@aljex.com> schreef in bericht news:pAVn7.10699$tL2.1...@news1.rdc1.nj.home.com...
>

(snip)

> > Finally: to my (limited) experience, this kind of trouble
> > is very hard to diagnosticize. If it doesn't come back
> > within a month of continuous operation, just forget it.
> > If it does, I think you ought to suspect hardware, and
> > start by swapping step for step, starting with memory
> > (easiest and cheapest), next CPU, next motherboard.
>
> while the memory might be the easiest and cheapest thing to replace, it is
> among the least likely to change with age. The most likely peice of hardware
> to look at for random problems after years of no problems, (barring the
> obvious, cpu and powersupply fans) is the power supply.

You are certainly right about the power supply.
I quite forget to mention it but is as good a candidate
as any for occasional trouble.

OTOH one ought really to start with RAM, IMHO.
It takes little effort to swap, hence costing less down time,
and if you have a good relation to your hardware suppliers
they might well lend you some for the test. And the connectors
do acquire some dust over the years, I've known this kind of
trouble to go away after swapping RAM and CPU, running
succesfully for a month of so, then swapping the originals back in.

(all of this applies to CPU's, too, of course.
And to a lesser degree to power supplies.
Swapping a motherboard is generally a hard job
so try postponing that as much as you can)

> if the power supply doesn't fix it, then all other components are about
> equally likely, but I'd start with the motherboard (it has capacitors that
> change value with age, as do some other components, but a dead cdrom drive
> generally doesn't affect the rest of the system)

The change of cap values won't, either.
Most are there as a security measure only, anyway.
Well, not really security, better power line reliability.
I once had a bet with my general manager:
all (tantalium) cap's were to be soldered out of a MB
and it would still run, though perhaps less reliably.
Guess who went home with a crate of beer that night?

> One thing I've seen, if the motherboard is old, try upping the voltage on
> the cpu by 1 or 2 tenths of a volt. I have seen this work. I don't know
> which of two most probable reasons it works. maybe a) the chip is "old" and
> has changed inside, and now needs a higher voltage to work reliably. or
> maybe b) the motherboard is "old" and when you set the jumpers for say, 2.9
> volts, it's no longer actually delivering that, so you need to set it to
> say, 3.1, to get 2.9. Either way, if this seems to fix the problem, do 2
> things immediately. 1) install a new, oversized cpu fan and heat sink, and
> use heatsink grease and remove any stickers or other foreign objects between
> the cpu face and the heat sink. 2) tell the owner of the computer to be
> ready to replace the server because this only buys you some time, it cannot
> be considered a real fix.

All of this is new to me yet seems to make good sense.
Thank you! I'm learning again!

> Another thing I've seen more than once: scsi ribbon cables that touch metal,
> where the insulation has worn away and the wire inside is just barely
> visible, meaning it sometimes but not always gets shorted to the chassis
> ground. You have to look close and careful for that, not just near sharp
> corners either. Systems run mostly ok, but get a random corrupted file here
> and there and messages in the log file and sometimes momentary server lockup
> or even unexpected reboot. Basically, If you see symptoms that make you
> think "hard drive is going south", look hard at the cable first.

Well such a story has never happened to me but it also makes a lot of sense.
If SCSI errors begin to appear for all devices on one chain
then certainly the cable is the first candidate.
With only a single SCSI HD, things are less obvious, of course.
Again, reseating the plugs might well solve the problem.

Kind greetings, Brian!
KA


0 new messages