I haven't done much troubleshooting on this one yet since I just started
seeing this one today. Is anyone else running 4.0.1-release on sun3 yet?
So far, I've noticed that the GENERIC kernel panics immediately on the
beginning of the boot sequence, and I'm getting random panics and lockups
when I roll my own. I wasn't having this issue with 3.1.1 prior to
upgrade, so I'm tentatively ruling out hardware.
Just curious if this is a known issue that I'm running into or not.
--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to news-...@muc.de
--Adam
What model and what panic messages?
I'm afraid no one has tried 4.x on 3/60 and
bus_space(9) or bus_dma(9) changes in it
might have bugs on some obio devices.
(IIRC 4.0 works on TME, which emulates VME based 3/150)
---
Izumi Tsutsui
It's a 3/60, and most of the panics have been MMU-related. I recently
reloaded 3.1.1 on it and have been using it again, but I can get 4.0.1
installed on a second drive pretty easily. I'll get that going later
tonight and report. If nothing else, I could dump a 4.0.1 GENERIC kernel
on the existing load and watch it crash.
> I'm afraid no one has tried 4.x on 3/60 and
> bus_space(9) or bus_dma(9) changes in it
> might have bugs on some obio devices.
> IIRC 4.0 works on TME, which emulates VME based 3/150)
aha. That may be part of the issue then. I'll report back in a little
bit and let you know.
Here's what I get on my 3/60 with 4.0.1:
1871128+143296 [144464+136056]=0x2307cc
starting program at 0x4000
console is ttya
trap type=0x0, code=0x155, v=0xce214454
kernel: Bus error trap
pid = 0, lid = 1, pc = 0E19CA60, ps = 2700, sfc = 1, dfc = 1
Registers:
0 1 2 3 4 5 6 7
dreg: 31DEBBAB 00000001 00001800 000015B1 0E1EFCD8 0E1F0004 0E213454 00000004
areg: CE214455 0E208010 0E2057E8 0E208008 0E1D9338 0E19CA7C 0E239F3C 0DFFFFFC
Kernel stack (0E239DE4):
239DE4: 0E160EFC 0E239E6C 00000080 00001800 000015B1 0E1EFCD8 0E1F0004 0E213454
239E04: 00000004 0E2057E8 0E208008 0E1D9338 0E19CA7C 00000000 00000000 00000000
239E24: 00000000 00000001 00000000 00000000 00000000 00000000 00000000 00000000
239E44: 00000000 00000000 00000000 00000000 0E239F3C 0E0040EC 0E239E6C 00000000
239E64: 00000155 CE214454 31DEBBAB 00000001 00001800 000015B1 0E1EFCD8 0E1F0004
239E84: 0E213454 00000004 CE214455 0E208010 0E2057E8 0E208008 0E1D9338 0E19CA7C
239EA4: 0E239F3C 0DFFFFFC 00000000 27000E19 CA60B008 1E6C0155 66FCD088 CE214454
239EC4: CE214454 0E208010 4A184A18 0E19CA66 0E19CA64 0E19CA62 CE2144FF 66FCFF0A
239EE4: 000F16EC 31DEBBAB 00000000 FFFFFFF1 FFFFFFF1 A2207003 0E239F10 00000000
239F04: 000015B1 CE214455 0E0D70E2 CE214454 0E1EFCD8 00000013 0E1F0004 00023450
239F24: 01000000 00000002 0E1EFFDC 0E1EFD0C 00000001 0EF16280 0E239F78 0E0D7B74
239F44: 0E1A6CC0 0E1F0004 00023450 0E213454 00021378 0E1D9338 0E1EFCD8 00000008
239F64: 0EF06990 00000005 00000000 0DFFFFFC 6646C8C3 0E239F8C 0E15D41E 00000001
239F84: 0E1EFCD8 0E2347CC 0E239FA8 0E0C88BE 00000008 0DFFFFFC 6646C8C3 00000000
239FA4: 00000000 00000000 0E00409C 00000000 00000000 00000000 00000000 00000000
239FC4: 00000000 00000000 00000000 00000000 00000000 00000000 00000000 00000000
panic: Bus error
Stopped in pid 0.1 () at 0xe15846e: unlk a6
db> tr
?(2700,0,e1ed1bc,e239df0,e239e54) at e15846e
?(e1bc2e0,1800,15b1,e1efcd8,e1f0004) at e0f6608
?(e239e6c,0,155,ce214454) at e160f18
?(e1a6cc0,e1f0004,23450,e213454,21378,e1d9338,e1efcd8) at e0040e6
?(1,e1efcd8,e2347cc) at e0d7b70
?(8,dfffffc,6646c8c3,0,0) at e15d418
?(0,0,0,0,0) at e0c88b8
A 4.0.1 kernel boots fine on a 3/80
--
Manuel Bouyer <bou...@antioche.eu.org>
NetBSD: 26 ans d'experience feront toujours la difference
--
:
> db> tr
> ?(2700,0,e1ed1bc,e239df0,e239e54) at e15846e
trap15()
> ?(e1bc2e0,1800,15b1,e1efcd8,e1f0004) at e0f6608
cpu_Debugger()
> ?(e239e6c,0,155,ce214454) at e160f18
panic()
> ?(e1a6cc0,e1f0004,23450,e213454,21378,e1d9338,e1efcd8) at e0040e6
trap()
> ?(1,e1efcd8,e2347cc) at e0d7b70
faultstkadj()
> ?(8,dfffffc,6646c8c3,0,0) at e15d418
ksyms_init()
> ?(0,0,0,0,0) at e0c88b8
consinit()
Ah, I'm afraid GENERIC is a bit too large for 3/60 PROM.
(current has already been hit the same problem even on 3/80)
Does it work if the GENERIC kernel is stripped?
---
Izumi Tsutsui
Attached is a config file that makes a bootable kernel from the netbsd-4
branch. I removed the VME stuff and some pseudo-devices. So size may
well be the issue.
Ok - good deal. I'll give that a try tonight.
I know that when I was able to boot/use the INSTALL kernel, I was able
to compile a kernel that would boot. Unfortunately, under heavy load (or if
I tried to run the pre-compiled pine from the 4.0 packages, for example), it
would either panic or hang. I don't know if you happen to have a machine
ready to go you can test that on, or I can certainly recreate it later
tonight.
> On Sat, Dec 27, 2008 at 02:41:18AM +0900, Izumi Tsutsui wrote:
> > Ah, I'm afraid GENERIC is a bit too large for 3/60 PROM.
> > (current has already been hit the same problem even on 3/80)
> >
> > Does it work if the GENERIC kernel is stripped?
>
> Attached is a config file that makes a bootable kernel from the netbsd-4
> branch. I removed the VME stuff and some pseudo-devices. So size may
> well be the issue.
Okay, adding "makeoptions COPTS="-Os" as -current would be
a simpler fix, but I wonder if it's worth to move bootloader
address a bit (though it doesn't work on sun2):
---
Index: stand/Makefile.inc
===================================================================
RCS file: /cvsroot/src/sys/arch/sun68k/stand/Makefile.inc,v
retrieving revision 1.13
diff -u -r1.13 Makefile.inc
--- stand/Makefile.inc 17 Sep 2006 06:15:40 -0000 1.13
+++ stand/Makefile.inc 27 Dec 2008 02:28:42 -0000
@@ -12,7 +12,11 @@
MDEC_DIR?=/usr/mdec
+.if ${MACHINE} == "sun2"
RELOC?= 240000
+.else
+RELOC?= 280000
+.endif
DEFS?= -Dsun3 -D_STANDALONE -D__daddr_t=int32_t
INCL?= -I. -I${.CURDIR} -I${.CURDIR}/../libsa -I${S}/lib/libsa -I${S}
---
---
Izumi Tsutsui
> I know that when I was able to boot/use the INSTALL kernel, I was able
> to compile a kernel that would boot. Unfortunately, under heavy load (or if
> I tried to run the pre-compiled pine from the 4.0 packages, for example), it
> would either panic or hang. I don't know if you happen to have a machine
> ready to go you can test that on, or I can certainly recreate it later
> tonight.
Hmm, I think 3/80 is stable enough. I didn't tried 4.x on it but
pkgsrc perl5 build worked on a kernel just before netbsd-4 was branched,
and upgrading 5.0_BETA and building toolchains (in ~3 days) also
worked fine.
There might be another 3/60 (or sun3, not sun3x) specific issue,
but exact panic messages always help for debugging.
---
Izumi Tsutsui
Sorry for the delay.. the holidays tend to be insanely busy for me.
I ended up blowing away my test 3.1.1 install and reinstalled 4.0.1 from
scratch. Running the INSTALL kernel and on running the 'mutt' package from
packages/4.0/sun3/All, the workstation crashes with the panic below. I
get this from nearly all the pre-compiled packages.
When I roll my own slimmed down kernel, I will get the same result from
pre-compiled packages, or the machine will lock; more often than not,
I'll get the panic. If I use a custom kernel _and_ compile things out of
pkgsrc instead of using pre-compiled, I'll often get random lockups
when under load.
I'm in the process of recompiling using Manuel's GENERIC kernel config
from last week to use as a control starting point as well. Let me know
if there's anything I can run to help with the troubleshooting. I'm not
much of a coder, but I can help whereever I can.
---
Logging in...vm_fault(0xe105238, 0x0, 0x1) -> 0xe
trap type=0x8, code=0x145, v=0x4
pid = 593, lid = 1, pc = 0E0BF7D6, ps = 2014, sfc = 1, dfc = 1
Registers:
0 1 2 3 4 5 6
7
dreg: 00000002 00000001 0000003A 00000002 00000466 0F8BFFB4 0F8BC118
0DFFBF54
areg: 0F8BC124 0F8BC000 00000000 0E10D148 0F8BC118 0F8BFFB4 0F8BFEDC
0DFFB16C
Kernel stack (0F8BFDA0):
8BFDA0: 0E0BCCF2 0F8BFE20 00000080 0000003A 00000002 00000466 0F8BFFB4
0F8BC118
8BFDC0: 0DFFBF54 00000000 0E10D148 0F8BC118 0F8BFFB4 00000000 00000000
00000001
8BFDE0: 00000000 00000000 00000000 00000002 00000000 00000000 00000008
00000000
8BFE00: 00000000 00000000 0F8BFEDC 0E0040EC 0F8BFE20 00000008 00000145
00000004
8BFE20: 00000002 00000001 0000003A 00000002 00000466 0F8BFFB4 0F8BC118
0DFFBF54
8BFE40: 0F8BC124 0F8BC000 00000000 0E10D148 0F8BC118 0F8BFFB4 0F8BFEDC
0DFFB16C
8BFE60: 00000000 20140E0B F7D6B008 1E2C0145 701FE1AB 00000004 00000004
00000002
8BFE80: 262A0004 0E0BF7DE 0E0BF7DC 0E0BF7DA E1AB2012 0004FF0B 000F1487
0000383C
8BFEA0: 00000000 0000383C 00000000 80200000 00000004 00000000 00048552
0E0BE858
8BFEC0: 0000003A 00000004 00000466 00000000 0E0BEB58 0F8BFF70 0F8BFFB4
0F8BFF2C
8BFEE0: 0E0BE870 0E10D148 00000000 00000002 0F8BC118 00000000 00000090
00000080
8BFF00: 0DFF000A 0008A4FE 0DFFBF54 0F863A84 0EE573B0 000063F8 0F8BFFB4
00000090
8BFF20: 0F863A84 0EE573B0 0F8BFF9C 0F8BFF9C 0E0BCE1E 0F8BFFB4 0F8BC040
0F8BFF70
8BFF40: 00000000 00000147 00000080 0DFF000A 0008A4FE 0DFFBF54 0DFFBF68
0009B4E0
8BFF60: 000063F8 0DFFDE10 00000000 000000B9 00000001 00000000 00000000
00000000
8BFF80: 00000000 00000000 00000000 00000010 00000000 00000000 00000000
0DFFB16C
panic: MMU fault
syncing disks... 2 1 done
dumping to dev 7,1 offset 213888
> Sorry for the delay.. the holidays tend to be insanely busy for me.
No problem.
> I ended up blowing away my test 3.1.1 install and reinstalled 4.0.1 from
> scratch. Running the INSTALL kernel and on running the 'mutt' package from
> packages/4.0/sun3/All, the workstation crashes with the panic below. I
> get this from nearly all the pre-compiled packages.
:
> Logging in...vm_fault(0xe105238, 0x0, 0x1) -> 0xe
> trap type=0x8, code=0x145, v=0x4
> pid = 593, lid = 1, pc = 0E0BF7D6, ps = 2014, sfc = 1, dfc = 1
Umm, if this is the stock INSTALL kernel from 4.0.1 distribution,
0x0e0bf7d6 is in fpu_implode() function in
sys/arch/m68k/fpe/fpu_implode.c.
Does your 3/60 actually have 68881? I.e. what does dmesg say?
FPU should be reported right after Model like this:
---
NetBSD 4.99.59 (GENERIC3X) #53: Sun Apr 13 06:12:29 JST 2008
tsutsui@mirage:/usr/src/sys/arch/sun3/compile/GENERIC3X
Model: sun3x 80
fpu: mc68882
total memory = 65536 KB
avail memory = 62224 KB
:
----
I'm not sure if FPE functions could be called even if
the machine has a real FPU, though.
> I'm in the process of recompiling using Manuel's GENERIC kernel config
> from last week to use as a control starting point as well. Let me know
> if there's anything I can run to help with the troubleshooting. I'm not
> much of a coder, but I can help whereever I can.
INSTALL kernel doesn't have ddb(4), so "trace" output on the ddb
prompt after panic with GENERIC kernel might help.
(see ddb(4) man page for details)
---
Izumi Tsutsui
Model: sun3 60
fpu: mc68882
This may be part of the problem. It appears that at some point in the
past life of this box, someone did the 68882 mod on it. I never caught
this before since it was never a problem on past SunOS or NetBSD
versions, but something here may be triggering it.
> INSTALL kernel doesn't have ddb(4), so "trace" output on the ddb
> prompt after panic with GENERIC kernel might help.
> (see ddb(4) man page for details)
Right.. that's what I'm working on next... sadly, kernel compiles on
this poor guy takes a little while. :) Due to the problem we talked
about earlier in the thread, GENERIC doesn't work on the 3/60 without a
recompile and some tweaks.
I compiled Manuel's GENERIC config to use as a baseline, and the system
booted fine, as expected. Using INSTALL, 'mutt' tends to consistently
cause a kernel panic, so I ran it again. The results are below, with
'tr' output:
---
vm_fault(0xe1b3518, 0x0, 0x1) -> 0xe
trap type=0x8, code=0x145, v=0x4
kernel: MMU fault trap
pid = 627, lid = 1, pc = 0E14C95C, ps = 2014, sfc = 1, dfc = 1
Registers:
0 1 2 3 4 5 6
7
dreg: 00000002 00000002 0000003A 00000004 00000466 0F9E9FB4 0F9E6118
00000114
areg: 0F9E6124 0E1BBE3C 00000000 0E1BBE3C 0F9E6118 0F9E9F70 0F9E9ED0
0DFFB16C
Kernel stack (0F9E9D94):
9E9D94: 0E149658 0F9E9E1C 00000080 0000003A 00000004 00000466 0F9E9FB4
0F9E6118
9E9DB4: 00000114 00000000 0E1BBE3C 0F9E6118 0F9E9F70 0E1B3518 00000000
00000001
9E9DF4: 00000008 00000000 00000000 00000000 0F9E9ED0 0E0040EC 0F9E9E1C
00000008
9E9E14: 00000145 00000004 00000002 00000002 0000003A 00000004 00000466
0F9E9FB4
9E9E34: 0F9E6118 00000114 0F9E6124 0E1BBE3C 00000000 0E1BBE3C 0F9E6118
0F9E9F70
9E9E54: 0F9E9ED0 0DFFB16C 00000000 20140E14 C95CB008 0E2C0145 701FE1AA
00000004
9E9E74: 00000004 00000002 242A0004 0E14C964 0E14C962 0E14C960 701FE1AE
0004FF0A
9E9E94: 000F1487 00003038 00004003 00003038 00000000 80200000 00000004
00000000
9E9EB4: 00049A5D 0E14B44E 0000003A 00000004 00000000 0E14BCE4 0F9E9FB4
0F9E9F24
9E9ED4: 0E14B466 0E1BBE3C 00000000 00000002 0F9E6118 0F9E6040 00000090
00000000
9E9EF4: 00000000 00000000 00000114 0EF57200 0F965A84 0F9E9FB4 0E0F49E8
00000090
9E9F14: 0EF57200 0F965A84 0F9E9FB4 0F9E9F9C 0F9E9F9C 0E14997E 0F9E9FB4
0F9E6040
9E9F34: 0F9E9F70 00000000 00000153 00000080 0DFF000A 0008A4FE 0DFFBF54
0DFFBF68
9E9F54: 0009B4A0 000063F8 0DFFDE10 0F965A84 0F9E9F7C 0E149AEA 0EF57200
00000001
9E9F74: 00000000 00000000 00000000 00000000 00000000 00000000 00000010
00000000
panic: MMU fault
Stopped in pid 627.1 (mutt) at netbsd:cpu_Debugger+0x6: unlk
a6
db> tr
cpu_Debugger(2000,8,ef57200,f9e9da0,f9e9e04) + 6
panic(e198d79,3a,4,466,f9e9fb4) + 11a
trap(f9e9e1c,8,145,4) + 244
fpu_implode(e1bbe3c,0,2,f9e6118) + ac
(f9e9fb4,f9e6040,f9e9f70) + 184e2
trap(f9e9fb4,10,0,0) + 548
fault() + 10
> Model: sun3 60
> fpu: mc68882
>
> This may be part of the problem. It appears that at some point in the
> past life of this box, someone did the 68882 mod on it. I never caught
> this before since it was never a problem on past SunOS or NetBSD
> versions, but something here may be triggering it.
According to the Sun3 Archive page:
http://www.sun3arc.org/hardpatches/3.60-tuning/fpu.phtml
68882 is software compatible with 68881 (except CPI),
so I wonder if it could cause FPU related errors.
Actually I used 68882 on 3/60 about 13 years ago (with 1.1 or 1.2?)
and it worked fine with X11 etc.
> I compiled Manuel's GENERIC config to use as a baseline, and the system
> booted fine, as expected. Using INSTALL, 'mutt' tends to consistently
> cause a kernel panic, so I ran it again.
Hmm, does only the mutt binary cause the problem?
How did you install it? From packages in ftp.NetBSD.org?
If so, in which path?
What command (including option etc.) did you try when the panic happens?
I.e. how can we reproduce it with a certain procedure?
I've tried
# pkg_add ftp://ftp.jp.netbsd.org/pub/pkgsrc/packages/NetBSD/m68k/4.0/All/mutt
# mutt
on 3/80 running -current but it looks working at least until
it shows the first screen.
(though I should also try it on TME emulating 3/150 and running 4.0.1)
> db> tr
> cpu_Debugger(2000,8,ef57200,f9e9da0,f9e9e04) + 6
> panic(e198d79,3a,4,466,f9e9fb4) + 11a
> trap(f9e9e1c,8,145,4) + 244
> fpu_implode(e1bbe3c,0,2,f9e6118) + ac
> (f9e9fb4,f9e6040,f9e9f70) + 184e2
> trap(f9e9fb4,10,0,0) + 548
> fault() + 10
This indicates:
- some userland code causes a FP related error
(fpu_implode() is in sys/arch/m68k/fpe/fpu_emulate.c and
it could only be called via fpu_emulate(), which is invoked
on "unimplemented FP instruction/data" trap)
- fpu_implode() seems to cause a NULL pointer dereference,
which should not happen even if unimplemented FP instructions
actually exist in binaries
I'm not sure unimlemented FPU trap could happen even with 68882.
One possibility is binaries compiled with -m68040 or 060 option,
but kernel shouldn't panic even in that case.
If you build a kernel without options FPU_EMULATE, I guess
it won't panic but shows "no floating point support",
as per src/sys/arch/sun3/sun3/trap.c.
---
Izumi Tsutsui
Off the top of my head...
The size of the floating point save context can be different.
If the kernel was only compiled with code for the 68881 floating point
save size, a larger 68882 context could overflow into something else.
It won't matter for executables that don't use floating point, because
of the only-save-if-used fp code.
> Hmm, does only the mutt binary cause the problem?
Mutt probably does some floating point arith for, which
makes a save-context be generated on context switch, and boom,
something dies.
To try and reproduce it just have something that does a floating
point op, and context switches -- ato[df] or printf("%f") might be
enough to trigger it.
Bolo -- Josef T. Burger
> The size of the floating point save context can be different.
> If the kernel was only compiled with code for the 68881 floating point
> save size, a larger 68882 context could overflow into something else.
>
> It won't matter for executables that don't use floating point, because
> of the only-save-if-used fp code.
Yes, I think that's the way how src/sys/arch/sun3/sun3/fpu.c
detects 68881 or 68882. I guess MI m68k code will handle
frame size by "fpu_type" variable since 3/80 (68030+68882)
uses the same fpu.c.
> Mutt probably does some floating point arith for, which
> makes a save-context be generated on context switch, and boom,
> something dies.
>
> To try and reproduce it just have something that does a floating
> point op, and context switches -- ato[df] or printf("%f") might be
> enough to trigger it.
Many simple commands (like ps(1)) uses FP ops so if FP instructions
don't work completely it's unlikely to boot up to multiuser.
(see "LC040 FPE problem" on mac68k port page)
In this case, unimplemented FP instruction trap happens
even though the machine has 68020+68882.
The trap invokes FP emulation functions, but I'm afraid
the FPE functions might not be initialized if the machine
has an FPU, so it might cause NULL pointer dereference.
We should confirm the following points:
1) whether unimplemented FP trap could happen even with 68881 or 68882
(if so we should check how it should be handled)
2) if it's vaild what code could cause the trap
3) whether 68020+68882 requires different handling from 68030+68882
---
Izumi Tsutsui
Yeah - and this is why I probably never caught this before. This
specific box has been running 2.0+ for years without issue... I only
started noticing oddities once I got 4.0.1 on it. Everything I'd read,
also, had said that it's full compatible and also faster... the last
part I've wondered about, however, since compared to other sun3s I've
used, this one doesn't seem to really react any faster on fpu-intensive
tasks.
> Hmm, does only the mutt binary cause the problem?
So far, yes - the first time I'd loaded this box, I was having problems
with most packages off of the ftp site. Now that the box has been
reloaded twice since, it appears to be just mutt. I'm compiling mutt out
of pkgsrc right now to compare, and also trying to see if I can find any
other apps that will cause the problem.
> How did you install it? From packages in ftp.NetBSD.org?
> If so, in which path?
Correct.
# env |grep PKG_PATH
PKG_PATH=ftp://ftp.netbsd.org/pub/NetBSD/packages/current-packages/NetBSD/m68k/4.0_2008Q2/All
> What command (including option etc.) did you try when the panic happens?
> I.e. how can we reproduce it with a certain procedure?
It happens consistently when I login to a IMAP server:
% mutt -f imap://mail.test.com
After login, it gathers message headers; when it begins to sort the
mailbox (user sees "Sorting mailbox..."), it immediately panics. For
reasons I don't fully understand, re-sorting local mailboxes don't cause
the panic.
> I'm not sure unimlemented FPU trap could happen even with 68882.
> One possibility is binaries compiled with -m68040 or 060 option,
> but kernel shouldn't panic even in that case.
> If you build a kernel without options FPU_EMULATE, I guess
> it won't panic but shows "no floating point support",
> as per src/sys/arch/sun3/sun3/trap.c.
Would it make sense, for troubleshooting-purposes, to recompile without
FPU_EMULATE?
An update for you from my last one...
I compiled mutt out of pkgsrc as a replacement to see what it'd do and
used all defaults (basically, "cd /usr/pkgsrc/mail/mutt && make && make
install"). The resultant binary appears to be stable; I can't seem to
crash the system with it.
I next decided to put the system under heavy load to get it to panic
again, which it had done in the past.. and it did. At the time of the
panic, I had a kernel compile running as well as 4 items out of pkgsrc.
The result:
vm_fault(0xe1b3518, 0x0, 0x1) -> 0xe
trap type=0x8, code=0x145, v=0x4
kernel: MMU fault trap
pid = 28842, lid = 1, pc = 0E14C95C, ps = 2014, sfc = 1, dfc = 1
Registers:
0 1 2 3 4 5 6
7
areg: 0F890124 0E1BBE3C 00000000 0E1BBE3C 0F890118 0F893F70 0F893ED0
0DFFDDF0
Kernel stack (0F893D94):
893D94: 0E149658 0F893E1C 00000080 0000003A 00000004 00000466 0F893FB4
0F890118
893DB4: 00000F0E 00000000 0E1BBE3C 0F890118 0F893F70 0E1B3518 00000000
00000001
893DD4: 00000000 00000001 00000000 00000000 00000000 00000002 00000000
00000000
893DF4: 00000008 00000000 00000000 00000000 0F893ED0 0E0040EC 0F893E1C
00000008
893E14: 00000145 00000004 00000002 00000002 0000003A 00000004 00000466
0F893FB4
893E34: 0F890118 00000F0E 0F890124 0E1BBE3C 00000000 0E1BBE3C 0F890118
0F893F70
893E54: 0F893ED0 0DFFDDF0 00000000 20140E14 C95CB008 0E2C0145 701FE1AA
00000004
893E74: 00000004 00000002 242A0004 0E14C964 0E14C962 0E14C960 701FE1AE
0004FF0A
893E94: 000F1487 00003038 00004006 00003038 00000000 80200000 00000004
00000000
893EB4: 0007FDEF 0E14B44E 0000003A 00000004 00000000 0E14BCE4 0F893FB4
0F893F24
893ED4: 0E14B466 0E1BBE3C 00000000 00000002 0F890118 0F890040 00000090
00000000
893EF4: 00000000 00000000 00000F0E 0EF56900 0F83DAD8 0F893FB4 0E0F49E8
00000090
893F14: 0EF56900 0F83DAD8 0F893FB4 0F893F9C 0F893F9C 0E14997E 0F893FB4
0F890040
893F34: 0F893F70 0000000D 00000000 00000000 00000000 00000000 00000000
00034024
893F54: 0000C000 00034004 0213BFA4 0F83DAD8 0F893F7C 0E149AEA 0EF56900
00000001
893F74: 00000000 00000000 00000000 00000000 00000000 00000000 00000010
00000000
panic: MMU fault
Stopped in pid 28842.1 (perl) at netbsd:cpu_Debugger+0x6:
unlk
a
6
cpu_Debugger(2000,8,ef56900,f893da0,f893e04) + 6
panic(e198d79,3a,4,466,f893fb4) + 11a
trap(f893e1c,8,145,4) + 244
fpu_implode(e1bbe3c,0,2,f890118) + ac
(f893fb4,f890040,f893f70) + 7c8bc
trap(f893fb4,10,0,0) + 548
fault() + 10
db>
> I next decided to put the system under heavy load to get it to panic
> again, which it had done in the past.. and it did. At the time of the
> panic, I had a kernel compile running as well as 4 items out of pkgsrc.
> The result:
>
> vm_fault(0xe1b3518, 0x0, 0x1) -> 0xe
> trap type=0x8, code=0x145, v=0x4
> kernel: MMU fault trap
Appearlenty this shows NULL pointer dereference again.
> Stopped in pid 28842.1 (perl) at netbsd:cpu_Debugger+0x6:
Is this perl binary from packages on ftp.NetBSD.org?
If so that m68k packages might have certain FP instructions
which can't be handled by 68020+68882, and
m68k FPE code didn't handle them properly either.
> cpu_Debugger(2000,8,ef56900,f893da0,f893e04) + 6
> panic(e198d79,3a,4,466,f893fb4) + 11a
> trap(f893e1c,8,145,4) + 244
> fpu_implode(e1bbe3c,0,2,f890118) + ac
> (f893fb4,f890040,f893f70) + 7c8bc
> trap(f893fb4,10,0,0) + 548
> fault() + 10
This shows:
unimplemented FP trap (fpfline() or fpunsupp() in locore.s)
-> fault() in src/sys/arch/m68k/m68k/trap_subr.s
-> trap() in src/sys/arch/sun3/sun3/trap.c
-> fpu_emulate() in src/sys/arch/m68k/fpe/fpu_emulate.c
-> fpu_emul_arith() (inlined into fpu_emulate() by gcc)
-> fpu_implode() called with res==NULL in fpu_implode.c
-> fpu_ftox() (inlined into fpu_implode()) with fp==NULL
offsetof(struct fpn, fp_sign) is 4, so fp->fp_sign with NULL fp
causes the reference to vaddr = 0x4.
In fpu_emulate.c:fpu_emul_arith(), I don't see an obvious code path
which could call fpu_implode() with NULL res.
Could you try this kernel (which has a debug printf in that path)?
http://www.ceres.dti.ne.jp/~tsutsui/netbsd/netbsd-sun3-FPETEST-4.0.1.gz
Index: sys/arch/m68k/fpe/fpu_emulate.c
===================================================================
RCS file: /cvsroot/src/sys/arch/m68k/fpe/fpu_emulate.c,v
retrieving revision 1.26.24.1
diff -u -r1.26.24.1 fpu_emulate.c
--- sys/arch/m68k/fpe/fpu_emulate.c 31 Mar 2007 15:40:39 -0000 1.26.24.1
+++ sys/arch/m68k/fpe/fpu_emulate.c 15 Jan 2009 14:21:39 -0000
@@ -918,6 +918,14 @@
sig = SIGILL;
} /* switch (word1 & 0x3f) */
+#if 1
+ if (res == NULL) {
+ printf("%s: FP instruction is not processed properly\n", __func__);
+ printf("%s: opcode=0x%x, word1=0x%x\n", __func__,
+ insn->is_opcode, insn->is_word1);
+ sig = SIGILL;
+ }
+#endif
if (!discard_result && sig == 0) {
fpu_implode(fe, res, FTYPE_EXT, &fpregs[regnum * 3]);
#if DEBUG_FPE
---
Izumi Tsutsui
Yeah - this is the packages binary - I didn't want to get invested in a
two-day perl compile in case the box was going to panic in the middle of
it. :)
> Could you try this kernel (which has a debug printf in that path)?
> http://www.ceres.dti.ne.jp/~tsutsui/netbsd/netbsd-sun3-FPETEST-4.0.1.gz
Sure thing... I installed it and ran the mutt package from the m68k
packages on ftp.netbsd.org... and it panic'd immediately, just as
before:
fpu_emul_arith: FP instruction is not processed properly
fpu_emul_arith: opcode=0xf200, word1=0x466
vm_fault(0xe1774e0, 0x0, 0x1) -> 0xe
trap type=0x8, code=0x145, v=0x4
kernel: MMU fault trap
pid = 678, lid = 1, pc = 0E120994, ps = 2000, sfc = 1, dfc = 1
Registers:
0 1 2 3 4 5 6
7
dreg: 00000008 00000006 00000004 00000004 00000466 0F8C3FB4 0F8C0118
000000F0
areg: 00000000 0E17F598 0E0D1554 00000000 0F8C3FB4 0F8C3F70 0F8C3ED8
0DFFB0F8
Kernel stack (0F8C3DAC):
8C3DAC: 0E11F5D4 0F8C3E34 00000080 00000004 00000004 00000466 0F8C3FB4
0F8C0118
8C3DCC: 000000F0 0E0D1554 00000000 0F8C3FB4 0F8C3F70 0E1774E0 00000000
00000001
8C3DEC: 00000000 00000001 00000000 00000000 00000000 00000002 00000000
00000000
8C3E0C: 00000008 00000000 00000000 00000000 0F8C3ED8 0E0040EC 0F8C3E34
00000008
8C3E2C: 00000145 00000004 00000008 00000006 00000004 00000004 00000466
0F8C3FB4
8C3E4C: 0F8C0118 000000F0 00000000 0E17F598 0E0D1554 00000000 0F8C3FB4
0F8C3F70
8C3E6C: 0F8C3ED8 0DFFB0F8 00000000 20000E12 0994B008 0E2C0145 670408C0
00000004
8C3E8C: 00000004 00000008 4AA80004 0E12099C 0E12099A 0E120998 670408CF
00040074
8C3EAC: 000F16EC 00000008 0E0CDDEE 00FFFFFF 000000FF 80200000 00000004
00000000
8C3ECC: 0E0CDD00 0F8C3ED4 00000004 0F8C3F24 0E121168 0E17F598 00000000
0F8C0040
8C3EEC: 00000090 00000000 00000000 00000000 000000F0 0EF17170 0F9233E4
0F8C3FB4
8C3F0C: 0E0D1554 00000090 0EF17170 0F9233E4 0F8C3FB4 0F8C3F9C 0F8C3F9C
0E11F8FA
8C3F2C: 0F8C3FB4 0F8C0040 0F8C3F70 00000000 00000155 00000080 0DFF000A
0008A4FE
8C3F4C: 0DFFBEE0 0DFFBEF4 0009B520 000063F8 0DFFDD9C 0F9233E4 0F8C3F7C
0E11FA66
8C3F6C: 0EF17170 00000001 00000000 00000000 00000000 00000000 00000000
00000000
8C3F8C: 00000010 00000000 00000000 00000000 0DFFB0F8 0E0040CC 0F8C3FB4
00000010
panic: MMU fault
Stopped in pid 678.1 (mutt) at netbsd:cpu_Debugger+0x6: unlk
a6
db> tr
cpu_Debugger(2000,8,ef17170,f8c3db8,f8c3e1c) + 6
panic(e15e641,4,4,466,f8c3fb4) + 11a
trap(f8c3e34,8,145,4) + 244
fpu_upd_fpsr(e17f598,0) + 18
fpu_emulate(f8c3fb4,f8c0040,f8c3f70) + 676
trap(f8c3fb4,10,0,0) + 548
fault() + 10
> > Could you try this kernel (which has a debug printf in that path)?
> > http://www.ceres.dti.ne.jp/~tsutsui/netbsd/netbsd-sun3-FPETEST-4.0.1.gz
>
> Sure thing... I installed it and ran the mutt package from the m68k
> packages on ftp.netbsd.org... and it panic'd immediately, just as
> before:
>
> fpu_emul_arith: FP instruction is not processed properly
> fpu_emul_arith: opcode=0xf200, word1=0x466
Okay, it looks an FDADD instruction which is available only on 040/060,
so I'm afraid all m68k 4.0 packages binaries on ftp might be built
with -m68040 or -m68060 and they won't run on 020/030 machines.
> vm_fault(0xe1774e0, 0x0, 0x1) -> 0xe
> trap type=0x8, code=0x145, v=0x4
> kernel: MMU fault trap
> pid = 678, lid = 1, pc = 0E120994, ps = 2000, sfc = 1, dfc = 1
:
> panic: MMU fault
> Stopped in pid 678.1 (mutt) at netbsd:cpu_Debugger+0x6: unlk
> a6
> db> tr
> cpu_Debugger(2000,8,ef17170,f8c3db8,f8c3e1c) + 6
> panic(e15e641,4,4,466,f8c3fb4) + 11a
> trap(f8c3e34,8,145,4) + 244
> fpu_upd_fpsr(e17f598,0) + 18
> fpu_emulate(f8c3fb4,f8c0040,f8c3f70) + 676
> trap(f8c3fb4,10,0,0) + 548
> fault() + 10
...but a kernel should not panic even in that case.
Maybe no one has tried such instructions on 020/030?
Maybe it's trivial to make those instructions cause
SIGILL properly (attached), but I'm not sure if
we should also emulate 040/060 instructions for
68881/68882 machines...
---
Index: sys/arch/m68k/fpe/fpu_emulate.c
===================================================================
RCS file: /cvsroot/src/sys/arch/m68k/fpe/fpu_emulate.c,v
retrieving revision 1.26.24.1
diff -u -r1.26.24.1 fpu_emulate.c
--- sys/arch/m68k/fpe/fpu_emulate.c 31 Mar 2007 15:40:39 -0000 1.26.24.1
+++ sys/arch/m68k/fpe/fpu_emulate.c 19 Jan 2009 11:38:47 -0000
@@ -753,8 +753,8 @@
* pointer to the result.
*/
- res = 0;
- switch (word1 & 0x3f) {
+ res = NULL;
+ switch (word1 & 0x7f) {
case 0x00: /* fmove */
res = &fe->fe_f2;
break;
@@ -910,7 +910,7 @@
discard_result = 1;
break;
- default:
+ default: /* possibly 040/060 instructions */
#ifdef DEBUG
printf("fpu_emul_arith: bad opcode=0x%x, word1=0x%x\n",
insn->is_opcode, insn->is_word1);
@@ -918,8 +918,15 @@
sig = SIGILL;
} /* switch (word1 & 0x3f) */
+ /* for sanity */
+ if (res == NULL)
+ sig = SIGILL;
+
if (!discard_result && sig == 0) {
fpu_implode(fe, res, FTYPE_EXT, &fpregs[regnum * 3]);
+
+ /* update fpsr according to the result of operation */
+ fpu_upd_fpsr(fe, res);
#if DEBUG_FPE
printf("fpu_emul_arith: %08x,%08x,%08x stored in FP%d\n",
fpregs[regnum*3], fpregs[regnum*3+1],
@@ -937,9 +944,6 @@
#endif
}
- /* update fpsr according to the result of operation */
- fpu_upd_fpsr(fe, res);
-
#if DEBUG_FPE
printf("fpu_emul_arith: FPSR = %08x, FPCR = %08x\n",
fe->fe_fpsr, fe->fe_fpcr);
---
Izumi Tsutsui
> John Carr wrote:
>
>>> Could you try this kernel (which has a debug printf in that path)?
>>> http://www.ceres.dti.ne.jp/~tsutsui/netbsd/netbsd-sun3-FPETEST-4.0.1.gz
>>
>> Sure thing... I installed it and ran the mutt package from the m68k
>> packages on ftp.netbsd.org... and it panic'd immediately, just as
>> before:
>>
>> fpu_emul_arith: FP instruction is not processed properly
>> fpu_emul_arith: opcode=0xf200, word1=0x466
>
> Okay, it looks an FDADD instruction which is available only on 040/060,
> so I'm afraid all m68k 4.0 packages binaries on ftp might be built
> with -m68040 or -m68060 and they won't run on 020/030 machines.
Maybe we should default m68k ports to compiling with -m68020-60 ?
--
David/absolute -- www.NetBSD.org: No hype required --
> Maybe we should default m68k ports to compiling with -m68020-60 ?
gcc's default seems -m68020.
-m68020-60 will disable 68881/68882 only instructions,
so it might make 040/060 a bit faster but
also might make 020/030 slower slightly?
Anyway, using -m68040 on building packages is too bad..
---
Izumi Tsutsui
> > fpu_emul_arith: FP instruction is not processed properly
> > fpu_emul_arith: opcode=0xf200, word1=0x466
>
> Okay, it looks an FDADD instruction which is available only on 040/060,
> so I'm afraid all m68k 4.0 packages binaries on ftp might be built
> with -m68040 or -m68060 and they won't run on 020/030 machines.
Now I can reproduce the same panic on 3/80 (68030+68882) running
a kernel with options FPU_EMULATE. (GENERIC3X doesn't have it)
Disabling options FPU_EMULATE (or to apply my previous patch)
could workaround it..
(though -m68040/060 binaries won't work anyway)
---
#include <stdio.h>
main()
{
printf("executing FDADD..\n");
__asm(".long 0xf2000466");
printf("done.\n");
> a...@NetBSD.org wrote:
>
>> Maybe we should default m68k ports to compiling with -m68020-60 ?
>
> gcc's default seems -m68020.
>
> -m68020-60 will disable 68881/68882 only instructions,
> so it might make 040/060 a bit faster but
> also might make 020/030 slower slightly?
It probably will slow down 020/030 slightly, but it should speed
up 040/060 by a march larger margin. I think it should be the
default for NetBSD m68k userland. Where a kernel is known to
be for a specific subset of 020/030 040/060 then it can set
more specific target options...
> Anyway, using -m68040 on building packages is too bad..
I believe that the performance penalty of -m68020 on 040 boxes
is great enough that there is real benefit for using -m68040.
Defaulting to -m68020-60 would reduce that need significantly...
--
David/absolute -- www.NetBSD.org: No hype required --
--
> > Anyway, using -m68040 on building packages is too bad..
>
> I believe that the performance penalty of -m68020 on 040 boxes
> is great enough that there is real benefit for using -m68040.
> Defaulting to -m68020-60 would reduce that need significantly...
As seen in this thread, -m68040 binaries trigger
a fatal FPU_EMULATE bug and panics on 020/030.
Even if FPE is fixed, they still cause SIGILL or wrong results
on 020/030 unless we implement FPE functions for 040/060 specific
FP instructions. (-m68020-60 is safe of course)
---
Izumi Tsutsui
> a...@NetBSD.org wrote:
>
>>> Anyway, using -m68040 on building packages is too bad..
>>
>> I believe that the performance penalty of -m68020 on 040 boxes
>> is great enough that there is real benefit for using -m68040.
>> Defaulting to -m68020-60 would reduce that need significantly...
>
> As seen in this thread, -m68040 binaries trigger
> a fatal FPU_EMULATE bug and panics on 020/030.
>
> Even if FPE is fixed, they still cause SIGILL or wrong results
> on 020/030 unless we implement FPE functions for 040/060 specific
> FP instructions. (-m68020-60 is safe of course)
I think we are in violent agtreement here - we should be
defaulting to -m68020-60 for userland and kernels which
need to be able to run on multiple processor types...
--
David/absolute -- www.NetBSD.org: No hype required --
--
Another way of looking at this is that the most common 68k boxes
are the 020 and 030 boxes. The 040s and 060s, though wonderful,
are comparatively rare.
The performance loss of instruction emulation will effect the slower
machines far more than than the faster boxes, it costs less to emulate
when a faster CPU is doing the emulation.
We ran a combined 020/030/040 userland (020 + 6888[12]) for a number
of years. The performance effects of the emulated instructions on
the 040 never made a real difference in performance -- they were
still the fast kids on the block by a large margin.
Specialized workloads which hit the emulated instructions a lot could
change that. However for ordinary system / unix / software development /
math / text processing type things there was no noticeable performance
loss with the fast machines emulating the "slow" instructions.
Bolo
> I think we are in violent agtreement here - we should be
> defaulting to -m68020-60 for userland and kernels which
> need to be able to run on multiple processor types...
Kernels are built with -msoft-float, so -m68040 (or -m68060) is
still safe even on 020/030 because no 040 specific FP insn is in kernel.
The problem is that 040 FP instructions in userland cause a kernel panic
(or SIGILL), i.e. userland binaries should use at least -m68020-60.
---
Izumi Tsutsui
> Another way of looking at this is that the most common 68k boxes
> are the 020 and 030 boxes. The 040s and 060s, though wonderful,
> are comparatively rare.
040/060 were rare among workstations (only HP had 040 ones?),
but they were/are a bit popular on (certain) consumer products.
(Amiga, Atari, X680x0... 060 support was contributed by those guys)
Nowadays, it will take more than three months for pkgsrc
bulk build (which doesn't support cross compile yet),
so even ~1% improvements on the fastest m68k (060?) machine
might also help all other slower NetBSD/m68k machines ;-)
---
Izumi Tsutsui
>>> Anyway, using -m68040 on building packages is too bad..
>>
>> I believe that the performance penalty of -m68020 on 040 boxes
>> is great enough that there is real benefit for using -m68040.
>> Defaulting to -m68020-60 would reduce that need significantly...
>
> Another way of looking at this is that the most common 68k boxes
> are the 020 and 030 boxes. The 040s and 060s, though wonderful,
> are comparatively rare.
>
> The performance loss of instruction emulation will effect the slower
> machines far more than than the faster boxes, it costs less to emulate
> when a faster CPU is doing the emulation.
>
> We ran a combined 020/030/040 userland (020 + 6888[12]) for a number
> of years. The performance effects of the emulated instructions on
> the 040 never made a real difference in performance -- they were
> still the fast kids on the block by a large margin.
>
> Specialized workloads which hit the emulated instructions a lot could
> change that. However for ordinary system / unix / software development /
> math / text processing type things there was no noticeable performance
> loss with the fast machines emulating the "slow" instructions.
I suspect the ratio has changed over time - a greater
proportion of the faster 040 and 060 machines (particularly 040
macs) have been kept than the 020 and 030, precicely because they
are faster.
The slowdown of -m68020 => -m68020-60 on 680[23]0 should be
much less than the slowdown of -m68020-60 => -m68020 on 680[46]0,
but the way to get a final conclusion is to run the test...
You should be able to put this into /etc/mk.conf
.if ${MACHINE_ARCH} == m68k
CPUFLAGS=-m68020
.endif
Then run './build.sh -U -m atari release' (can be done on a nice
fast box like a beefy AMD/Intel NetBSD or Linux system) to build
a complete m68020 optimised distribution. Install it and then time
something relevant - like building a kernel or run apachebench
against a website, whatever you care about.
Then replace -m68020 with -m68020-60 and repeat
At this point I would still prefer -m68020-60 as the default...
--
David/absolute -- www.NetBSD.org: No hype required --
--