Google Groups no longer supports new Usenet posts or subscriptions. Historical content remains viewable.
Dismiss

P4/Netburst architecture is dead

59 views
Skip to first unread message

Douglas Siebert

unread,
May 7, 2004, 4:58:10 PM5/7/04
to
Intel confirmed information The Inquirer had written several months ago,
they are cancelling the P4 (effective with the next rev that was supposed
to be out in 2005, Tejas) in favor of their P6 architecture based Pentium
M core, for both desktops and servers. I would assume Prescott will stick
around long enough for them to get a 64 bit enabled Pentium M core ready.
Since MS is now looking at Q4 for SP2/Win64, they may not miss the market
by much on this.

The talk now is of dual Pentium M cores on desktop CPUs by the end of
next year. That makes sense, they are low power enough and small enough
to work well. One wonders what will become of hyperthreading? With the
ability to do dual core CPUs in 90nm and probably quad core in 65nm, it
doesn't look like there will be much need for all the work that would be
required to retrofit Pentium M for HT, not unless they can get a larger
performance benefit than HT was for the P4. Doesn't seem to be any word
yet on what will happen to the BTX form factor they were pushing to manage
all the heat output from Prescott and especially Tejas, and the migration
to socket 775 they are attempting to jumpstart soon. I'll bet it'll be
another dead end like socket 423.

Whatever anyone may say about AMD's impact on the marketplace in terms
of market share, they surely seem to be having an affect on the market
in terms of making Intel dance so far in 2004!

IA64 aficionados may want to note that The Inquirer mentioned today that
they are now hearing rumblings that some parts of the IA64 roadmap are
getting cancelled as well. Guess we'll see if their sources for that
are as good as their sources for the P4's cancellation were proven to be.

--
Douglas Siebert dsie...@excisethis.khamsin.net

When hiring, avoid unlucky people, they are a risk to the firm. Do this by
randomly tossing out 90% of the resumes you receive without looking at them.

Robert Myers

unread,
May 7, 2004, 8:10:38 PM5/7/04
to
Douglas Siebert wrote:

<snip>

>
> Whatever anyone may say about AMD's impact on the marketplace in terms
> of market share, they surely seem to be having an affect on the market
> in terms of making Intel dance so far in 2004!
>

AMD gets the credit, or physics?

In any case, just the time for an upbeat puff piece on Craig Barrett

http://msnbc.msn.com/id/4892329/

from which it would be hard to tell that the end of the road for Intel's
current cash cow is even a pebble in Intel's shoe. "It Isn’t Just About
the PC," the article tells us, "Making microchips for the aging PC
industry isn’t a windfall anymore." Neither is making microchips for
servers, to make a reasonable inference from the chart that accompanies
that text.

To narrow the search down a bit further (yes, I've been cheating and
using Google news search on "Intel"), I tried "Intel and leakage." More
bad news not attributable to AMD, and not to heat, either: Soft errors
are apparently back in style:

http://www.eetimes.com/semi/news/showArticle.jhtml?articleID=19400052

or, if you don't want to be nailed down as to just what it is that is
causing the problems, "signal integrity:"

http://arstechnica.com/news/posts/1083010432.html

INTC up $0.49 on the day. Guess the news wasn't as bad as the Street
expected, or it hasn't yet really figured out what's going on.

RM

Andrew Reilly

unread,
May 7, 2004, 9:47:03 PM5/7/04
to
On Fri, 07 May 2004 20:58:10 +0000, Douglas Siebert wrote:

> Intel confirmed information The Inquirer had written several months ago,
> they are cancelling the P4 (effective with the next rev that was supposed
> to be out in 2005, Tejas) in favor of their P6 architecture based Pentium
> M core, for both desktops and servers.

Are you sure that the Inquirer isn't quoting one N. Maclaren, from this
very august journal?

> The talk now is of dual Pentium M cores on desktop CPUs by the end of
> next year.

Yep. That's our Nick...

--
Andrew

del cecchi

unread,
May 7, 2004, 10:24:55 PM5/7/04
to

"Robert Myers" <rmyer...@comcast.net> wrote in message
news:2AVmc.1030$iF6.152879@attbi_s02...

>
> or, if you don't want to be nailed down as to just what it is that is
> causing the problems, "signal integrity:"
>
> http://arstechnica.com/news/posts/1083010432.html
>
> INTC up $0.49 on the day. Guess the news wasn't as bad as the Street
> expected, or it hasn't yet really figured out what's going on.
>
> RM
>

Or they will save bucks by using one core design instead of two. more
EPS.

The arstechnica article was pretty specific about signal integrity from
coupling and stuff being a problem, and I didn't see any mention of soft
errors. They really aren't much of a problem in logic, although latch
design needs to be careful. And Arrays just have to waste a little
space to keep the Qcrit up or put in ECC, which they should have anyway.

Look into the effect of adjacent wires on capacitance. Now figure that
moving the aggressor doubles the effect. And add that since the wire is
very narrow, less than 200 nm, and there might be 8 levels, the
capacitance to substrate(ground) is quite low.

del cecchi


Robert Myers

unread,
May 7, 2004, 11:46:48 PM5/7/04
to
del cecchi wrote:

<snip>

>
> The arstechnica article was pretty specific about signal integrity from
> coupling and stuff being a problem, and I didn't see any mention of soft
> errors. They really aren't much of a problem in logic, although latch
> design needs to be careful. And Arrays just have to waste a little
> space to keep the Qcrit up or put in ECC, which they should have anyway.
>

You have a bundle of experience at your disposal that I don't.
Everybody seems to be experiencing unpleasant surprises moving to 90nm.
I don't have a way of evaluating whether the unpleasant surprises
people are experiencing are what should be truly regarded as surprises
or just another day at the office for people working at the smallest
scales in production.

I do have some experience with how creative people can be in making
explanations when things start to go wrong. :-).

RM

Douglas Siebert

unread,
May 8, 2004, 3:50:49 AM5/8/04
to
Robert Myers <rmyer...@comcast.net> writes:

>Douglas Siebert wrote:

><snip>

>>
>> Whatever anyone may say about AMD's impact on the marketplace in terms
>> of market share, they surely seem to be having an affect on the market
>> in terms of making Intel dance so far in 2004!
>>

>AMD gets the credit, or physics?


Depends on how you look at it. Physics played a role, but if AMD was
not as strong of a second fiddle as they are right now, and their best
effort was only say a 2400+ or so, then Intel wouldn't have any need to
push so hard for faster stuff. They could just add 100MHz every quarter
and keep well ahead of AMD and not worry about what happens when they
hit 4 or 5 GHz. So yes, I think AMD plays a part in this, by managing
to not have bankrupted themselves as many expected would happen in the
pre-Athlon days. Just like if AMD hadn't done a 64 bit part, Intel would
never have done a 64 bit x86 extension, and would have successfully
pushed everyone to IA64 over the next five years just by Moore's Law
expanding the minimum DRAM beyond the limit of a 32 bit OS on x86.

Look at it compared to MS. They upgrade the heck out of their software
when there are competitors. They may not produce quality, but they are
damn good at quantity in terms of features! Once the competition is
gone, it is just in maintenance mode. Compare all the upgrades to IE
back when Netscape was a threat, versus the last few years when they are
a tired brand name barely remembered by the masses.

Stephen Sprunk

unread,
May 8, 2004, 4:51:12 AM5/8/04
to
"Robert Myers" <rmyer...@comcast.net> wrote in message
news:2AVmc.1030$iF6.152879@attbi_s02...
> INTC up $0.49 on the day. Guess the news wasn't as bad as the Street
> expected, or it hasn't yet really figured out what's going on.

I see it as good news: Intel has reduced the number of parallel development
teams, which should result in higher EPS -- assuming they can get dual-core
and 64-bit PM chips out before the P4 line comes to an end.

Innovation is usually a positive, but in the wake of the long-running
Itanium fiasco and mounting scaling problems with NetBurst, the Street has
to be pushing for Intel to start following the strategies of other companies
(e.g. AMD and IBM) that have had more success.

S

--
Stephen Sprunk "Stupid people surround themselves with smart
CCIE #3723 people. Smart people surround themselves with
K5SSS smart people who disagree with them." --Aaron Sorkin

Yousuf Khan

unread,
May 8, 2004, 11:58:19 AM5/8/04
to
Robert Myers wrote:
> INTC up $0.49 on the day. Guess the news wasn't as bad as the
> Street expected, or it hasn't yet really figured out what's going
> on.

Likely the latter.

Yousuf Khan

Klaus Fehrle

unread,
May 8, 2004, 12:19:31 PM5/8/04
to
Firstnam...@tiscali.co.uk
"Douglas Siebert" <dsie...@excisethis.khamsin.net> schrieb im Newsbeitrag
news:c7gt91$e08$1...@narsil.avalon.net...


<snip >


> The talk now is of dual Pentium M cores on desktop CPUs by the end of
> next year. That makes sense, they are low power enough and small enough
> to work well.

Well, for the timefram it would make sense if they started working on it
quite a while ago.
If they only begin now, end of next year sounds like a very, very ambitious
target, even
just for a Dual-Core Dothan design. If Intel intended to implement 64-bit
capabilities
and a memory controller, it would appear nothing short of impossible to have
it ready
by end of next year. Let alone manufacturable for volume.

One wonders what will become of hyperthreading? With the
> ability to do dual core CPUs in 90nm and probably quad core in 65nm, it
> doesn't look like there will be much need for all the work that would be
> required to retrofit Pentium M for HT, not unless they can get a larger
> performance benefit than HT was for the P4.

Multicore CPUs allow for better parallelization than Hyperthreading, without
its downsides.
As you said, no need for all the work.

Doesn't seem to be any word
> yet on what will happen to the BTX form factor they were pushing to manage
> all the heat output from Prescott and especially Tejas, and the migration
> to socket 775 they are attempting to jumpstart soon. I'll bet it'll be
> another dead end like socket 423.

Well, as socket 775 and Prescott is all they have for the next two years,
that would
make just an average lifetime of Intel-platforms.

> Whatever anyone may say about AMD's impact on the marketplace in terms
> of market share, they surely seem to be having an affect on the market
> in terms of making Intel dance so far in 2004!

2004 it will only be slow waltz to dance for Intel.
2005 it will be Cha-cha-cha. Nothing much impressive in mss-terms, but in
terms of ASP and earnings.
2006, MSS-Foxtrott will be played when Fab30 is on capacity of one or two
mature 90nm processes.
Pace of this dance will accelerate while Fab-36 will be ramping.
2007, Fab36 could be on capacity of 5000 300mmWSPW already.
Better Intel has a competitive design again by then. Otherwise Tango will be
the dance.

KF


Mike Haertel

unread,
May 8, 2004, 3:11:47 PM5/8/04
to
On 2004-05-08, Stephen Sprunk <ste...@sprunk.org> wrote:
> Intel has reduced the number of parallel development
> teams, which should result in higher EPS

The cost of an extra design team here and there is just a drop
in Intel's bucket of $$$$, lost in the noise compared to their
infrastructure and production costs.

Douglas Siebert

unread,
May 8, 2004, 4:24:48 PM5/8/04
to
"Klaus Fehrle" <nos...@t-online.de> writes:

>> The talk now is of dual Pentium M cores on desktop CPUs by the end of
>> next year. That makes sense, they are low power enough and small enough
>> to work well.

>Well, for the timefram it would make sense if they started working on it
>quite a while ago.
>If they only begin now, end of next year sounds like a very, very ambitious
>target, even
>just for a Dual-Core Dothan design. If Intel intended to implement 64-bit
>capabilities
>and a memory controller, it would appear nothing short of impossible to have
>it ready
>by end of next year. Let alone manufacturable for volume.


That depends on how long ago they started on it. The Inquirer had rumors
of a 64 bit Pentium M skunkworks project since the beginning of the year,
and if true, who knows how long they would have been working on it. I
doubt Intel would announce such a major change in strategy without having
some fairly good ideas of the timelines involved and some initial work
done to prove them. There are always some delays you don't plan on, but
that's a problem hardly limited to Intel.

I haven't heard anything remotely concrete about an on-die memory
controller for Intel, other than claims about it being the real reason
for the 775 pins in the new socket. For the dual Pentium M, there's no
reason they'd have to do that, certainly not in their first iteration.
After all, Pentium Ms run pretty well with 1600 MB/s FSB in today's
laptops. Give them the 6.4 GB/s FSB in today's P4s, or by the end of
2005 8 GB/s or even 9.6 GB/s for DDR2-553 or DDR2-667, and I think those
two cores, even if they were twice as fast as today's top end Pentium
Ms, would be quite well fed memory wise. They don't need the bandwidth
a P4 does since they don't have the long pipeline and extra high clock
rates.

I think on die memory controllers get more interesting for Intel if/when
they get FBDRAM going to allow for plenty of memory channels for those
quad core CPUs they will probably be making in 65nm, without needing 2000
pin packages. And of course if they push FBDRAM I think AMD would surely
follow on that, if they don't end up explicitly partnering with Intel and
others to make it happen.

Felger Carbon

unread,
May 8, 2004, 5:04:01 PM5/8/04
to
"Douglas Siebert" <dsie...@excisethis.khamsin.net> wrote in message
news:c7jfmg$ah0$1...@narsil.avalon.net...

>
> I haven't heard anything remotely concrete about an on-die memory
> controller for Intel, other than claims about it being the real
reason
> for the 775 pins in the new socket. For the dual Pentium M, there's
no
> reason they'd have to do that, certainly not in their first
iteration.
> After all, Pentium Ms run pretty well with 1600 MB/s FSB in today's
> laptops. Give them the 6.4 GB/s FSB in today's P4s, or by the end
of
> 2005 8 GB/s or even 9.6 GB/s for DDR2-553 or DDR2-667, and I think
those
> two cores, even if they were twice as fast as today's top end
Pentium
> Ms, would be quite well fed memory wise. They don't need the
bandwidth
> a P4 does since they don't have the long pipeline and extra high
clock
> rates.

I believe the main point of on-die memory controllers is reduced
latency, not improved bandwidth.


> I think on die memory controllers get more interesting for Intel
if/when
> they get FBDRAM going to allow for plenty of memory channels for
those
> quad core CPUs they will probably be making in 65nm, without needing
2000
> pin packages.

You seem to believe that each core on the quad-core chip will have an
independent memory controller/channel. While I have no definite
information to the contrary, it seems unlikely. Am I missing
something here?


Stephen Sprunk

unread,
May 8, 2004, 5:23:28 PM5/8/04
to
"Felger Carbon" <fms...@jfoops.net> wrote in message
news:5Xbnc.12569$Hs1....@newsread2.news.pas.earthlink.net...

> "Douglas Siebert" <dsie...@excisethis.khamsin.net> wrote in message
> news:c7jfmg$ah0$1...@narsil.avalon.net...
> > I think on die memory controllers get more interesting for Intel
> > if/when they get FBDRAM going to allow for plenty of memory channels
> > for those quad core CPUs they will probably be making in 65nm, without
> > needing 2000 pin packages.
>
> You seem to believe that each core on the quad-core chip will have an
> independent memory controller/channel. While I have no definite
> information to the contrary, it seems unlikely. Am I missing
> something here?

AMD's K8 sports one memory controller with two core interfaces. Given
Intel's new strategy of copying AMD, that's probably what the PM will end up
with :-)

Having one controller per core also implies NUMA within a single chip, which
must bring "interesting" performance implications.

Google doesn't turn up much on FBDRAM, but it appears it'll have the same
pincount as DDR2, so I doubt we'll be getting past two channels (per chip)
any time soon, regardless of how many cores we can cram into a die.

Andy Glew

unread,
May 8, 2004, 5:55:23 PM5/8/04
to

"Douglas Siebert" <dsie...@excisethis.khamsin.net> wrote in message
news:c7gt91$e08$1...@narsil.avalon.net...

> Intel confirmed information The Inquirer had written several months ago,
> they are cancelling the P4 (effective with the next rev that was supposed
> to be out in 2005, Tejas) in favor of their P6 architecture based Pentium
> M core, for both desktops and servers.

Although, as one of the P6 architects, I might be happy to hear this,
I am not sure that all of these extrapolations are true.

The Tejas project has been in trouble for years - IMHO beginning with
when they decided not to make an out-of-order x86 chip that could also
run Itanium code (by converting the Itanium VLIW to uops that could run
OOO). Tejas lost people all over the place, such as McDermott, who
set up the Intel Austin facility, Sprangle (who left Texas for Intel Oregon,
I'm guessing when it became apparent that Tejas was not going to do
any new microarchitecture), Brad Burgess and Doug Beard (now at AMD).
I think even Marvin Denman is gone.

Also, it became obvious that Intel's Texas Design Center was not going to be
the next major processor group, when Intel acquired the Massachusetts
Alpha 21464 group that is now doing Tanglewood/Tukwila. Given the sorry
state
of Itanic, I wonder when we will hear an announcement involving THAT group.

Prescott's problems may just have been the last straw leading to the
cancellation
of Tejas.

I'm not so sure that we shoukd rule the Pentium 4 microarchitecture dead
yet,
though. Or, at least, some of its key ideas are still valid: mainly,
eliminate unnecessary
logic to make things run fast. Personally, I think the aggressive circuit
stuff for the
fireball was overkill.

The Willamette/Pentium 4 microarchitecture, IMHO, had some good ideas
mixed up with a whole slew of bad ones. Its main badness was
design-by-committee.
Many people I know at Intel said that Willamette was a failure by the CPU
architects,
rescued by heroic circuit and process engineers.

However, it would be a pity if the good ideas were dragged down by the bad
ones.

===

Moreover, exercising my paranoia: I'm not so sure that I want *MY* employer
to draw unjustified conclusions from the cancelation of Tejas.


Stephen Sprunk

unread,
May 8, 2004, 6:27:38 PM5/8/04
to
"Andy Glew" <glew2pub...@sbcglobal.net> wrote in message
news:fHcnc.46320$dJ3....@newssvr29.news.prodigy.com...

> The Willamette/Pentium 4 microarchitecture, IMHO, had some good ideas
> mixed up with a whole slew of bad ones. Its main badness was
> design-by-committee. Many people I know at Intel said that Willamette was
> a failure by the CPU architects, rescued by heroic circuit and process
engineers.
>
> However, it would be a pity if the good ideas were dragged down by the bad
> ones.

Is the trace cache something that (a) could and (b) should be retrofitted
onto the PM core, or is its decoder fast enough not to affect the critical
path? That's one of the few ideas in the P4 core I thought had a lot of
promise...

Yousuf Khan

unread,
May 8, 2004, 7:21:39 PM5/8/04
to
Douglas Siebert wrote:

> Robert Myers <rmyer...@comcast.net> writes:
>> AMD gets the credit, or physics?
>
> Depends on how you look at it. Physics played a role, but if AMD was
> not as strong of a second fiddle as they are right now, and their best
> effort was only say a 2400+ or so, then Intel wouldn't have any need
> to push so hard for faster stuff. They could just add 100MHz every
> quarter and keep well ahead of AMD and not worry about what happens
> when they hit 4 or 5 GHz. So yes, I think AMD plays a part in this,
> by managing to not have bankrupted themselves as many expected would
> happen in the pre-Athlon days. Just like if AMD hadn't done a 64 bit
> part, Intel would never have done a 64 bit x86 extension, and would
> have successfully pushed everyone to IA64 over the next five years
> just by Moore's Law expanding the minimum DRAM beyond the limit of a
> 32 bit OS on x86.

Basically what you're saying is that physics played a role because Intel was
the one who was pursuing faster and faster speeds, therefore they hit the
wall first. :-)

Yousuf Khan


Klaus Fehrle

unread,
May 8, 2004, 7:25:42 PM5/8/04
to
>I doubt Intel would announce such a major change in strategy without having
> some fairly good ideas of the timelines involved and some initial work
> done to prove them. There are always some delays you don't plan on, but
> that's a problem hardly limited to Intel.

Doug, just my gut feeling, Intels course-correction announced yesterday
sounds more
like an(other) emergency plan than like a strategy.

Put it that way: You can always drop the ball - even if you have loads of
money in the bank. ;-)

> I haven't heard anything remotely concrete about an on-die memory
> controller for Intel, other than claims about it being the real reason
> for the 775 pins in the new socket. For the dual Pentium M, there's no
> reason they'd have to do that, certainly not in their first iteration.

Hmm. I am not sure about that - from a performance point of view, that is.

> After all, Pentium Ms run pretty well with 1600 MB/s FSB in today's
> laptops. Give them the 6.4 GB/s FSB in today's P4s, or by the end of
> 2005 8 GB/s or even 9.6 GB/s for DDR2-553 or DDR2-667, and I think those
> two cores, even if they were twice as fast as today's top end Pentium
> Ms, would be quite well fed memory wise. They don't need the bandwidth
> a P4 does since they don't have the long pipeline and extra high clock
> rates.

I completely agree with Felgers comment on that.

> I think on die memory controllers get more interesting for Intel if/when
> they get FBDRAM going to allow for plenty of memory channels for those
> quad core CPUs they will probably be making in 65nm, without needing 2000
> pin packages. And of course if they push FBDRAM I think AMD would

Dsurelyis


> follow on that, if they don't end up explicitly partnering with Intel and
> others to make it happen.

As for 65nm, I wont hold my breath waiting for it. Certainly, the industry
will get there.
But not anytime soon, from today's viewpoint. As for future memory specs, I
admit
this is a somewhat blind-spot for me. So I can only see the track record of
Intels
DRAM-approaches: RDRAM in the past, DDR-2 in the present, looking like a
non-starter.
FBDRAM? Yeah, sure, if its worthwhile doing it i am confident AMD will
implement it.
Follow??? Rather leading the way to it.

KF


Andi Kleen

unread,
May 8, 2004, 8:31:23 PM5/8/04
to
"Stephen Sprunk" <ste...@sprunk.org> writes:

> Is the trace cache something that (a) could and (b) should be retrofitted
> onto the PM core, or is its decoder fast enough not to affect the critical
> path? That's one of the few ideas in the P4 core I thought had a lot of
> promise...

One thing that I always found strange about the P4 trace cache is that
it was reversing the trend towards bigger caches. Normally software
code gets more bloated and needs bigger icaches and gets them eventually.

But the trace cache is a lot smaller than a more conventional icache,
probably because it is much less die efficient. e.g. compare the 12k
entry P4 trace cache to the 64K l1 icache of K7/K8. Assuming an
average length of 3 bytes/instruction the 64K cache could in theory
hold ~21k instructions, which is nearly twice as much.

Now of course a lot of software will thrash even an 64K icache,
because they do not have a small inner loop. The only cache that has
any chance holding these codes is the big L2 or L3 cache. This means
you need an fast L2/L3 decoder anyways to perform well on these.

Given that requirement is it really that useful to have the trace
cache compared to a big L1/L2 with decoding hints? When you spend a
lot of transistors to make the decoder fast aren't the transistors
spent on the rather die inefficient trace cache then wasted?

I understand that a x86 decoder is a complex beast and likely to
contain frequency limiting speed paths, while the trace cache may
have this problem less. But I see no way around having a fast decoder to
work well on bloated software.

It will be interesting to see how big the icache or trace cache of the
P-M based Prescott successor will be.

-Andi

Stephen Sprunk

unread,
May 8, 2004, 10:07:14 PM5/8/04
to
"Andi Kleen" <fre...@alancoxonachip.com> wrote in message
news:m3brkym...@averell.firstfloor.org...

> "Stephen Sprunk" <ste...@sprunk.org> writes:
> > Is the trace cache something that (a) could and (b) should be
retrofitted
> > onto the PM core, or is its decoder fast enough not to affect the
critical
> > path? That's one of the few ideas in the P4 core I thought had a lot of
> > promise...
>
> One thing that I always found strange about the P4 trace cache is that
> it was reversing the trend towards bigger caches. Normally software
> code gets more bloated and needs bigger icaches and gets them eventually.

Smaller caches provide lower latency, which was supposedly the justification
for the anemic L1D and trace caches in Willamette/Northwood. Prescott
Doubles the L1D size and the associativity, but at the cost of increasing
the latency from one (two?) cycles to four cycles. Based on performance
results to date, this is appears to be a wash.

> But the trace cache is a lot smaller than a more conventional icache,
> probably because it is much less die efficient. e.g. compare the 12k
> entry P4 trace cache to the 64K l1 icache of K7/K8. Assuming an
> average length of 3 bytes/instruction the 64K cache could in theory
> hold ~21k instructions, which is nearly twice as much.

Do the P4 trace cache and the K8 L1 icache take the same amount of die space
or transistors? Got to keep something constant if we're going to compare...

> Now of course a lot of software will thrash even an 64K icache,
> because they do not have a small inner loop. The only cache that has
> any chance holding these codes is the big L2 or L3 cache. This means
> you need an fast L2/L3 decoder anyways to perform well on these.

On cache-busting applications, there's no easy solution; the decoders will
always be in the critical path. However, for applications that DO fit in
the icache, taking the decoders out of the critical path seems like it could
reduce the pipeline length (and thus branch penalties, etc) by a couple
stages. And, you need an L1 cache of some sort to hold the results of those
fast decoders anyways, why not store the instructions as uops instead of x86
instructions?

> It will be interesting to see how big the icache or trace cache of the
> P-M based Prescott successor will be.

Indeed.

Samuel

unread,
May 8, 2004, 11:01:38 PM5/8/04
to
NetBUST is dead... YAY!!!

Another screw up added to Intel's list RamBus, IA-64 and not going with 64
bit X86.

"Douglas Siebert" <dsie...@excisethis.khamsin.net> wrote in message
news:c7gt91$e08$1...@narsil.avalon.net...

> Intel confirmed information The Inquirer had written several months ago,
> they are cancelling the P4 (effective with the next rev that was supposed
> to be out in 2005, Tejas) in favor of their P6 architecture based Pentium
> M core, for both desktops and servers.

Does anyone have the performance numbers of Pentium M? It just doesn't make
sense to me using the same core that is meant for low end mobile computing
in high end server applications, unless they no longer care about Spec Int
numbers and focus on TPCC numbers, much like IBM and Sun, that way they can
connect 4 cores and above and support SMT per core (yes it's called SMT, not
HyperThreading). If that is the case what is going to happen to technical
applications and games that need the Powerfull spec int performance when
everyone go the other route?

This change is really a big deal, one of Intel's strength was frequency and
they used it well as a marketing tool so well. One of the reasons that the
PowerPC didn't win in the desktop market WAS the fact that it could's keep
up with frequency against the Pentiums, thus performance. Now with Intel
going to more Low power multi core approach just levels the field for other
processors to compete better, in fact, other processors already ahead in the
game of designing chips for multi-threading like AMD's Dual Core K8 and K9,
IBM's Power4, 5, 6 and Sun's Rock and Niagra.

What I'm wondering right now is how on earth a dual core Pentium M processor
can beat a dual core K8 and K9?

Greg Lindahl

unread,
May 9, 2004, 12:18:45 AM5/9/04
to
In article <m3brkym...@averell.firstfloor.org>,
Andi Kleen <fre...@alancoxonachip.com> wrote:

>But the trace cache is a lot smaller than a more conventional icache,
>probably because it is much less die efficient. e.g. compare the 12k
>entry P4 trace cache to the 64K l1 icache of K7/K8. Assuming an
>average length of 3 bytes/instruction the 64K cache could in theory
>hold ~21k instructions, which is nearly twice as much.

I don't think that's a good assumption. PathScale's compiler is the
best compiler for AMD64, and while I don't have a simulator in hand to
tell you the actual data for a benchmark like SPEC, from staring at a
lot of floating point code, I think our average is around 5
bytes/instruction -- remember that 64 bit instructions often have an
extra byte. That means that on occasion, even the aggressive K8
decoder can't issue 3 instructions in a cycle, because it can't get
enough bytes.

A trace cache never has that problem. And if it wastes some
transistors, who cares as long as it doesn't limit cycle time?
And it's decoupled from the instruction parser... so it's not
like there's any weird complexity increase...

There are also other benefits to a trace cache, such as the potential
for smaller branch bubbles. Isn't the P4 better than K8 in that area?

-- greg
(disclaimer: I work for PathScale, but don't speak for them.)

Nick Maclaren

unread,
May 9, 2004, 5:36:03 AM5/9/04
to
In article <c7jfmg$ah0$1...@narsil.avalon.net>,

Douglas Siebert <dsie...@excisethis.khamsin.net> wrote:
>
>That depends on how long ago they started on it. The Inquirer had rumors
>of a 64 bit Pentium M skunkworks project since the beginning of the year,
>and if true, who knows how long they would have been working on it. I
>doubt Intel would announce such a major change in strategy without having
>some fairly good ideas of the timelines involved and some initial work
>done to prove them. There are always some delays you don't plan on, but
>that's a problem hardly limited to Intel.

And some of us were hpyothesising such developments a year before that.
If I, as a complete outsider and not even a hardware person, can do
the relevant sums, I am absolutely sure that Intel could. Unless the
management were COMPLETELY incompetent, they would have realised that
there was a significant chance of the current problems NOT being
soluble (in time, effectively, etc.) and would have set up a backup
scheme.

If this were a sweepstake, I would bet on 1Q03 for the start.


Regards,
Nick Maclaren.

Andi Kleen

unread,
May 9, 2004, 5:35:47 AM5/9/04
to
lin...@pbm.com (Greg Lindahl) writes:

> In article <m3brkym...@averell.firstfloor.org>,
> Andi Kleen <fre...@alancoxonachip.com> wrote:
>
>>But the trace cache is a lot smaller than a more conventional icache,
>>probably because it is much less die efficient. e.g. compare the 12k
>>entry P4 trace cache to the 64K l1 icache of K7/K8. Assuming an
>>average length of 3 bytes/instruction the 64K cache could in theory
>>hold ~21k instructions, which is nearly twice as much.
>
> I don't think that's a good assumption. PathScale's compiler is the
> best compiler for AMD64, and while I don't have a simulator in hand to
> tell you the actual data for a benchmark like SPEC, from staring at a
> lot of floating point code, I think our average is around 5
> bytes/instruction -- remember that 64 bit instructions often have an

I was talking about 32bit integer code. 64bit floating point code
is totally different because it uses SSE2, which is much bigger
than normal x86 instructions. For x87 code it is true too.

But it is an interesting theory. Did Intel add the trace
cache because it was the only way to get their SSE2 FPU
fed quickly enough?

> extra byte. That means that on occasion, even the aggressive K8
> decoder can't issue 3 instructions in a cycle, because it can't get
> enough bytes.

I just ran some quick statistics on my gcc 3.2 generated 32bit
/usr/bin, and it gives on average 3.356 bytes. So my number was not
too far off for 32bit x86. Of course that is with x87 and not
particularly FP intensive.

x86-64 is different because of the REX prefix bytes, but has on
average less instructions for a given C function because of the more
registers and less spill code, which offsets this. Overall the code
length for a given C function are usually in the same league compared
to 32bit with SSE2 [all this with gcc; i don't know how your compiler or the
Microsoft compiler do ..., still waiting for the GPL release of yours]

I mention SSE2 as a special case because 64bit code using SSE2 is bigger
than x87 using 32bit code simply because SSE2 is a lot bigger than
FP stack code. Modern optimized 32bit code will use SSE2
anyways, but shipping production code often does not because running
on non SSE2 supporting CPUs is still important. 64bit code always
uses SSE2. Simple comparisons can be misleading.

It also depends on whether the code is optimized for K7/K8/P3/P-M or
P4. P4 optimized code can be shorter because the trace cache does not
need much extra alignments unlike the other CPUs (as long as you do
not need thrash it). It also can use shorter function prologues/epilogues
because it can execute lots of push and pops in parallel instead
of requiring the code size wasting tricks Opteron needs for this
to be fast.

On the other hand this all assumes that your code actually usually
hits the trace cache, which may not be the case in a lot of software.

If your code thrashes the trace cache it is possible that the P4 even
needs a different code generation strategy than what is recommended in
the Intel optimization guide, with more branch target alignment and
different function prologues. Would be interesting to benchmark this
out a bit.

> A trace cache never has that problem. And if it wastes some
> transistors, who cares as long as it doesn't limit cycle time?
> And it's decoupled from the instruction parser... so it's not
> like there's any weird complexity increase...

My point was that you need to have the fast decoder for "modern"
bloated software anyways; so why bother with the trace cache too?

> There are also other benefits to a trace cache, such as the potential
> for smaller branch bubbles. Isn't the P4 better than K8 in that area?

Hmm, let's see. The Opteron optimization guide says 1 cycle latency
for a fully predicted branch. A trace cache hit is equivalent to
"fully predicted" right? Otherwise it is not too likely for the target
to be in the trace cache, except for very small codes.

I don't have the number for a predicted branch for P4, but it is
unlikely that it is better than 1 cycle latency.

If you mean non predicted branches with branch bubbles then I do not
have any numbers; but I assume that a trace cache will not help
much with these anyways.

What I know is that the P4 gets *extremly* slow when it has to flush
the trace cache and replay the instruction stream, and it does this
far too often :-(

-Andi

Andi Kleen

unread,
May 9, 2004, 6:07:17 AM5/9/04
to
Andi Kleen <fre...@alancoxonachip.com> writes:
>
> I just ran some quick statistics on my gcc 3.2 generated 32bit
> /usr/bin, and it gives on average 3.356 bytes. So my number was not
> too far off for 32bit x86. Of course that is with x87 and not
> particularly FP intensive.
>
> x86-64 is different because of the REX prefix bytes, but has on
> average less instructions for a given C function because of the more
> registers and less spill code, which offsets this. Overall the code
> length for a given C function are usually in the same league compared
> to 32bit with SSE2 [all this with gcc; i don't know how your compiler or the
> Microsoft compiler do ..., still waiting for the GPL release of yours]

Addendum: for a gcc 3.3 compiled 64bit /usr/bin it is 3.481 bytes average.
So roughly comparable.

-Andi

Nick Maclaren

unread,
May 9, 2004, 6:17:51 AM5/9/04
to
In article <slrnc9qc7j...@ducky.net>,

Yes, indeed, if you mean the earlier stages. But, if you mean the
final stages, that is not so. Intel can't afford to put an indefinite
number of designs through chipset integration, validation and all that.
I don't know which Stephen Sprunk meant.

This was the whole lunacy about "Yamhill". The interesting question
never was whether there was a 64-bit extension project( we know there
were several), but whether any had got the "go ahead" for the later
and more expensive stages of development.


Regards,
Nick Maclaren.

ando_san

unread,
May 9, 2004, 10:03:27 AM5/9/04
to
Dear Andy,

Where did Eric go? This is a personal question. You may reply to my
address.

Best Regards,

H.Ando


"Andy Glew" <glew2pub...@sbcglobal.net> wrote in message
news:fHcnc.46320$dJ3....@newssvr29.news.prodigy.com...
>

Bengt Larsson

unread,
May 9, 2004, 12:10:27 PM5/9/04
to
"Samuel" <sam...@austin.rr.com> wrote:

>NetBUST is dead... YAY!!!
>
>Another screw up added to Intel's list RamBus, IA-64 and not going with 64
>bit X86.
>

>...

Just a quiet question: why all the Intel hatred?

Yousuf Khan

unread,
May 9, 2004, 1:20:35 PM5/9/04
to
Stephen Sprunk wrote:
> On cache-busting applications, there's no easy solution; the decoders
> will always be in the critical path. However, for applications that
> DO fit in the icache, taking the decoders out of the critical path
> seems like it could reduce the pipeline length (and thus branch
> penalties, etc) by a couple stages. And, you need an L1 cache of
> some sort to hold the results of those fast decoders anyways, why not
> store the instructions as uops instead of x86 instructions?

How would you eliminate branching even in a trace cache? Isn't branching
just as much of an atomic instruction as any other micro-op?

Yousuf Khan


Yousuf Khan

unread,
May 9, 2004, 1:30:42 PM5/9/04
to
Samuel <sam...@austin.rr.com> wrote:
> Does anyone have the performance numbers of Pentium M? It just
> doesn't make sense to me using the same core that is meant for low
> end mobile computing in high end server applications, unless they no
> longer care about Spec Int numbers and focus on TPCC numbers, much
> like IBM and Sun, that way they can connect 4 cores and above and
> support SMT per core (yes it's called SMT, not HyperThreading). If
> that is the case what is going to happen to technical applications
> and games that need the Powerfull spec int performance when everyone
> go the other route?

In general, it would seem to me that server chips in the x86 world are just
bigger-cache versions of desktop and mobile chips. In fact, in the low-power
blade server world, the Pentium M is already used.

> This change is really a big deal, one of Intel's strength was
> frequency and they used it well as a marketing tool so well. One of
> the reasons that the PowerPC didn't win in the desktop market WAS the
> fact that it could's keep up with frequency against the Pentiums,
> thus performance. Now with Intel going to more Low power multi core
> approach just levels the field for other processors to compete
> better, in fact, other processors already ahead in the game of
> designing chips for multi-threading like AMD's Dual Core K8 and K9,
> IBM's Power4, 5, 6 and Sun's Rock and Niagra.

Actually the main reason that PowerPC didn't win in the desktop market was
not because of performance, but because it simply didn't run x86 software,
and especially because it was cubbyholed into the small Macintosh world. It
wouldn't have mattered if PowerPC was an order of magnitude faster than any
x86 chip, the great mass of software was concentrated in the x86 world --
unless PowerPC could run that stuff, then it wasn't going anywhere.

> What I'm wondering right now is how on earth a dual core Pentium M
> processor can beat a dual core K8 and K9?

Unless Pentium-M grows an onboard memory controller in a few months, the
only option Intel has to hope to stay on pace with AMD is to add tons of L2
cache. Of course AMD can do the same, but it can afford to add less cache
than Intel to stay on par, due to its onboard memory controller.

Yousuf Khan


Yousuf Khan

unread,
May 9, 2004, 1:40:44 PM5/9/04
to
Andi Kleen <fre...@alancoxonachip.com> wrote:
> I was talking about 32bit integer code. 64bit floating point code
> is totally different because it uses SSE2, which is much bigger
> than normal x86 instructions. For x87 code it is true too.
>
> But it is an interesting theory. Did Intel add the trace
> cache because it was the only way to get their SSE2 FPU
> fed quickly enough?

Those 128-bit registers of SSE and those 80-bit registers of x87 are fed by
the D-cache aren't they? So why should SSE or x87 FPU instructions have a
larger footprint in an instruction or trace cache?

BTW, is there an officially accepted general term to describe either a trace
cache or an instruction cache? I would think just calling them both
instruction caches should be sufficient?

>> extra byte. That means that on occasion, even the aggressive K8
>> decoder can't issue 3 instructions in a cycle, because it can't get
>> enough bytes.
>
> I just ran some quick statistics on my gcc 3.2 generated 32bit
> /usr/bin, and it gives on average 3.356 bytes. So my number was not
> too far off for 32bit x86. Of course that is with x87 and not
> particularly FP intensive.
>
> x86-64 is different because of the REX prefix bytes, but has on
> average less instructions for a given C function because of the more
> registers and less spill code, which offsets this. Overall the code
> length for a given C function are usually in the same league compared
> to 32bit with SSE2 [all this with gcc; i don't know how your compiler
> or the Microsoft compiler do ..., still waiting for the GPL release
> of yours]

I can see the extra registers of AMD64 would likely reduce instruction
lengths, because you eliminate that bad habit in x86 code of doing
operations directly in memory rather than in a register because of a lack of
available registers.

Yousuf Khan


Terje Mathisen

unread,
May 9, 2004, 2:23:32 PM5/9/04
to
Yousuf Khan wrote:

> Andi Kleen <fre...@alancoxonachip.com> wrote:
>>x86-64 is different because of the REX prefix bytes, but has on
>>average less instructions for a given C function because of the more
>>registers and less spill code, which offsets this. Overall the code
>>length for a given C function are usually in the same league compared
>>to 32bit with SSE2 [all this with gcc; i don't know how your compiler
>>or the Microsoft compiler do ..., still waiting for the GPL release
>>of yours]
>
> I can see the extra registers of AMD64 would likely reduce instruction
> lengths, because you eliminate that bad habit in x86 code of doing
> operations directly in memory rather than in a register because of a lack of
> available registers.

Using x86 operate_from_mem style instructions is almost always fine, it
saves both registers, instructions and code space.

Read-modify-write is another case totally. :-(

Terje

--
- <Terje.M...@hda.hydro.com>
"almost all programming can be viewed as an exercise in caching"

Daniel Gustafsson

unread,
May 9, 2004, 2:25:05 PM5/9/04
to
"Samuel" <sam...@austin.rr.com> wrote in message news:<mahnc.71878$NR5....@fe1.texas.rr.com>...

> "Douglas Siebert" <dsie...@excisethis.khamsin.net> wrote in message
> news:c7gt91$e08$1...@narsil.avalon.net...
> > Intel confirmed information The Inquirer had written several months ago,
> > they are cancelling the P4 (effective with the next rev that was supposed
> > to be out in 2005, Tejas) in favor of their P6 architecture based Pentium
> > M core, for both desktops and servers.
>
> Does anyone have the performance numbers of Pentium M? It just doesn't make
> sense to me using the same core that is meant for low end mobile computing
> in high end server applications, unless they no longer care about Spec Int
> numbers and focus on TPCC numbers, much like IBM and Sun, that way they can
> connect 4 cores and above and support SMT per core (yes it's called SMT, not
> HyperThreading). If that is the case what is going to happen to technical
> applications and games that need the Powerfull spec int performance when
> everyone go the other route?

The Pentium M has similarities with Pentium 3 and those did not had
bad SPECint numbers. Besides, the current Pentium M's has such low
power usage they may currently be clocked down to meet specific power
requirements. Intel may have figured out that by powering up the
current Pentium M cores a bit and give them a year of development then
they may be very competetive.

Whether they will add SMT to these cores is I think still a secret.

( Although I get what you say, Sun does not focus on TPCC numbers ;) )

Regards
Daniel Gustafsson

Yousuf Khan

unread,
May 9, 2004, 3:01:51 PM5/9/04
to
Terje Mathisen <terje.m...@hda.hydro.com> wrote:

> Yousuf Khan wrote:
>> I can see the extra registers of AMD64 would likely reduce
>> instruction lengths, because you eliminate that bad habit in x86
>> code of doing operations directly in memory rather than in a
>> register because of a lack of available registers.
>
> Using x86 operate_from_mem style instructions is almost always fine,
> it saves both registers, instructions and code space.
>
> Read-modify-write is another case totally. :-(

You're right, they shouldn't be much of a problem either. I was actually
thinking of those x86 operand-embedded-in-instruction style instructions.

Yousuf Khan


Terje Mathisen

unread,
May 9, 2004, 4:03:28 PM5/9/04
to
Yousuf Khan wrote:

Huh?

Which opcodes would that be?

Those with immediate data? Implicit registers.

The only thing that comes close afaik would seem to be a couple of the
MMX/SSE permute operations where the actual operation to perform can be
a runtime variable?

Bengt Larsson

unread,
May 9, 2004, 5:10:18 PM5/9/04
to
Douglas Siebert <dsie...@excisethis.khamsin.net> wrote:

>IA64 aficionados may want to note that The Inquirer mentioned today that
>they are now hearing rumblings that some parts of the IA64 roadmap are
>getting cancelled as well. Guess we'll see if their sources for that
>are as good as their sources for the P4's cancellation were proven to be.

I guess I'm one of those aficionados, sort of. On purely technical
grounds I'd prefer IA-64 over x86-64. As a programmer, I don't look
forward to 20 more years of x86. IA-64 is a bit more modern, more
RISC-like, has more registers and so on.

It's true that Intel would have a monopoly on IA-64 but there are two
comments one can make on that:

1. Intel want to compete with IBM on processors for the high end, and
IBM show no signs of stopping, so there will be competition.

2. If Intel were to go too far and milk their monopoly too much there
is a very simple remedy: force them to license IA-64 to someone. The
Anti-trust remedy writes itself, unlike cases vs. Microsoft and IBM.

I have nothing against PowerPC, but it's domineered/dominated by IBM.
There will never realistically be competition on PowerPC-based systems
vs IBM.

The ideal would be processors from Intel and AMD, systems from other
people (like Dell, HP, IBM...), operating systems from yet other
people (like Linux) etc. All to promote competition.

Samuel

unread,
May 9, 2004, 5:23:52 PM5/9/04
to

"Bengt Larsson" <bengt...@telia.NOSPAMcom> wrote in message
news:tqls90lb2bs8e6h1i...@text.giganews.com...

>
> Just a quiet question: why all the Intel hatred?

I guess it shows, huh? ;)

Samuel

unread,
May 9, 2004, 5:42:35 PM5/9/04
to

"Yousuf Khan" <news.tal...@spamgourmet.com> wrote in message
news:6Vtnc.8296$pp....@news04.bloor.is.net.cable.rogers.com...

> Samuel <sam...@austin.rr.com> wrote:
> Actually the main reason that PowerPC didn't win in the desktop market was
> not because of performance, but because it simply didn't run x86 software,

Sure there is no arguing that, but there was a time when G4 PowerPC was
lagging WAY behind in Frequency and Spec Int performance, particularly with
the G3 and the G4. At one time the Pentium4 was 2-3X the frequency of the G4
which got Steve Jobs very worried. Then Jobs pulled the G5 development from
Motorola granted it to IBM.

Anyhow, with IBM clocking the PPC976 @ 3.5 GHz and Intel falling back on a 2
GHz core, I see Apple would hardly be worried about frequency wars in the
next couple of year at least. Adding to that the Power6 core will be a speed
demon which is a 180 degree shift from the Power4, Power5 approach, Apple
would be even more comfortable with frequency in the future (assuming IBM
don't screw up with Power6 like Intel screwed up with NetBUST)


Bengt Larsson

unread,
May 9, 2004, 5:44:06 PM5/9/04
to
"Samuel" <sam...@austin.rr.com> wrote:

Yeah it does. But why?

Samuel

unread,
May 9, 2004, 6:04:13 PM5/9/04
to

"Daniel Gustafsson" <dan...@mimer.se> wrote in message

> The Pentium M has similarities with Pentium 3 and those did not had
> bad SPECint numbers.

Yes and I was a very good wrestler when I was in High School.

PIIIs were great at some point, probably even before the first Athlon came
out. The PIIIs were competing very well with K6 and the four stage pipelined
G3 PPC. When the Athlon came out it took the lead in Spec int performance
and the PIII cores started to show their age. Today it's a different world,
AMD is doing a great job and they are ahead with multi core design with
cores that are superior to the good old PIIIs.

I don't really see how the "Bach to the Future" approach would make Intel
processors competative in performance with AMD.

> ( Although I get what you say, Sun does not focus on TPCC numbers ;) )

From reading about the Rock, Niagra and follow ups it seems like it's all
they care about now.

>
> Regards
> Daniel Gustafsson


Samuel

unread,
May 9, 2004, 6:06:07 PM5/9/04
to

"Bengt Larsson" <bengt...@telia.NOSPAMcom> wrote in message
> Yeah it does. But why?

That should be a topic of it's own :)

Norbert Juffa

unread,
May 9, 2004, 6:14:54 PM5/9/04
to

"Andi Kleen" <fre...@alancoxonachip.com> wrote in message news:m37jvmx...@averell.firstfloor.org...

> lin...@pbm.com (Greg Lindahl) writes:
>
> > In article <m3brkym...@averell.firstfloor.org>,
> > Andi Kleen <fre...@alancoxonachip.com> wrote:
> >
> >>But the trace cache is a lot smaller than a more conventional icache,
> >>probably because it is much less die efficient. e.g. compare the 12k
> >>entry P4 trace cache to the 64K l1 icache of K7/K8. Assuming an
> >>average length of 3 bytes/instruction the 64K cache could in theory
> >>hold ~21k instructions, which is nearly twice as much.
> >
> > I don't think that's a good assumption. PathScale's compiler is the
> > best compiler for AMD64, and while I don't have a simulator in hand to
> > tell you the actual data for a benchmark like SPEC, from staring at a
> > lot of floating point code, I think our average is around 5
> > bytes/instruction -- remember that 64 bit instructions often have an
>
> I was talking about 32bit integer code. 64bit floating point code
> is totally different because it uses SSE2, which is much bigger
> than normal x86 instructions. For x87 code it is true too.
[...]

It's not clear to me what "for x87 code it is true too" refers to.
Could you clarify please?

When I last checked several years ago, 32-bit integer code on x86
took up approximately 3.8 bytes/instruction. IIRC, x87 intensive
code had a _shorter_ average instruction length due to its tight
encoding, where one register operand (ST0) is implicit. Also, x87
code in general did not need prefixes (e.g. 0x0f, 0x66 etc). FWIW,
16-bit x86 integer code (ca 1995) ran about 2.8 bytes/instruction.

Of course, average instruction length is somewhat a function of
instruction mix issues by a particular compiler, e.g. use of
load-execute instructions versus separate load and reg-to-reg
instruction.

-- Norbert


Bengt Larsson

unread,
May 9, 2004, 6:18:00 PM5/9/04
to
"Samuel" <sam...@austin.rr.com> wrote:

>Anyhow, with IBM clocking the PPC976 @ 3.5 GHz and Intel falling back on a 2
>GHz core, I see Apple would hardly be worried about frequency wars in the
>next couple of year at least. Adding to that the Power6 core will be a speed
>demon which is a 180 degree shift from the Power4, Power5 approach, Apple
>would be even more comfortable with frequency in the future (assuming IBM
>don't screw up with Power6 like Intel screwed up with NetBUST)

I never bought the reasoning that NetBurst was high-frequency for
marketing. There were technical papers that showed a performance
advantage up to 50 stages (the P4 had 20). If Intel wanted to they
could have marketed NetBurst as capable of adding integers at 6.4 GHz,
but they never bothered. It wouldn't have worked as marketing. It
would have been true, in a sense, but obviously not for any real
applications - and people would have noticed.

AMD bypassed the whole thing with their xxxx+ marketing anyway.

Why people look for conspiracy explanations when there are natural
explanations I will never understand. Are conspiracy explanations more
interesting?

Samuel

unread,
May 9, 2004, 6:39:21 PM5/9/04
to

"Bengt Larsson" <bengt...@telia.NOSPAMcom> wrote in message
news:jtat90paet5l02li3...@text.giganews.com...

> I never bought the reasoning that NetBurst was high-frequency for
> marketing. There were technical papers that showed a performance
> advantage up to 50 stages (the P4 had 20). If Intel wanted to they
> could have marketed NetBurst as capable of adding integers at 6.4 GHz,
> but they never bothered. It wouldn't have worked as marketing. It
> would have been true, in a sense, but obviously not for any real
> applications - and people would have noticed.

I'm sure frequency was not the ONLY target with P4 design, they also wanted
to get performacne out of driving frequency and they thought they can scale
frequency better with P4 and follow ups that they can do with a P3 like
architecture. However, it's hard to deny that there was an abvious frequency
wars in the desktop market and the average consumor only looks at the mega
Hz numbers when shopping for a new PC for Christmas. It seemed to me and to
many people at the time that they were playing this as a marketing tool.

> AMD bypassed the whole thing with their xxxx+ marketing anyway.

That was a marketing genius, wasn't it?

> Why people look for conspiracy explanations when there are natural
> explanations I will never understand. Are conspiracy explanations more
> interesting?

I guess I'm a conspiracy theorist Intel hater. LOL


Bengt Larsson

unread,
May 9, 2004, 6:54:56 PM5/9/04
to
"Samuel" <sam...@austin.rr.com> wrote:

>"Bengt Larsson" <bengt...@telia.NOSPAMcom> wrote in message
>news:jtat90paet5l02li3...@text.giganews.com...

>I guess I'm a conspiracy theorist Intel hater. LOL

You said it.

Stephen Sprunk

unread,
May 9, 2004, 7:17:18 PM5/9/04
to
"Bengt Larsson" <bengt...@telia.NOSPAMcom> wrote in message
news:jtat90paet5l02li3...@text.giganews.com...
> I never bought the reasoning that NetBurst was high-frequency for
> marketing. There were technical papers that showed a performance
> advantage up to 50 stages (the P4 had 20). If Intel wanted to they
> could have marketed NetBurst as capable of adding integers at 6.4 GHz,
> but they never bothered. It wouldn't have worked as marketing. It
> would have been true, in a sense, but obviously not for any real
> applications - and people would have noticed.

Did you miss when Intel demonstrated 10GHz ALUs last year? The P4 strategy
was always based on marketing GHz over performance, but their recent
setbacks in increasing clock speed have caused them to kill the entire
product.

S

--
Stephen Sprunk "Stupid people surround themselves with smart
CCIE #3723 people. Smart people surround themselves with
K5SSS smart people who disagree with them." --Aaron Sorkin

Stefan Monnier

unread,
May 9, 2004, 7:35:47 PM5/9/04
to
> Did you miss when Intel demonstrated 10GHz ALUs last year? The P4 strategy
> was always based on marketing GHz over performance, but their recent
> setbacks in increasing clock speed have caused them to kill the entire
> product.

There's no question that Intel's marketing has played pretty heavily the
Ghz song. But this newsgroup is not about marketing, so the real question
is whether the marketing drove the microarchitecture or not.

I personally don't believe it did. The P4 is a pretty good performer if
you ask me, so there seem to have been valid technical reasons to go
this route.


Stefan

Yousuf Khan

unread,
May 9, 2004, 8:13:40 PM5/9/04
to
Bengt Larsson <bengt...@telia.NOSPAMcom> wrote:
> It's true that Intel would have a monopoly on IA-64 but there are two
> comments one can make on that:
>
> 1. Intel want to compete with IBM on processors for the high end, and
> IBM show no signs of stopping, so there will be competition.

Everybody else gave up the ghost at least five years ago, at the mere
thought of having to compete against IA64 before there was even a working
IA64, and instead embraced it wholeheartedly -- bye-bye MIPS, Alpha,
PA-RISC, etc. What you're merely saying is that we don't have to worry about
lack of competition because there were at least a few corporations that were
not stupid enough to give up their own processor architectures. What if they
had _all_ decided to give up at the mere mention of competition from Intel?

> 2. If Intel were to go too far and milk their monopoly too much there
> is a very simple remedy: force them to license IA-64 to someone. The
> Anti-trust remedy writes itself, unlike cases vs. Microsoft and IBM.

It's that simple, huh? Intel has managed to make life very difficult for its
x86 competitors to sell their processors to OEMs, and yet Intel manages to
keep away from the anti-trust authorities, because it never ever writes down
its threats.

> I have nothing against PowerPC, but it's domineered/dominated by IBM.
> There will never realistically be competition on PowerPC-based systems
> vs IBM.

There was Motorola too.

> The ideal would be processors from Intel and AMD, systems from other
> people (like Dell, HP, IBM...), operating systems from yet other
> people (like Linux) etc. All to promote competition.

Yes, that would be the ideal. However that's what's happening already, but
in a lopsided fashion.

Yousuf Khan


Yousuf Khan

unread,
May 9, 2004, 8:03:39 PM5/9/04
to
Terje Mathisen <terje.m...@hda.hydro.com> wrote:
> Yousuf Khan wrote:
>> You're right, they shouldn't be much of a problem either. I was
>> actually thinking of those x86 operand-embedded-in-instruction style
>> instructions.
>
> Huh?
>
> Which opcodes would that be?
>
> Those with immediate data? Implicit registers.

Yes, the immediate data would be the one. Things such as:

mov eax, 0x00000001

Where that final 32-bit number "1" would occupy a full 4 bytes in the
instruction stream.

But remind me, which instruction forms are the implicit registers?

> The only thing that comes close afaik would seem to be a couple of the
> MMX/SSE permute operations where the actual operation to perform can
> be a runtime variable?

I wasn't really talking about the SIMD instructions, just the good old
fashioned x86 ones. Not familiar enough with the newer instructions. I used
to program in assembly back in the 386 days.

Yousuf Khan


Yousuf Khan

unread,
May 9, 2004, 8:23:42 PM5/9/04
to
Samuel <sam...@austin.rr.com> wrote:
> "Yousuf Khan" <news.tal...@spamgourmet.com> wrote in message
>> Actually the main reason that PowerPC didn't win in the desktop
>> market was not because of performance, but because it simply didn't
>> run x86 software,
>
> Sure there is no arguing that, but there was a time when G4 PowerPC
> was lagging WAY behind in Frequency and Spec Int performance,
> particularly with the G3 and the G4. At one time the Pentium4 was
> 2-3X the frequency of the G4 which got Steve Jobs very worried. Then
> Jobs pulled the G5 development from Motorola granted it to IBM.

Jobs needn't have worried. Just like there was no way PC people were ever
going to switch to a Macintosh processor, no matter what the performance,
similarly there was no way that Macintosh people would've ever switched to a
PC processor. You're just stuck in the environment that you're stuck in.

> Anyhow, with IBM clocking the PPC976 @ 3.5 GHz and Intel falling back
> on a 2 GHz core, I see Apple would hardly be worried about frequency
> wars in the next couple of year at least. Adding to that the Power6
> core will be a speed demon which is a 180 degree shift from the
> Power4, Power5 approach, Apple would be even more comfortable with
> frequency in the future (assuming IBM don't screw up with Power6 like
> Intel screwed up with NetBUST)

IBM isn't at 3.5 Ghz yet, and it would seem rather optimistic that they'll
even touch that speed even with 90nm. What is the pipeline length of the PPC
97x? About 10 stages? I'd say AMD would be closer to 3.5 Ghz with its 12
stage pipeline than the PPC. AMD is already nearing 2.5 Ghz with a 130nm
process.

Sure clock frequency isn't everything, but if you are going to make it to
certain frequency, the pipeline has to accomodate it.

Yousuf Khan


Yousuf Khan

unread,
May 9, 2004, 8:23:42 PM5/9/04
to
Bengt Larsson <bengt...@telia.NOSPAMcom> wrote:
> AMD bypassed the whole thing with their xxxx+ marketing anyway.

Which was fortunate for them. The previous attempt at equating true
performance against an Intel processor ended up hurting the manufacturer
that tried to pass it off (i.e. Cyrix and even AMD to a certain extent).

Yousuf Khan


Yousuf Khan

unread,
May 9, 2004, 8:23:43 PM5/9/04
to
Samuel <sam...@austin.rr.com> wrote:
> PIIIs were great at some point, probably even before the first Athlon
> came out. The PIIIs were competing very well with K6 and the four
> stage pipelined G3 PPC. When the Athlon came out it took the lead in
> Spec int performance and the PIII cores started to show their age.
> Today it's a different world, AMD is doing a great job and they are
> ahead with multi core design with cores that are superior to the good
> old PIIIs.
>
> I don't really see how the "Bach to the Future" approach would make
> Intel processors competative in performance with AMD.

Well, it's all Intel has got right now at the moment. And it's not like as
if it is a completely unmodified P3 core, it's got all of that wonderful
power savings feature.

Yousuf Khan


Douglas Siebert

unread,
May 10, 2004, 12:01:05 AM5/10/04
to
"Stephen Sprunk" <ste...@sprunk.org> writes:

>Google doesn't turn up much on FBDRAM, but it appears it'll have the same
>pincount as DDR2, so I doubt we'll be getting past two channels (per chip)
>any time soon, regardless of how many cores we can cram into a die.


Search under "fully buffered DRAM" and you'll probably find more...

FBDRAM uses far fewer pins than DDR2, but plans are that it will use the
DDR2 socket (at least initially) for cost/compatibility reasons. FBDRAM
uses existing DRAM chips. It could use DDR2 or some future thing like
DDR3 without changing the memory controller -- great for on die memory
controllers! A FBDRAM DIMM would look the same with the addition of a
single chip (the buffer) that makes it a FBDRAM DIMM. It is probably
doable to set things up so that you could have a motherboard that supported
FBDRAM but could detect if you plugged regular DIMMs in and use them
instead.

IIRC there are only 69 pins required per FBDRAM controller, so you can
support more than twice the channels with FBDRAM, and it allows for a
larger number of DIMMs per channel.

--
Douglas Siebert dsie...@excisethis.khamsin.net

When hiring, avoid unlucky people, they are a risk to the firm. Do this by
randomly tossing out 90% of the resumes you receive without looking at them.

Douglas Siebert

unread,
May 10, 2004, 12:10:10 AM5/10/04
to
"Samuel" <sam...@austin.rr.com> writes:

>Does anyone have the performance numbers of Pentium M? It just doesn't make
>sense to me using the same core that is meant for low end mobile computing
>in high end server applications, unless they no longer care about Spec Int
>numbers and focus on TPCC numbers, much like IBM and Sun, that way they can
>connect 4 cores and above and support SMT per core (yes it's called SMT, not
>HyperThreading). If that is the case what is going to happen to technical
>applications and games that need the Powerfull spec int performance when
>everyone go the other route?


Intel doesn't want to sell you a desktop CPU stuff for stuff that needs
"Powerfull spec int performance", they want you to buy Itanium for that.

Douglas Siebert

unread,
May 10, 2004, 12:16:48 AM5/10/04
to
"Felger Carbon" <fms...@jfoops.net> writes:

>"Douglas Siebert" <dsie...@excisethis.khamsin.net> wrote in message

>news:c7jfmg$ah0$1...@narsil.avalon.net...
>>
>> I haven't heard anything remotely concrete about an on-die memory
>> controller for Intel, other than claims about it being the real
>reason
>> for the 775 pins in the new socket. For the dual Pentium M, there's
>no
>> reason they'd have to do that, certainly not in their first
>iteration.
>> After all, Pentium Ms run pretty well with 1600 MB/s FSB in today's
>> laptops. Give them the 6.4 GB/s FSB in today's P4s, or by the end
>of
>> 2005 8 GB/s or even 9.6 GB/s for DDR2-553 or DDR2-667, and I think
>those
>> two cores, even if they were twice as fast as today's top end
>Pentium
>> Ms, would be quite well fed memory wise. They don't need the
>bandwidth
>> a P4 does since they don't have the long pipeline and extra high
>clock
>> rates.

>I believe the main point of on-die memory controllers is reduced
>latency, not improved bandwidth.


I believe the post I was responding to was implying that a dual core
Pentium M would demand twice as much bandwidth and therefore suffer
performance-wise. But you are correct, on die reduces latency, though
that reduced latency does give a small bandwidth benefit as part of
the deal.


>> I think on die memory controllers get more interesting for Intel
>if/when
>> they get FBDRAM going to allow for plenty of memory channels for
>those
>> quad core CPUs they will probably be making in 65nm, without needing
>2000
>> pin packages.

>You seem to believe that each core on the quad-core chip will have an
>independent memory controller/channel. While I have no definite
>information to the contrary, it seems unlikely. Am I missing
>something here?


No, I don't believe that at all. The cores will share the same memory
controllers. But since a P4 performs better with a 800 MHz FSB than
with 400 MHz, and an A64 better with dual channels rather than one, it
stands to reason that a single core with X amount of memory bandwidth
will perform worse than dual cores with X amount of memory bandwidth.
And it only gets worse with quad cores. Clearly it wouldn't be cost
effective to provide quad channel DDR (let alone the even more pin hungry
DDR2 which will be mainstream when 65 nm stuff comes out in 2006) But
with FBDRAM, you need fewer pins for quad channels than dual channel DDR
requires today.

Bill Todd

unread,
May 10, 2004, 1:19:34 AM5/10/04
to

"Yousuf Khan" <news.tal...@spamgourmet.com> wrote in message
news:vFznc.4839$n7P1...@twister01.bloor.is.net.cable.rogers.com...

... Things such as:

>
> mov eax, 0x00000001
>
> Where that final 32-bit number "1" would occupy a full 4 bytes in the
> instruction stream.

Oh, my - it's been a *long* time. But though the details have faded from
memory ISTR that the x86 instruction set provides mechanisms for compressing
immediate operands that will fit into 1 or 2 bytes and zero- or
sign-extending them to full destination width (leaving aside explicit
mechanisms such as MOVZX and MOVSX - which may not take immediate source
operands - multi-instruction sequences and creative use of instructions such
as LEA).

- bill

Andi Kleen

unread,
May 10, 2004, 2:00:26 AM5/10/04
to
"Norbert Juffa" <ju...@earthlink.net> writes:

> It's not clear to me what "for x87 code it is true too" refers to.
> Could you clarify please?

x87 code is much shorter than SSE2 code, especially 64bit SSE2 code
which has additional REX prefixes.

What I attempted to say was that you should not compare 32bit-with-x87
to 64bit-with-SSE2, but 32bit-with-SSE2 to 64bit-with-SSE2.
And Greg's experiences with heavy SSE2 floating point code are somewhat
of an exceptional case for code length comparisons.

Or rather if you do such comparisons compare both, one applies to
optimized software and the other to packaged software, where shipping
x87 code will be probably the norm for some more years until all the
non SSE2 x86s are throughly obsolete. It probably does not make that
much difference, because floating point heavy code is usually rare
(and when it is not you are more likely to work with "optimized"
instead of "generic" code). The Intel compiler actually has options to
generate paths for both, but comparing to that would be really unfair.

> When I last checked several years ago, 32-bit integer code on x86
> took up approximately 3.8 bytes/instruction. IIRC, x87 intensive

I get 3.2 bytes/instructions with gcc 3.2 (and 3.4 bytes for x86-64
with gcc 3.3-hammer, but with less instructions). The gcc 3.2 code was
mostly optimized for the P6 core, the 64bit code was optimized for
the K8 (including big function prologues). It probably depends on
the compiler a lot.

> code had a _shorter_ average instruction length due to its tight
> encoding, where one register operand (ST0) is implicit. Also, x87
> code in general did not need prefixes (e.g. 0x0f, 0x66 etc). FWIW,
> 16-bit x86 integer code (ca 1995) ran about 2.8 bytes/instruction.

Thanks for the information.

-Andi

Yousuf Khan

unread,
May 10, 2004, 2:00:06 AM5/10/04
to

Yes, obviously you could've replaced that entire "mov eax, ..." stuff with a
"mov ah, 0x01" and that would've compressed it down nicely. But that's not
really the point I was trying to make. The point I was trying to make was
that immediate values would clog up the instruction cache not the data
cache. And a big 32-bit immediate would clog up an Icache more than an
old-fashioned 8-bit immediate.

If we're worried about the sizes of variable-length instructions occupying
too much room in Icaches, an instruction with an immediate value would be
one of the largest instructions available.

Yousuf Khan


Terje Mathisen

unread,
May 10, 2004, 2:25:53 AM5/10/04
to
Stefan Monnier wrote:

May I suggest you all take the 70-90 minutes required to watch Bob
Colwell (+ Andy 'Crazy' Glew at one point) explain all this stuff in a
Stanford lecture?

http://stanford-online.stanford.edu/courses/ee380/040218-ee380-100.asx

_Very_ short version: Yes, the P4 was intentionally a GHz speed demon,
but tempered with the need to deliver some actual performance that the
engineers could be comfortable with.

Bob also makes the same argument that I made in a conference
presentation last year: The P4 is quite brittle, i.e. it is too easy to
get stuck with very non-optimal performance for too long.

Terje
PS. Thanks to RM for sending me the link!

Terje Mathisen

unread,
May 10, 2004, 4:13:00 AM5/10/04
to
Yousuf Khan wrote:

> Terje Mathisen <terje.m...@hda.hydro.com> wrote:
>
>>Yousuf Khan wrote:
>>
>>>You're right, they shouldn't be much of a problem either. I was
>>>actually thinking of those x86 operand-embedded-in-instruction style
>>>instructions.
>>
>>Huh?
>>
>>Which opcodes would that be?
>>
>>Those with immediate data? Implicit registers.
>
>
> Yes, the immediate data would be the one. Things such as:
>
> mov eax, 0x00000001
>
> Where that final 32-bit number "1" would occupy a full 4 bytes in the
> instruction stream.

Actually, it would not: Values from -128 to +127 are encoded as a single
byte, making the instruction two or three bytes shorter.

However, what's the problem???

Instructions with immediate data are a staple of pretty much every cpu
architecture afaik!


>
> But remind me, which instruction forms are the implicit registers?

MUL/DIV/SH*/SAR/R*R/R*L/LOOP*/CBW/LODS/STOS/MOVS/IN/OUT/...

Grumble

unread,
May 10, 2004, 4:37:31 AM5/10/04
to
Stephen Sprunk wrote:

> Smaller caches provide lower latency, which was supposedly
> the justification for the anemic L1D and trace caches in
> Willamette/Northwood. Prescott Doubles the L1D size and the
> associativity, but at the cost of increasing the latency from
> one (two?) cycles to four cycles. Based on performance results
> to date, this is appears to be a wash.

IA-32 Optimization Reference Manual
http://intel.com/design/pentium4/manuals/24896610.pdf
Table 1-1 Pentium 4 and Intel Xeon Processor Cache Parameters

Northwood
L1 = 8 KB, 4-way, 2/9 cycles INT/FP latency, write-through
L2 = 512 KB, 8-way, 9/16 cycles INT/FP latency, write-back

Prescott
L1 = 16 KB, 8-way, 4/12 cycles INT/FP latency, write-through
L2 = 1 MB, 8-way, 22/30 cycles INT/FP latency, write-back


L2 latency took a hit too :-)

Nick Maclaren

unread,
May 10, 2004, 4:56:04 AM5/10/04
to

In article <c7ndid$ln9$1...@osl016lin.hda.hydro.com>,

Terje Mathisen <terje.m...@hda.hydro.com> writes:
|>
|> However, what's the problem???
|>
|> Instructions with immediate data are a staple of pretty much every cpu
|> architecture afaik!

Hmm. That's SLIGHTLY overstating it, because there was a long
period when there were a lot of architectures that didn't use
them (or not much). The System/370, for example, was one of
those.

However, I agree that what's the problem? It is a well-known and
fairly problem-free technology.


Regards,
Nick Maclaren.

Grumble

unread,
May 10, 2004, 7:32:59 AM5/10/04
to
Yousuf Khan wrote:

> Yes, obviously you could've replaced that entire "mov eax, ..."
> stuff with a "mov ah, 0x01" and that would've compressed it down
> nicely.

Partial register write :-)

> If we're worried about the sizes of variable-length instructions
> occupying too much room in Icaches, an instruction with an
> immediate value would be one of the largest instructions available.

add [eax + 4*ecx + 0x10000], 0x12345678

81 84 88 00 00 01 00 78 56 34 12

You could also throw a few instruction prefixes in the mix :-)

Chris Morgan

unread,
May 10, 2004, 10:27:12 AM5/10/04
to
Bengt Larsson <bengt...@telia.NOSPAMcom> writes:

> "Samuel" <sam...@austin.rr.com> wrote:
>
> >NetBUST is dead... YAY!!!
> >
> >Another screw up added to Intel's list RamBus, IA-64 and not going with 64
> >bit X86.


> >
> >...
>
> Just a quiet question: why all the Intel hatred?

I don't hate intel, but I can't resist some schadenfreude when the
biggest propaganda merchant in the industry, the one that calls
Itanium the "industry-standard" architecture, and SPARC "proprietary"
admits a miscalculation of this size.

On the other hand, I think they have done the right thing and that
Pentium-M is a fine product. It's nearly as nice as an AMD CPU.

Chris
--
Chris Morgan
"Post posting of policy changes by the boss will result in
real rule revisions that are irreversible"

- anonymous correspondent

Bengt Larsson

unread,
May 10, 2004, 10:42:50 AM5/10/04
to
"Yousuf Khan" <news.tal...@spamgourmet.com> wrote:

>Bengt Larsson <bengt...@telia.NOSPAMcom> wrote:
>> It's true that Intel would have a monopoly on IA-64 but there are two
>> comments one can make on that:
>>
>> 1. Intel want to compete with IBM on processors for the high end, and
>> IBM show no signs of stopping, so there will be competition.
>
>Everybody else gave up the ghost at least five years ago, at the mere
>thought of having to compete against IA64 before there was even a working
>IA64, and instead embraced it wholeheartedly

Well, I was talking about the present.

> -- bye-bye MIPS, Alpha,
>PA-RISC, etc.

No "etc.". HP always wanted to switch to IA-64 from HP-PA - They
collaborated with Intel in designing iA-64. That leaves two, MIPS and
Alpha, of which Alpha is the most notable. SPARC and PowerPC
continued. And AMD of course, with x86.

>What you're merely saying is that we don't have to worry about
>lack of competition because there were at least a few corporations that were
>not stupid enough to give up their own processor architectures. What if they
>had _all_ decided to give up at the mere mention of competition from Intel?

But they didn't. Why worry about the what-if? The remarkable thing is
that they didn't demand second-source.

>> 2. If Intel were to go too far and milk their monopoly too much there
>> is a very simple remedy: force them to license IA-64 to someone. The
>> Anti-trust remedy writes itself, unlike cases vs. Microsoft and IBM.
>
>It's that simple, huh? Intel has managed to make life very difficult for its
>x86 competitors to sell their processors to OEMs, and yet Intel manages to
>keep away from the anti-trust authorities, because it never ever writes down
>its threats.

It's a lot simpler than the other anti-trust cases, which was my
point. The anti-trust remedy I was talking about was for the IP
(intellectual property) that Intel/HP has/have.

>> I have nothing against PowerPC, but it's domineered/dominated by IBM.
>> There will never realistically be competition on PowerPC-based systems
>> vs IBM.
>
>There was Motorola too.

Not at the high end. I can't see anyone else doing high-end PowerPC
processors, other than IBM, who would have a real chance of suceeding
with it. Motorola, last I looked, only wanted to do embedded.

>> The ideal would be processors from Intel and AMD, systems from other
>> people (like Dell, HP, IBM...), operating systems from yet other
>> people (like Linux) etc. All to promote competition.
>
>Yes, that would be the ideal. However that's what's happening already, but
>in a lopsided fashion.

Except it will be x86-64. It will work, but I can't bring myself to
cheering for it.

Allan Sandfeld Jensen

unread,
May 10, 2004, 11:52:35 AM5/10/04
to
Yousuf Khan wrote:

And scheduling grouping (instruction merging) like Power4. It has improved
its IPC since the PIII days.

`Allan

krw

unread,
May 10, 2004, 12:32:49 PM5/10/04
to
In article <iYznc.5135$n7P1.3091
@twister01.bloor.is.net.cable.rogers.com>, news.tally.bbbl67
@spamgourmet.com says...

> Samuel <sam...@austin.rr.com> wrote:
> > "Yousuf Khan" <news.tal...@spamgourmet.com> wrote in message
> >> Actually the main reason that PowerPC didn't win in the desktop
> >> market was not because of performance, but because it simply didn't
> >> run x86 software,
> >
> > Sure there is no arguing that, but there was a time when G4 PowerPC
> > was lagging WAY behind in Frequency and Spec Int performance,
> > particularly with the G3 and the G4. At one time the Pentium4 was
> > 2-3X the frequency of the G4 which got Steve Jobs very worried. Then
> > Jobs pulled the G5 development from Motorola granted it to IBM.
>
> Jobs needn't have worried. Just like there was no way PC people were ever
> going to switch to a Macintosh processor, no matter what the performance,
> similarly there was no way that Macintosh people would've ever switched to a
> PC processor. You're just stuck in the environment that you're stuck in.
>
> > Anyhow, with IBM clocking the PPC976 @ 3.5 GHz and Intel falling back
> > on a 2 GHz core, I see Apple would hardly be worried about frequency
> > wars in the next couple of year at least. Adding to that the Power6
> > core will be a speed demon which is a 180 degree shift from the
> > Power4, Power5 approach, Apple would be even more comfortable with
> > frequency in the future (assuming IBM don't screw up with Power6 like
> > Intel screwed up with NetBUST)
>
> IBM isn't at 3.5 Ghz yet, and it would seem rather optimistic that they'll
> even touch that speed even with 90nm. What is the pipeline length of the PPC
> 97x? About 10 stages?

I wish Book-4 were public, but from the MPF 2002 PPC970 presentation:
http://www306.ibm.com/chips/techlib/techlib.nsf/techdocs/A1387A29AC1C2A
E087256C5200611780/$file/PPC970_MPF2002.pdf (ref: page 8)

9 fetch/decode stages
5-13 OoO execute stages
+ 2-3 dispatch/completion
------
16-25 stages in the complete instruction pipe

By counting the stages in the diagram, the VF pipe would be the longest
at 25 stages and FX1, FX2, BR, and CR being the shortest at 16. The FP
pipes are 21 stages.

--
Keith

Daniel Gustafsson

unread,
May 10, 2004, 12:48:07 PM5/10/04
to

"Samuel" <sam...@austin.rr.com> wrote in message
news:xVxnc.73069$Dn1....@fe2.texas.rr.com...
>
> "Daniel Gustafsson" <dan...@mimer.se> wrote in message
> > The Pentium M has similarities with Pentium 3 and those did not had
> > bad SPECint numbers.
>
> Yes and I was a very good wrestler when I was in High School.

>
> PIIIs were great at some point, probably even before the first Athlon came
> out. The PIIIs were competing very well with K6 and the four stage
pipelined
> G3 PPC. When the Athlon came out it took the lead in Spec int performance
> and the PIII cores started to show their age. Today it's a different
world,
> AMD is doing a great job and they are ahead with multi core design with
> cores that are superior to the good old PIIIs.

Yes, AMD is doing a great job, whether they will ship a dual core CPU before
Intel remains to be seen. It is also difficult to compare the current
Pentium M's with other current desktop/server chips because they have slower
busses. Who knows what a desktop/server Pentium M with a bus similar to what
the current Pentium 4 have could do in terms of speed.

> I don't really see how the "Bach to the Future" approach would make Intel
> processors competative in performance with AMD.
>

> > ( Although I get what you say, Sun does not focus on TPCC numbers ;) )
>
> From reading about the Rock, Niagra and follow ups it seems like it's all
> they care about now.

You are ofcourse correct that Rock is intended to compete in that market.
I was a bit unclear, but Sun does not publish TPCC benchmarks for various
reasons.

Regards
Daniel Gustafsson


Niels Jørgen Kruse

unread,
May 10, 2004, 1:04:03 PM5/10/04
to
I artiklen <iYznc.5135$n7P1...@twister01.bloor.is.net.cable.rogers.com> ,
"Yousuf Khan" <news.tal...@spamgourmet.com> skrev:

> Samuel <sam...@austin.rr.com> wrote:
>> "Yousuf Khan" <news.tal...@spamgourmet.com> wrote in message
>>> Actually the main reason that PowerPC didn't win in the desktop
>>> market was not because of performance, but because it simply didn't
>>> run x86 software,
>>
>> Sure there is no arguing that, but there was a time when G4 PowerPC
>> was lagging WAY behind in Frequency and Spec Int performance,
>> particularly with the G3 and the G4. At one time the Pentium4 was
>> 2-3X the frequency of the G4 which got Steve Jobs very worried. Then
>> Jobs pulled the G5 development from Motorola granted it to IBM.
>
> Jobs needn't have worried. Just like there was no way PC people were ever
> going to switch to a Macintosh processor, no matter what the performance,
> similarly there was no way that Macintosh people would've ever switched to a
> PC processor. You're just stuck in the environment that you're stuck in.

This is an exaggeration. There are switchers back and forth and people who
take one of each. Percieved performance does matter for marginal sales.
(Moderated by the fact that many applications have dropped under the Forrest
curve.)

>> Anyhow, with IBM clocking the PPC976 @ 3.5 GHz and Intel falling back
>> on a 2 GHz core, I see Apple would hardly be worried about frequency
>> wars in the next couple of year at least. Adding to that the Power6
>> core will be a speed demon which is a 180 degree shift from the
>> Power4, Power5 approach, Apple would be even more comfortable with
>> frequency in the future (assuming IBM don't screw up with Power6 like
>> Intel screwed up with NetBUST)
>
> IBM isn't at 3.5 Ghz yet, and it would seem rather optimistic that they'll
> even touch that speed even with 90nm. What is the pipeline length of the PPC
> 97x? About 10 stages? I'd say AMD would be closer to 3.5 Ghz with its 12
> stage pipeline than the PPC. AMD is already nearing 2.5 Ghz with a 130nm
> process.

Samuel seems to have been reading rumor sites. Throwing around clock rates
for future cores is pointless, it is far too frequent to miss targets. BTW,
the 970 has a 12 clock branch mispredict penalty.

Part of Intels trouble with Prescott may be due to the use of strained
silicon. When transistors are strained in place, it seems [speculation] that
the strain must extend into the substrate, increasing leakage. This has to
be much worse without SOI. SSDOI should not have this problem (I think).

The current PPC970FX as shipped in the Xserve is actually produced without
strained silicon, contrary to early reports. At 1.0V the unstrained 970FX
achieves 1.4 GHz at 12.3W, whereas the strained 970FX consumes 15W at 625
MHz (powersave at 1.0V). Dynamic power should be less than half at 625 MHz,
so leakage must be up in a serious way.

I wonder if Dothan is unstrained.

--
Mvh./Regards, Niels Jørgen Kruse, Vanløse, Denmark

Brian Hurt

unread,
May 10, 2004, 2:29:10 PM5/10/04
to
"Bill Todd" <bill...@metrocast.net> wrote in message news:<mPSdnfoSNOl...@metrocast.net>...

One byte, but not two byte. At least, gasm won't output the 2-byte
format. The one in
mov eax, 0x01
takes only one byte, but the 256 in:
mov eax, 256
takes four.

A side comment to another post in this thread. The x87 stack based
architecture does allow for shorter instruction lengths- assuming you
don't count the fxchg instructions you need to insert to get any sort
of super scalar performance out of the chip. Instructions you don't
need on the average RISC processor (or on SSE math).

Brian

>
> - bill

Yousuf Khan

unread,
May 10, 2004, 4:44:06 PM5/10/04
to
Grumble <inv...@kma.eu.org> wrote:
>> If we're worried about the sizes of variable-length instructions
>> occupying too much room in Icaches, an instruction with an
>> immediate value would be one of the largest instructions available.
>
> add [eax + 4*ecx + 0x10000], 0x12345678
>
> 81 84 88 00 00 01 00 78 56 34 12
>
> You could also throw a few instruction prefixes in the mix :-)

Okay, you got me, that's much longer. :-)

Yousuf Khan


Yousuf Khan

unread,
May 10, 2004, 5:04:09 PM5/10/04
to
Bengt Larsson <bengt...@telia.NOSPAMcom> wrote:

> "Yousuf Khan" <news.tal...@spamgourmet.com> wrote:
>> What you're merely saying is that we don't have to worry about
>> lack of competition because there were at least a few corporations
>> that were not stupid enough to give up their own processor
>> architectures. What if they had _all_ decided to give up at the mere
>> mention of competition from Intel?
>
> But they didn't. Why worry about the what-if? The remarkable thing is
> that they didn't demand second-source.

Not so remarkable if you think about the fact that most of these companies
are scared sh*tless of Intel, and the remainder merely just worry all of the
time about Intel. IBM was able to demand a second source for the original
IBM PC processors from Intel back in early 80's, when Intel was still just a
nobody (and thus led to a collaboration with AMD). But now, Intel just has
to tell them that they got about a dozen plants throughout the world, they
don't need no stinkin second source.

In fact, their original argument about the Itanium was that they could
probably outproduce all second sources put together. That's what made all of
these companies so scared sh*tless in the first place, they looked at
Intel's supposed economies of scale and gave up the ghost.

>>> 2. If Intel were to go too far and milk their monopoly too much
>>> there is a very simple remedy: force them to license IA-64 to
>>> someone. The Anti-trust remedy writes itself, unlike cases vs.
>>> Microsoft and IBM.
>>
>> It's that simple, huh? Intel has managed to make life very difficult
>> for its x86 competitors to sell their processors to OEMs, and yet
>> Intel manages to keep away from the anti-trust authorities, because
>> it never ever writes down its threats.
>
> It's a lot simpler than the other anti-trust cases, which was my
> point. The anti-trust remedy I was talking about was for the IP
> (intellectual property) that Intel/HP has/have.

I suppose they could've given some contract fabber (TSMC, UMC, SMIC, etc.) a
token contract to produce Itaniums. And then not given them any assistance
in fabbing the buggers. And then when all customers demanded GenuineIntel
Itaniums, they could've pointed to the problems these fabbers had in
producing Itaniums as the excuse for not ever buying alternate sourced
Itaniums from them. And then the customers would say, "besides, Intel
produces more than enough Itaniums, for our needs, why bother with a second
source?".

>>> I have nothing against PowerPC, but it's domineered/dominated by
>>> IBM. There will never realistically be competition on PowerPC-based
>>> systems vs IBM.
>>
>> There was Motorola too.
>
> Not at the high end. I can't see anyone else doing high-end PowerPC
> processors, other than IBM, who would have a real chance of suceeding
> with it. Motorola, last I looked, only wanted to do embedded.

Wasn't Motorola the sole source of PowerPC's for Apple for the past decade,
until the G5?

>>> The ideal would be processors from Intel and AMD, systems from other
>>> people (like Dell, HP, IBM...), operating systems from yet other
>>> people (like Linux) etc. All to promote competition.
>>
>> Yes, that would be the ideal. However that's what's happening
>> already, but in a lopsided fashion.
>
> Except it will be x86-64. It will work, but I can't bring myself to
> cheering for it.

I was referring to the bit about processors from Intel and AMD: happening
already, but a few system vendors refuse to deal with anybody other than
Intel for whatever reason (Dell), while most others try to avoid it as much
as possible (just in case Intel gets mad).

Systems from other people (Dell, HP, IBM, etc.): happening already, but some
vendors get special treatment from their suppliers (eg.. Dell from Intel).

Operating systems from other people (Linux): happening already, but
Microsoft puts severe pressure on system vendors to not do it.

Yousuf Khan


Yousuf Khan

unread,
May 10, 2004, 5:14:10 PM5/10/04
to
krw <k...@att.biz> wrote:
>> IBM isn't at 3.5 Ghz yet, and it would seem rather optimistic that
>> they'll even touch that speed even with 90nm. What is the pipeline
>> length of the PPC 97x? About 10 stages?
>
> I wish Book-4 were public, but from the MPF 2002 PPC970 presentation:
> http://www306.ibm.com/chips/techlib/techlib.nsf/techdocs/A1387A29AC1C2A
> E087256C5200611780/$file/PPC970_MPF2002.pdf (ref: page 8)
>
> 9 fetch/decode stages
> 5-13 OoO execute stages
> + 2-3 dispatch/completion
> ------
> 16-25 stages in the complete instruction pipe
>
> By counting the stages in the diagram, the VF pipe would be the
> longest at 25 stages and FX1, FX2, BR, and CR being the shortest at
> 16. The FP pipes are 21 stages.

Hmm, that's almost P4 Northwood territory.

Yousuf Khan


Yousuf Khan

unread,
May 10, 2004, 5:14:10 PM5/10/04
to
Daniel Gustafsson <dan...@mimer.se> wrote:
> Yes, AMD is doing a great job, whether they will ship a dual core CPU
> before Intel remains to be seen. It is also difficult to compare the
> current Pentium M's with other current desktop/server chips because
> they have slower busses. Who knows what a desktop/server Pentium M
> with a bus similar to what the current Pentium 4 have could do in
> terms of speed.

I'm sure AMD will get it out before Intel, they've already semi-officially
talked about bringing it out (AMD CEO interview a few weeks back). But it's
likely only for the server market, so it'll only be an Opteron. But if Intel
does manage to convince the world that desktops need dual-core, then I'm
sure AMD can do another Athlon 64FX job and transplant an Opteron (dual-core
this time) into the desktop.

> You are ofcourse correct that Rock is intended to compete in that
> market.
> I was a bit unclear, but Sun does not publish TPCC benchmarks for
> various reasons.

Probably because they sucked at those benchmarks prior to Opteron.

Yousuf Khan


Yousuf Khan

unread,
May 10, 2004, 5:29:22 PM5/10/04
to
Niels Jørgen Kruse <nj_k...@get2net.dk> wrote:
> Part of Intels trouble with Prescott may be due to the use of strained
> silicon. When transistors are strained in place, it seems
> [speculation] that the strain must extend into the substrate,
> increasing leakage. This has to be much worse without SOI. SSDOI
> should not have this problem (I think).
>
> The current PPC970FX as shipped in the Xserve is actually produced
> without strained silicon, contrary to early reports. At 1.0V the
> unstrained 970FX achieves 1.4 GHz at 12.3W, whereas the strained
> 970FX consumes 15W at 625 MHz (powersave at 1.0V). Dynamic power
> should be less than half at 625 MHz, so leakage must be up in a
> serious way.
>
> I wonder if Dothan is unstrained.

There was some talk about five years back about using isotopically pure
silicon wafers, where non-Si-28 isotopes are removed so that the remaining
Si-28 isotopes congeal into a very constistent lattice structure because
there are no other sized atoms to mess up the lattice. Looking at the ideas
behind strained silicon, it's also there to get a very consistent lattice
structure.

If you're saying that straining the silicon introduces negative effects, and
if the need for a consistent lattice structure are still there, then perhaps
Si-28's time has come finally?

Yousuf Khan


J Ahlstrom

unread,
May 10, 2004, 6:40:34 PM5/10/04
to
Yousuf Khan wrote:

--snip snip


>
> Not so remarkable if you think about the fact that most of these companies
> are scared sh*tless of Intel, and the remainder merely just worry all of the
> time about Intel. IBM was able to demand a second source for the original
> IBM PC processors from Intel back in early 80's, when Intel was still just a
> nobody (and thus led to a collaboration with AMD). But now, Intel just has
> to tell them that they got about a dozen plants throughout the world, they
> don't need no stinkin second source.
>
>

> Yousuf Khan
>
>
>

Can anyone confirm this statement that it
was IBM's insistence on a 2nd source for 8088s
that led to the AMD second-source agreement?
If this was not the reason for that agreement,
what was?

JKA

Rob Warnock

unread,
May 10, 2004, 7:48:04 PM5/10/04
to
Nick Maclaren <nm...@cus.cam.ac.uk> wrote:
+---------------

| Terje Mathisen <terje.m...@hda.hydro.com> writes:
| |> Instructions with immediate data are a staple of pretty much every cpu
| |> architecture afaik!
|
| Hmm. That's SLIGHTLY overstating it, because there was a long
| period when there were a lot of architectures that didn't use
| them (or not much). The System/370, for example, was one of those.
+---------------

Well, that's SLIGHTLY overstating it yourself, because people used
the "load effective address" (I forget the opcode mnemonic, "LA"?,
but it used RX format) with index register 0 as a 12-bit "load
immediate" all over the place! ;-}


-Rob

-----
Rob Warnock <rp...@rpw3.org>
627 26th Avenue <URL:http://rpw3.org/>
San Mateo, CA 94403 (650)572-2607

Bruce Hoult

unread,
May 10, 2004, 9:29:13 PM5/10/04
to
In article
<d7Snc.14907$n7P1...@twister01.bloor.is.net.cable.rogers.com>,
"Yousuf Khan" <news.tal...@spamgourmet.com> wrote:

> Wasn't Motorola the sole source of PowerPC's for Apple for the past decade,
> until the G5?

Not even close.

I could look up more exact data if needed, but:

- all the original PPC601's in the 6100/7100/8100 came from IBM

- the 603's and 604's were a mixed bag

- I don't recall with the initial G3's

- Motorola was the sole source for G4's, until they couldn't make enough
fast enough and Apple got IBM into making them too.

- while Motorola was struggling with G4's, Apple was getting faster and
faster G3's from IBM, for iMacs and (especially) iBooks.


The only real Motorola-exclusive was the G4 with Altivec. The design of
VMX/AltiVec was initiated by IBM, and worked on cooperatively by both
companies, but then IBM decided it wasn't needed by their market and
they didn't make any until very recently.

-- Bruce

KR Williams

unread,
May 10, 2004, 10:58:51 PM5/10/04
to
In article <CgSnc.14956$n7P1.10850

What made you think otherwise? The G3 was in the 4-5 clock
range, but the 970 has always be advertised as a long-pipe
machine.

--
Keith


Douglas Siebert

unread,
May 11, 2004, 12:34:50 AM5/11/04
to
"Yousuf Khan" <news.tal...@spamgourmet.com> writes:

>Not so remarkable if you think about the fact that most of these companies
>are scared sh*tless of Intel, and the remainder merely just worry all of the
>time about Intel. IBM was able to demand a second source for the original
>IBM PC processors from Intel back in early 80's, when Intel was still just a
>nobody (and thus led to a collaboration with AMD). But now, Intel just has
>to tell them that they got about a dozen plants throughout the world, they
>don't need no stinkin second source.


I wonder if more than one of their fabs produces Itaniums?

Douglas Siebert

unread,
May 11, 2004, 12:52:29 AM5/11/04
to
"Yousuf Khan" <news.tal...@spamgourmet.com> writes:

>Niels Jørgen Kruse <nj_k...@get2net.dk> wrote:
>> Part of Intels trouble with Prescott may be due to the use of strained
>> silicon. When transistors are strained in place, it seems
>> [speculation] that the strain must extend into the substrate,
>> increasing leakage. This has to be much worse without SOI. SSDOI
>> should not have this problem (I think).
>>
>> The current PPC970FX as shipped in the Xserve is actually produced
>> without strained silicon, contrary to early reports. At 1.0V the
>> unstrained 970FX achieves 1.4 GHz at 12.3W, whereas the strained
>> 970FX consumes 15W at 625 MHz (powersave at 1.0V). Dynamic power
>> should be less than half at 625 MHz, so leakage must be up in a
>> serious way.
>>
>> I wonder if Dothan is unstrained.

>There was some talk about five years back about using isotopically pure
>silicon wafers, where non-Si-28 isotopes are removed so that the remaining
>Si-28 isotopes congeal into a very constistent lattice structure because
>there are no other sized atoms to mess up the lattice. Looking at the ideas
>behind strained silicon, it's also there to get a very consistent lattice
>structure.


AMD purchased some isotopically pure wafers from Isonics a few years back,
nothing was ever made public about how well they worked. Isonics recently
said they'd (themselves/their customer, not mentioning AMD specifically)
made a lot of progress with them in customer evaluations and said this may
lead to taking it out of testing and into actual production. They didn't
say which customer or how much volume this production would be, so this
may be some botique low volume product like defense rather than a
mainstream CPU.

I was under the impression that Dothan was strained, but I'm not positive.
I quickly browsed a couple reviews today and one noted that while Dothan
used less power at 2.0 GHz (21w) than Banias at 1.7 GHz (24.5w) when both
were dropped to their lowest speedstep rate of 600 MHz at slightly under
1v, Dothan used a bit more power. That sounds like higher leakage to me,
but it was pretty close (6w to 6.5w or something like that)

So it sounds like it is doable to make good quality product at 90nm, and
Prescott is probably a design problem rather than a technology problem.
That's got to be good news for Itanium lovers, as well as AMD, IBM, and
Apple fans.

Terje Mathisen

unread,
May 11, 2004, 1:55:26 AM5/11/04
to
Brian Hurt wrote:

> "Bill Todd" <bill...@metrocast.net> wrote in message news:<mPSdnfoSNOl...@metrocast.net>...

>>Oh, my - it's been a *long* time. But though the details have faded from
>>memory ISTR that the x86 instruction set provides mechanisms for compressing
>>immediate operands that will fit into 1 or 2 bytes and zero- or
>>sign-extending them to full destination width (leaving aside explicit
>>mechanisms such as MOVZX and MOVSX - which may not take immediate source
>>operands - multi-instruction sequences and creative use of instructions such
>>as LEA).
>
> One byte, but not two byte. At least, gasm won't output the 2-byte
> format. The one in
> mov eax, 0x01
> takes only one byte, but the 256 in:
> mov eax, 256
> takes four.

You _can_ sometimes still do it in three by using a 16-bit prefix and a
16-bit constant. However, since prefix bytes used to take a full cycle
each on the 486 and Pentium, the usage really wasn't encouraged.


>
> A side comment to another post in this thread. The x87 stack based
> architecture does allow for shorter instruction lengths- assuming you
> don't count the fxchg instructions you need to insert to get any sort
> of super scalar performance out of the chip. Instructions you don't
> need on the average RISC processor (or on SSE math).

Even with a FXCHG between every regular Fops, which you don't really
need, you'd still most probably end up with shorter code. In fact, all
compilers I've seen over the last several years have inserted those
FXCHGs anyway, so code size comparisons would have been done with them
included.

Yousuf Khan

unread,
May 11, 2004, 3:19:09 AM5/11/04
to
Terje Mathisen <terje.m...@hda.hydro.com> wrote:
>> A side comment to another post in this thread. The x87 stack based
>> architecture does allow for shorter instruction lengths- assuming you
>> don't count the fxchg instructions you need to insert to get any sort
>> of super scalar performance out of the chip. Instructions you don't
>> need on the average RISC processor (or on SSE math).
>
> Even with a FXCHG between every regular Fops, which you don't really
> need, you'd still most probably end up with shorter code. In fact, all
> compilers I've seen over the last several years have inserted those
> FXCHGs anyway, so code size comparisons would have been done with them
> included.

Since all calculations are done against the x87's ST(0) register, and not
having done any assembly since the 386/387 days, since before FXCHG existed;
I am curious, how (or more appropriately "when") exactly do you use FXCHG to
get superscalar performance? I mean do you push the first two operands into
the stack, begin the calculation, and then while it's calculating, FXCHG the
currently calculating ST(0) to another location, and then begin the process
of pushing the next two operands onto the stack to begin the next parallel
calculation sequence? Or do you push all of the operand values in at once,
begin the first set of calculations, FXCHG it while it's chugging away,
begin the next set of calculations, FXCHG it, etc., etc.? I hope I made
myself clear.

I'll demonstrate with FPU pseudo-ops, where the FCALC pseudo-instruction
represents any x87 calculation operation.

Scenario 1:
fpush A ;push A in first, it will go to ST(1) eventually
fpush B ;push B in second, it will go to ST(0)
fcalc ; begin first set of calculations on A & B
fxch ST(2) ;exchange top of stack with next empty stack element
fpush C ;push C in third, it will go to ST(1) eventually
fpush D ;push D in fourth, it will go to ST(0)
fcalc ;begin second set of calculations on C & D

Scenario 2:
fpush A
fpush B
fpush C
fpush D
fcalc ; begin first set of calculations on C & D
fxchg ST(2)
fcalc ; begin second set of calculation on A & B now

Which one would be the more likely way, if either?

Yousuf Khan


glen herrmannsfeldt

unread,
May 11, 2004, 3:30:40 AM5/11/04
to
Rob Warnock wrote:

(snip)

> | Hmm. That's SLIGHTLY overstating it, because there was a long
> | period when there were a lot of architectures that didn't use
> | them (or not much). The System/370, for example, was one of those.

> Well, that's SLIGHTLY overstating it yourself, because people used


> the "load effective address" (I forget the opcode mnemonic, "LA"?,
> but it used RX format) with index register 0 as a 12-bit "load
> immediate" all over the place! ;-}

Not to mention the real immediate instructions, such as MVI, CLI,
XI, OI, NI, TM, though those are only storage immediate.

Yes, LA is the immediate instruction for loading a 12 bit unsigned
constant into a register.

-- glen

Ralph Schmidt

unread,
May 11, 2004, 9:05:45 AM5/11/04
to
Bruce Hoult wrote:

>
> The only real Motorola-exclusive was the G4 with Altivec.  The design of
> VMX/AltiVec was initiated by IBM, and worked on cooperatively by both
> companies, but then IBM decided it wasn't needed by their market and
> they didn't make any until very recently.
>

Can you give any pointers to your opinion that VMX was initiated by
IBM ? My knowledge is that it was mostly some Apple/Motorola thing and the
pre release VMX docs i've seen 97 or so only indicated Motorola here.

Regards

Terje Mathisen

unread,
May 11, 2004, 9:49:48 AM5/11/04
to
Yousuf Khan wrote:
> Scenario 1:
> fpush A ;push A in first, it will go to ST(1) eventually
> fpush B ;push B in second, it will go to ST(0)
> fcalc ; begin first set of calculations on A & B
> fxch ST(2) ;exchange top of stack with next empty stack element
> fpush C ;push C in third, it will go to ST(1) eventually
> fpush D ;push D in fourth, it will go to ST(0)
> fcalc ;begin second set of calculations on C & D
>
> Scenario 2:
> fpush A
> fpush B
> fpush C
> fpush D
> fcalc ; begin first set of calculations on C & D
> fxchg ST(2)
> fcalc ; begin second set of calculation on A & B now
>
> Which one would be the more likely way, if either?

Closer to the first.

Assume you want to sum together an array of (single prec) fp values:

Doing it the naive way

for (i = 0; i < len; i++)
sum += a[i];

could lead to code like this (assume array has at least some elements!)

fld dword ptr [esi]
add esi,4
dec ecx
next:
fadd [esi]
add esi,4
dec ecx
jnz next

The latency of this would be limited by the dependent FADDs, right?

By doing three FADDs in parallel, the throughput increases:

fld dword ptr [esi] ; 0
fld dword ptr [esi+4] ; 1 0
fld dword ptr [esi+8] ; 2 1 0
fxchg st(2) ; 0 1 2 Swap st(0) with st(2)
add esi,12
sub ecx,3
next6:
fadd dword ptr [esi]
fxch st(1) ; 1 0 2
fadd dword ptr [esi+4]
fxchg st(2) ; 2 0 1
fadd dword ptr [esi+8]
fxchg st(1) ; 0 2 1
fadd dword ptr [esi+12]
fxch st(2) ; 1 2 0
fadd dword ptr [esi+16]
fxchg st(2) ; 2 1 0
fadd dword ptr [esi+20]
fxchg st(1) ; 0 1 2
add esi,24
sub ecx,6
ja next6

faddp st,st(1) ; 1 2
faddp st,st(1) ; 2

Modulo handling any remaining items etc, this version will run three
times faster than the first.

Daniel Gustafsson

unread,
May 11, 2004, 12:38:02 PM5/11/04
to
"Yousuf Khan" <news.tal...@spamgourmet.com> wrote in message
news:CgSnc.14957$n7P1...@twister01.bloor.is.net.cable.rogers.com...

> Daniel Gustafsson <dan...@mimer.se> wrote:
>
> > You are ofcourse correct that Rock is intended to compete in that
> > market.
> > I was a bit unclear, but Sun does not publish TPCC benchmarks for
> > various reasons.
>
> Probably because they sucked at those benchmarks prior to Opteron.

Not necessarily, but lets drop this now, it has been discussed before...

Regards
Daniel Gustafsson


Tim McCaffrey

unread,
May 11, 2004, 1:25:37 PM5/11/04
to
In article <SuSnc.15028$n7P1...@twister01.bloor.is.net.cable.rogers.com>,
news.tal...@spamgourmet.com says...
>

>There was some talk about five years back about using isotopically pure
>silicon wafers, where non-Si-28 isotopes are removed so that the remaining
>Si-28 isotopes congeal into a very constistent lattice structure because
>there are no other sized atoms to mess up the lattice. Looking at the ideas
>behind strained silicon, it's also there to get a very consistent lattice
>structure.
>
>If you're saying that straining the silicon introduces negative effects, and
>if the need for a consistent lattice structure are still there, then perhaps
>Si-28's time has come finally?
>

Isotopically pure silicon was being looked at because it had some wonderful
heat transfer characteritics, I don't think there was any great difference
with respect to electrical properties (IIRC).

- Tim

Zak

unread,
May 11, 2004, 1:35:36 PM5/11/04
to
Douglas Siebert wrote:

> Once the competition is
> gone, it is just in maintenance mode. Compare all the upgrades to IE
> back when Netscape was a threat, versus the last few years when they are
> a tired brand name barely remembered by the masses.

Or for that matter the days of DOS, where very little would change until
Digital Research DOS appeared, with improvements like on line
documentation and better error messages.

Suddenly, MS-DOS improved as well.


Thomas

Russell Wallace

unread,
May 11, 2004, 2:28:12 PM5/11/04
to
On Mon, 10 May 2004 21:29:22 GMT, "Yousuf Khan"
<news.tal...@spamgourmet.com> wrote:

>There was some talk about five years back about using isotopically pure
>silicon wafers, where non-Si-28 isotopes are removed so that the remaining
>Si-28 isotopes congeal into a very constistent lattice structure because
>there are no other sized atoms to mess up the lattice.

I would have expected the atom size to be determined only by the
atomic number, not the atomic weight - can anyone explain why the
atomic weight makes a difference?

--
"Sore wa himitsu desu."
To reply by email, remove
the small snack from address.
http://www.esatclear.ie/~rwallace

Duane Rettig

unread,
May 11, 2004, 2:51:44 PM5/11/04
to
"Yousuf Khan" <news.tal...@spamgourmet.com> writes:

[ Fear of Intel, systems moving, or not, to AMD64 ...]

> Operating systems from other people (Linux): happening already, but
> Microsoft puts severe pressure on system vendors to not do it.

Huh? Or perhaps this is a case of "do as I say, not as I do", eh?
MS certainly have an XP port to AMD64:

http://www.microsoft.com/windowsxp/64bit/downloads/upgrade.asp

so if they are indeed pressuring others not to port, perhaps it
is for competitive reasons...

--
Duane Rettig du...@franz.com Franz Inc. http://www.franz.com/
555 12th St., Suite 1450 http://www.555citycenter.com/
Oakland, Ca. 94607 Phone: (510) 452-2000; Fax: (510) 452-0182

Douglas Siebert

unread,
May 11, 2004, 2:52:08 PM5/11/04
to
wallacet...@eircom.net (Russell Wallace) writes:

>On Mon, 10 May 2004 21:29:22 GMT, "Yousuf Khan"
><news.tal...@spamgourmet.com> wrote:

>>There was some talk about five years back about using isotopically pure
>>silicon wafers, where non-Si-28 isotopes are removed so that the remaining
>>Si-28 isotopes congeal into a very constistent lattice structure because
>>there are no other sized atoms to mess up the lattice.

>I would have expected the atom size to be determined only by the
>atomic number, not the atomic weight - can anyone explain why the
>atomic weight makes a difference?


Si-28 is a silicon fullerene, that's why you get a very consistent and
more dense atomic lattice.

Russell Crook - Computer Systems - System Engineer

unread,
May 11, 2004, 4:40:21 PM5/11/04
to
Douglas Siebert wrote:
> wallacet...@eircom.net (Russell Wallace) writes:
>
>
>>On Mon, 10 May 2004 21:29:22 GMT, "Yousuf Khan"
>><news.tal...@spamgourmet.com> wrote:
>
>
>>>There was some talk about five years back about using isotopically pure
>>>silicon wafers, where non-Si-28 isotopes are removed so that the remaining
>>>Si-28 isotopes congeal into a very constistent lattice structure because
>>>there are no other sized atoms to mess up the lattice.
>
>
>>I would have expected the atom size to be determined only by the
>>atomic number, not the atomic weight - can anyone explain why the
>>atomic weight makes a difference?


The elctrons have angular momentum around the nucleus. If the mass
changes, so does the radius (slightly).

It is for this reason that diamonds of pure carbon 13 are slightly
harder (5%?) than diamonds of (the usual) carbon 12. The bonds
are shorter.


>
>
>
> Si-28 is a silicon fullerene, that's why you get a very consistent and
> more dense atomic lattice.

Eh? Si-28 would be a single *atom*, not a large empty cage molecule.
What were you trying to say?
>


--
Russell Crook, Technology Specialist, Computer Systems
Sun Microsystems of Canada
Sufficiently advanced cluelessness is indistiguishable from malice.
-- J. Porter Clark, NASA/MSFC Flight Data Systems Branch 1994/11/16

Yousuf Khan

unread,
May 11, 2004, 5:31:32 PM5/11/04
to
Tim McCaffrey <t...@spamfilter.asns.tr.unisys.com> wrote:
> Isotopically pure silicon was being looked at because it had some
> wonderful heat transfer characteritics, I don't think there was any
> great difference with respect to electrical properties (IIRC).

Yes, heat transfer properties, but I also recall reading on their website
about pure crystalline lattices with few defects.

Yousuf Khan


Yousuf Khan

unread,
May 11, 2004, 5:31:33 PM5/11/04
to
Russell Wallace <wallacet...@eircom.net> wrote:
> On Mon, 10 May 2004 21:29:22 GMT, "Yousuf Khan"
> <news.tal...@spamgourmet.com> wrote:
>
>> There was some talk about five years back about using isotopically
>> pure silicon wafers, where non-Si-28 isotopes are removed so that
>> the remaining Si-28 isotopes congeal into a very constistent lattice
>> structure because there are no other sized atoms to mess up the
>> lattice.
>
> I would have expected the atom size to be determined only by the
> atomic number, not the atomic weight - can anyone explain why the
> atomic weight makes a difference?

I would assume that while it is being cooled down from liquid state to solid
to make the wafers, the other-sized atoms would have slightly different
vibrational properties which would make it settle somewhat differently.

Yousuf Khan


Yousuf Khan

unread,
May 11, 2004, 5:41:35 PM5/11/04
to
Duane Rettig <du...@franz.com> wrote:
> "Yousuf Khan" <news.tal...@spamgourmet.com> writes:
>
> [ Fear of Intel, systems moving, or not, to AMD64 ...]
>
>> Operating systems from other people (Linux): happening already, but
>> Microsoft puts severe pressure on system vendors to not do it.
>
> Huh? Or perhaps this is a case of "do as I say, not as I do", eh?
> MS certainly have an XP port to AMD64:
>
> http://www.microsoft.com/windowsxp/64bit/downloads/upgrade.asp
>
> so if they are indeed pressuring others not to port, perhaps it
> is for competitive reasons...

No, no, people aren't being pressured not to port, they are being pressured
not to offer anything other than MS OSes to their customers.

Yousuf Khan


Tim Olson

unread,
May 11, 2004, 6:30:58 PM5/11/04
to

| Bruce Hoult wrote:
|
| >
| > The only real Motorola-exclusive was the G4 with Altivec.  The design of
| > VMX/AltiVec was initiated by IBM, and worked on cooperatively by both
| > companies, but then IBM decided it wasn't needed by their market and
| > they didn't make any until very recently.
|

Ralph Schmidt <la...@t-online.de> replied:



| Can you give any pointers to your opinion that VMX was initiated by
| IBM ? My knowledge is that it was mostly some Apple/Motorola thing and the
| pre release VMX docs i've seen 97 or so only indicated Motorola here.

VMX/Altivec was originally proposed by Apple (Keith Diefendorff), and
was defined by a group composed of members from all three companies.

-- Tim Olson

Bruce Hoult

unread,
May 11, 2004, 9:12:28 PM5/11/04
to
In article <c7qj32$gki$01$1...@news.t-online.com>,
Ralph Schmidt <la...@t-online.de> wrote:

I could well be misremembering the "initiated" part (and I wasn't
there), but I do recall seeing VMX papers from IBM, not Motorola, in the
early days. By the May '98 WWDC it looked to be all Motorola though.

Tim will know :-)

-- Bruce

John Miller

unread,
May 12, 2004, 2:16:46 AM5/12/04
to
On Mon, 10 May 2004 22:40:34 GMT, J Ahlstrom <Ahlst...@comcast.net>
wrote:


I hadn't heard before that it IBM's insistence on
a second source that lead to the AMD agreement
(but it is not outrageously unbelievable)

I was merely a teenager in the 70s, but it was my
understanding that if you wanted a part to be
accepted, you pretty well had to make sure that
there was a second source available (and third
or fourth sources would be even better).

As chip companies grew larger (than most
of their customers), they started saying "Look,
we aren't going to go under. This chip is
really, really great, and you can only get it
from us."

Peter Dickerson

unread,
May 12, 2004, 6:59:07 AM5/12/04
to
"Russell Wallace" <wallacet...@eircom.net> wrote in message
news:40a11b16....@news.eircom.net...

The atomic weight does make a small contribution to the size through a
change in the reduced mass. From a simple non-QM viewpoint this is because
the centre of mass between electrons an nucleus has changed. This affect is
tiny, roughly one part in 28*29*1800 for Si 29, so it unlikely this is
important.

However, one of the ways that heat is conducted through the crystal is by
lattice vibrations. Clearly, different isotopes will have different
resonance frequencies. So Si 29 (say) atoms will be a source of scattering,
reducing energy flow - and so reduce thermal conductivity.

Another affect, which may or may not be important is the fact that Si 28
nuclei are bosons with (I guess) no magnetic dipole moment, while Si 29
nuclei are fermions - dipole moment affects, exclusion principal affects,
spin-splitting affects... some of these may have an affect on electrical
conductivity too.

Peter


Russell Wallace

unread,
May 12, 2004, 10:35:02 AM5/12/04
to
On Wed, 12 May 2004 11:59:07 +0100, "Peter Dickerson"
<first{dot}sur...@ukonline.co.uk> wrote:

>The atomic weight does make a small contribution to the size through a
>change in the reduced mass. From a simple non-QM viewpoint this is because
>the centre of mass between electrons an nucleus has changed. This affect is
>tiny, roughly one part in 28*29*1800 for Si 29, so it unlikely this is
>important.

Hmm, from a QM viewpoint could you say it's because the atom as a
whole has a slightly shorter wavelength due to increased mass?


>
>However, one of the ways that heat is conducted through the crystal is by
>lattice vibrations. Clearly, different isotopes will have different
>resonance frequencies. So Si 29 (say) atoms will be a source of scattering,
>reducing energy flow - and so reduce thermal conductivity.

Ah! That makes sense.

It is loading more messages.
0 new messages