| Does anyone know why the Pythagorean method of determining winning
| percentages work so well in baseball? Granted, football is a much
| shorter season, but I calculated the expected winning percentages
| for the 1996 season and they were way off. Then I did one of the
| basketball divisions for this year and the Pythagorean method didn't
| seem to work there, either. If anyone wants that data, I can submit
| it. Right now, it's hand written on scrap paper in my office.
Well, the scoring environments are very different in baseball and
basketball (for example). In basketball, it is not unusual for the
winning team to score only 5% more points than the losing team,
whereas in baseball the winner will often score four times as many
points as the loser.
I'd have to think about exactly what the implications would be, but my
guess is that a small difference in points scored vs allowed would
make a bigger difference in basketball than in baseball. This would
also help to explain why it's not unusual for a basketball team to win
75+% of their games, while a baseball team almost never does.
Football's a weird case; sometimes teams win 28-27, and sometimes they
win 35-3. So I dunno how I'd expect that to compare to baseball.
--
Dan Schmidt -> df...@harmonixmusic.com, df...@alum.mit.edu
Honest Bob & the http://www2.thecia.net/users/dfan/
Factory-to-Dealer Incentives -> http://www2.thecia.net/users/dfan/hbob/
Gamelan Galak Tika -> http://web.mit.edu/galak-tika/www/
In baseball you always score in units (granted, you may get four units
at once). In football, you can score in 1's, 2's, 3's, or 7's. That's
probably enough to make any relationship between scoring and winning
pretty odd right there.
Mike Jones | jon...@rpi.edu
Actually, from what I have been able to piece together about the
Khmer Rouge, they were a strange blend of communist, rousseauist, and
environmentalist.Not too far from Al Gore.
- Brian K Yoder (byo...@netcom.com), alt.philosophy.objectivism,
August 1993.
>Does anyone know why the Pythagorean method of determining winning
percentages work so well in baseball? <
There is a mathematical relationship among runs scored, runs allowed, and the
number of additional runs required to win one additional ball game or winning
percentage. But why is that?
Bill James expesses the relationship in the Pythagorean Theorem. Pete Palmer
uses a run differential approach. An interesting thread apppeared here last
January exploring these approaches and the relationships between them capped by
an excellent post by John Franjione.
http://x24.deja.com/[ST_rn=ps]/getdoc.xp?AN=433418110&CONTEXT=926957241.12
26244107&hitnum=4
Read the thread from its beginning to follow the progression of ideas.
Each run or point added at random to a season will have a greater effect the
fewer the number of games and will have a greater effect the lower scoring the
games.
I suspect that the relationship between the level of scoring and the length
of the baseball season is the reason why these formulas work for baseball and
not for other sports.
Think of the impact of one additional score in soccer, baseball, and
basketball. It's not at all the same, is it? These fomulas were devised and
refined to fit baseball.
Don
Facts are stubborn things, but statistics are much more pliable.
--Laurence J. Peter
Jonp...@aol.com wrote:
> Does anyone know why the Pythagorean method of determining winning
> percentages work so well in baseball? Granted, football is a much shorter
> season, but I calculated the expected winning percentages for the 1996 season
> and they were way off. Then I did one of the basketball divisions for this
> year and the Pythagorean method didn't seem to work there, either. If anyone
> wants that data, I can submit it. Right now, it's hand written on scrap
> paper in my office.
I played around with exponents once for a football Pythagorean formula:
I think 2.5 worked pretty well, but (as you might expect) teams with real high
winning percentages (or real low) tended to be a bit "lucky". For example,
a team which goes undefeated or even 15-1 (.93%) probably does NOT
have a true quality over .900, and in fact most such teams tended to have Pyth.
winning percentages of .750-.850 (which is still pretty good).
Hockey, due to its rough similarity to baseball in terms of scoring, would
probably also work well IF you could figure out what to do with tie games...
Basketball is probably a lost cause...the exponent would probably need to
be close to 4 or 5 to give you even a ghost of a chance of accurate estimates...
John DiFool
Dfmccjr wrote in message <7hq5jo$o4$1...@fpage2.ba.best.com>...
> I suspect that the relationship between the level of scoring and the
length
>of the baseball season is the reason why these formulas work for baseball
and
>not for other sports.
>
>Think of the impact of one additional score in soccer, baseball, and
>basketball. It's not at all the same, is it? These fomulas were devised
and
>refined to fit baseball.
Could the scoring explosion change the dynamics so much that things like the
Pythagorean method don't work anymore? And how big would the scoring
increase have to be to do that? I'm not enough of a statistician to post
actual answers on here, but I have lots of questions...many of which might
not even be relevant or make sense to people who know what they're talking
about.
davidb
--------------CA7BDCFE58AB5AADB1BE880D
Content-Type: text/plain; charset=us-ascii
Content-Transfer-Encoding: 7bit
David Brazeal wrote:
> Could the scoring explosion change the dynamics so much that things like the
> Pythagorean method don't work anymore? And how big would the scoring
> increase have to be to do that? I'm not enough of a statistician to post
> actual answers on here, but I have lots of questions...many of which might
> not even be relevant or make sense to people who know what they're talking
> about.
Oh, go ahead and ask. This is a discussion group!
I just pulled down W/L and Run Scored/Allowed data from I can't remember where,
to study the Ptyhagorean Exponent in difference eras. I restricted attention to
the period 1901 (founding of the AL) to 1998, and used all league-seasons,
including the Federal League (1914-15) and strike-shortened years. I modeled
the calculation as a linear regression:
LOG (W/L) = Exponent x LOG (Scored/Allowed) and solved for EXPONENT using
Lotus.
Since the universe was 1,912 team/seasons, I divided it into groups of about
250 to study different eras. Here is what I got:
YEARS TEAMS EXPONENT SD(EXP) R/G
1990-1998 248 1.867 .048 9.25
1980-1989 260 1.938 .049 8.61
1970-1979 246 1.740 .042 8.32
1957-1969 246 1.878 .041 8.21
1942-1956 240 1.796 .035 8.78
1927-1941 240 1.978 .038 9.83
1914-1926 224 1.889 .038 8.62
1901-1913 208 1.796 .036 8.18
ALL YEARS 1912 1.854 .0138 8.73
My Comments:
1. Generally speaking, the exponent appears correlated with Runs/Game. This
makes sense. The exponent should go up the fewer close games there are, and the
higher the scoring the fewer close games there should be.
2. The standard deviations calculated by Lotus may not be accurate because they
don't recognize that leagues (until recently) are closed systems, and that total
wins and losses and runs scored/allowed must be equal within a league/season. I
don't know what affect this might have on the SD's.
3. Standard Deviations have been going up over time, especially since 1980. I
suspect this is related to the development of the relief closer, which would
increase the spread in the ability of teams to win close games.
4. There are only two eras in which the exponent appears statistically
significantly different from the overall exponent of 1.854. The 1927-1941
period is quite high, probably related to its high scoring. The 1970's exponent
is very low; it was a low-offense era, but it's exponent is still much lower
than the exponent in the other low-scoring eras.
5. The 1990's are the most average era in the bunch.
--------------CA7BDCFE58AB5AADB1BE880D
Content-Type: text/html; charset=us-ascii
Content-Transfer-Encoding: 7bit
<HTML>
David Brazeal wrote:
<BLOCKQUOTE TYPE=CITE>
<P>Could the scoring explosion change the dynamics so much that things
like the
<BR>Pythagorean method don't work anymore? And how big would the
scoring
<BR>increase have to be to do that? I'm not enough of a statistician
to post
<BR>actual answers on here, but I have lots of questions...many of which
might
<BR>not even be relevant or make sense to people who know what they're
talking
<BR>about.</BLOCKQUOTE>
Oh, go ahead and ask. This is a discussion group!
<P>I just pulled down W/L and Run Scored/Allowed data from I can't remember
where, to study the Ptyhagorean Exponent in difference eras. I restricted
attention to the period 1901 (founding of the AL) to 1998, and used all
league-seasons, including the Federal League (1914-15) and strike-shortened
years. I modeled the calculation as a linear regression:
<P>LOG (W/L) = Exponent x LOG (Scored/Allowed) and solved for EXPONENT
using Lotus.
<P>Since the universe was 1,912 team/seasons, I divided it into groups
of about 250 to study different eras. Here is what I got:
<BR>
<P><TT>YEARS TEAMS EXPONENT SD(EXP) R/G</TT>
<BR><TT>1990-1998 248 1.867 .048
9.25</TT>
<BR><TT>1980-1989 260 1.938 .049
8.61</TT>
<BR><TT>1970-1979 246 1.740 .042
8.32</TT>
<BR><TT>1957-1969 246 1.878 .041
8.21</TT>
<BR><TT>1942-1956 240 1.796 .035
8.78</TT>
<BR><TT>1927-1941 240 1.978 .038
9.83</TT>
<BR><TT>1914-1926 224 1.889 .038
8.62</TT>
<BR><TT>1901-1913 208 1.796 .036
8.18</TT>
<BR><TT>ALL YEARS 1912 1.854 .0138 8.73</TT><TT></TT>
<P><TT>My Comments:</TT><TT></TT>
<P><TT>1. Generally speaking, the exponent appears correlated with
Runs/Game. This makes sense. The exponent should go up the
fewer close games there are, and the higher the scoring the fewer close
games there should be.</TT><TT></TT>
<P><TT>2. The standard deviations calculated by Lotus may not be
accurate because they don't recognize that leagues (until recently) are
closed systems, and that total wins and losses and runs scored/allowed
must be equal within a league/season. I don't know what affect this
might have on the SD's.</TT><TT></TT>
<P><TT>3. Standard Deviations have been going up over time, especially
since 1980. I suspect this is related to the development of the relief
closer, which would increase the spread in the ability of teams to win
close games.</TT><TT></TT>
<P><TT>4. There are only two eras in which the exponent appears statistically
significantly different from the overall exponent of 1.854. The 1927-1941
period is quite high, probably related to its high scoring. The 1970's
exponent is very low; it was a low-offense era, but it's exponent is still
much lower than the exponent in the other low-scoring eras.</TT>
<P>5. The 1990's are the most average era in the bunch.</HTML>
--------------CA7BDCFE58AB5AADB1BE880D--
>Could the scoring explosion change the dynamics so much that things like the
>Pythagorean method don't work anymore? And how big would the scoring
>increase have to be to do that? I'm not enough of a statistician to post
>actual answers on here, but I have lots of questions...many of which might
>not even be relevant or make sense to people who know what they're talking
>about.
It's unlikely to. Despite complaints about the "scoring explosion",
present day scores are well within historic norms. Scoring in the AL is
certainly lower than it was in the 1930's AL, and scoring in the NL is a
bit lower than it was in the 1950's, much less the 1920's or 1890's. The
pythagorean approach, even with the exact same exponent, worked just fine
in those past seasons where scoring was higher, so there's no reason to
think that it won't continue to work well today.
Even if things changed drastically there's no particular reason to think
that the general _approach_ used in the pythagorean method will fail.
Contrary to some of the comments made here, the general approach does work
quite nicely for other sports. You just need a different exponent to get
the results to work out. This shouldn't be terribly surprising, as the
pythagorean approach represents a very simple mathematical model of
winning and losing.
--
Raj (r...@alumni.caltech.edu)
Master of Meaningless Trivia (626) 585-0144
http://www.alumni.caltech.edu/~raj/
Roger Moore wrote in message <7iammj$66f$1...@fpage2.ba.best.com>...
>
>"David Brazeal" <nospamd...@sockets.net> writes:
>
>>Could the scoring explosion change the dynamics so much that things like
the
>>Pythagorean method don't work anymore? And how big would the scoring
>>increase have to be to do that? I'm not enough of a statistician to post
>>actual answers on here, but I have lots of questions...many of which might
>>not even be relevant or make sense to people who know what they're talking
>>about.
>
>It's unlikely to. Despite complaints about the "scoring explosion",
>present day scores are well within historic norms. Scoring in the AL is
>certainly lower than it was in the 1930's AL, and scoring in the NL is a
>bit lower than it was in the 1950's, much less the 1920's or 1890's.
That's interesting... could you let me know where I could find that? Is
there a good online site that would have league scoring data from those
eras? I'm still trying to complete my list of stathead bookmarks as I learn
some of this stuff.
davidb
>Roger Moore wrote in message <7iammj$66f$1...@fpage2.ba.best.com>...
>>It's unlikely to. Despite complaints about the "scoring explosion",
>>present day scores are well within historic norms. Scoring in the AL is
>>certainly lower than it was in the 1930's AL, and scoring in the NL is a
>>bit lower than it was in the 1950's, much less the 1920's or 1890's.
>That's interesting... could you let me know where I could find that? Is
>there a good online site that would have league scoring data from those
>eras? I'm still trying to complete my list of stathead bookmarks as I learn
>some of this stuff.
I don't have an explicit web reference because I've figured the stuff
myself using Sean Lahman's baseball database (http://www.baseball1.com)
rather than depending on pre-formed data. I can certainly post a
worksheet containing runs scored by league and season in tab-delimited
format if anyone wants; r.s.bb.data would be the obvious place, though,
rather than r.s.bb.a.
As a sidelight, I think that some misperception of league-wide scoring
levels is a product of using ERA as a proxy for runs per game. It's
certainly easier to look up league average ERA in Total Baseball or
wherever, and it does give you a good basis for comparing this season's
scoring with last season's. The problem is that it's actually a very
misleading way of comparing current scoring with that of fifty, seventy,
or a hundred years ago, because of the large change in error rate. As an
example, I saw a piece complaining that this year's AL was on a pace to
break the record (from the 1930 NL, IIRC) for highest seasonal ERA, and
that this was a sign that scoring was at an all time high.
I've found that, as a general rule, the percentage of unearned runs is
about five times the percentage of errors. That is to say that:
1 - ERA/RA = 5 * ( 1 - FA ) = 5 * E/TC
This formula (on a league wide basis) has actually held true for most of
the time since the NL was founded. The only notable exception was the
dead ball era, when the rate of unearned runs remained higher than
expected. I find that pretty remarkable, given that error rates have
changed by close to an order of magnitude over that timespan.
--
Raj (r...@alumni.caltech.edu)
Master of Meaningless Trivia (626) 585-0144
http://www.alumni.caltech.edu/~raj/
I haven't had enough caffeine yet this morning to follow the algebra
myself. Is it accurate to extrapolate from this that the ratio of errors
to unearned runs is 5 to 1 ? Or roughly, that 20% of errors lead to an
unearned run? (I know this isn't precise, because some errors lead to
more than 1 UER, but is this the general principle?)
Sean.
| Roger Moore wrote:
| > I've found that, as a general rule, the percentage of unearned runs is
| > about five times the percentage of errors. That is to say that:
| >
| > 1 - ERA/RA = 5 * ( 1 - FA ) = 5 * E/TC
|
| I haven't had enough caffeine yet this morning to follow the algebra
| myself. Is it accurate to extrapolate from this that the ratio of
| errors to unearned runs is 5 to 1 ? Or roughly, that 20% of errors
| lead to an unearned run? (I know this isn't precise, because some
| errors lead to more than 1 UER, but is this the general principle?)
No, that doesn't follow.
The idea is that if 2% of all chances are errors, then approximately
5 x 2% = 10% of all runs will be unearned. To rephrase the original
equation to turn it into straight ratios:
UR / R = 5 * E / TC (UR = unearned runs, R = runs)
So if you want a ratio of errors to unearned runs, you get this:
E / UR = TC / 5 * R
which is not a particularly interesting equation.
--
Dan Schmidt -> df...@harmonixmusic.com, df...@alum.mit.edu
Honest Bob & the http://www2.thecia.net/users/dfan/
Factory-to-Dealer Incentives -> http://www2.thecia.net/users/dfan/hbob/
Gamelan Galak Tika -> http://web.mit.edu/galak-tika/www/
>Roger Moore wrote:
>> I've found that, as a general rule, the percentage of unearned runs is
>> about five times the percentage of errors. That is to say that:
>>
>> 1 - ERA/RA = 5 * ( 1 - FA ) = 5 * E/TC
>I haven't had enough caffeine yet this morning to follow the algebra
>myself. Is it accurate to extrapolate from this that the ratio of errors
>to unearned runs is 5 to 1 ? Or roughly, that 20% of errors lead to an
>unearned run? (I know this isn't precise, because some errors lead to
>more than 1 UER, but is this the general principle?)
I don't think that either of these is accurate. The best result I can
come up with is to use a relationship between TC and IP. If you assume
that each inning requires k total chances (i.e. TC = k * IP), then you
can manipulate:
1 - ERA/RA = UER/R = 5 * E/TC
= 5 * E/(k*IP)
= 5/k * E/IP
UER/E = 5/k * R/IP
= 45/k * RA
UER/E = 1/k' * RA
What this implies is that the number of unearned runs per error increases
proportionally with scoring. This makes some sense; scoring goes up as a
result of high on base and slugging averages, and high OBA and SLG are
likely to increase the chance of a man who reached on an error of scoring,
or of an error which resulted in extra bases in allowing a man already on
base to score. OTOH, this implies that as the number of TC per inning
goes up, the number of unearned runs per error goes down. That seems
backward to me.
--
Raj (r...@alumni.caltech.edu)
Master of Meaningless Trivia (626) 585-0144
http://www.alumni.caltech.edu/~raj/
Roger Moore wrote:
> I don't think that either of these is accurate. The best result I can
> come up with is to use a relationship between TC and IP. If you assume
> that each inning requires k total chances (i.e. TC = k * IP), then you
> can manipulate:
>
> 1 - ERA/RA = UER/R = 5 * E/TC
> = 5 * E/(k*IP)
> = 5/k * E/IP
> UER/E = 5/k * R/IP
> = 45/k * RA
> UER/E = 1/k' * RA
>
<SNIP>
> OTOH, this implies that as the number of TC per inning
> goes up, the number of unearned runs per error goes down. That seems
> backward to me.
I think you've reached the point where the algebraic manipulations of the
formula are not useful. The formula says that a team that fields .980 will
allow 10% unearned runs, .970 will allow 15%, etc. This may work as a general
rule, but looking at
UER / RA = 5 * E / TC
we see that a team that allows additonal earned runs will lower the left side
of the equation but not change the right side. So, I'd be skeptical about
using this formula to do "related rate" problems.
Doug Olson wrote:
> > OTOH, this implies that as the number of TC per inning
> > goes up, the number of unearned runs per error goes down. That seems
> > backward to me.
>
> I think you've reached the point where the algebraic manipulations of the
> formula are not useful. The formula says that a team that fields .980 will
> allow 10% unearned runs, .970 will allow 15%, etc. This may work as a general
> rule, but looking at
>
> UER / RA = 5 * E / TC
>
> we see that a team that allows additonal earned runs will lower the left side
> of the equation but not change the right side. So, I'd be skeptical about
> using this formula to do "related rate" problems.
I ran numbers on errors per 9 innings vs. UER% (using major league team data), and
while my math background is too ancient to plot the graph anymore, there was a
definite "S" shaped curve to the relation, rather than a straight line. As
errors/9 innings pass 7, the UER% hits close to 67, and 9 errors/9 innings doesn't
raise the UER% much beyond that (it gets into the low 70s). As errors/9 innings
reaches .7, the UER% drops to 7 also, but doesn't drop much lower (it CAN'T drop
MUCH lower). It actually drops below 4% when the error rate drops to .51 errors/9
innings (Baltimore 1994), but that's a much flatter curve than when we're moving
between 1.5 and 3 errors per game per team.
Actual figures can be seen in a chart at
http://home.swbell.net/rockpf/baseball/uer_pct.html
>I ran numbers on errors per 9 innings vs. UER% (using major league team
>data), and while my math background is too ancient to plot the graph
>anymore, there was a definite "S" shaped curve to the relation, rather
>than a straight line. As errors/9 innings pass 7, the UER% hits close to
>67, and 9 errors/9 innings doesn't raise the UER% much beyond that (it
>gets into the low 70s). As errors/9 innings reaches .7, the UER% drops to
>7 also, but doesn't drop much lower (it CAN'T drop MUCH lower). It
>actually drops below 4% when the error rate drops to .51 errors/9 innings
>(Baltimore 1994), but that's a much flatter curve than when we're moving
>between 1.5 and 3 errors per game per team.
I did the plot (using Excel) and the graph doesn't look so much sigmoidal
as like a two population graph. As E/9 goes above about 2, the data
really flattens out, as if there's an elbow in the data. Still, a
regression indicates that there's pretty good correlation: a formula of
UER% = 0.0955 * E/9IP gives an R^2 of 0.676, which is pretty good. Even
better, though, is the relation between E/9IP and UERA. I found that
UERA = 0.610 * E/9IP give an R^2 of 0.839, which is pretty darn good.
--
Raj (r...@alumni.caltech.edu)
Master of Meaningless Trivia (626) 585-0144
http://www.alumni.caltech.edu/~raj/