Google Groups no longer supports new Usenet posts or subscriptions. Historical content remains viewable.
Dismiss

Predicting by record and by projected record

1 view
Skip to first unread message

Bob-Nob

unread,
Oct 4, 2004, 5:08:04 PM10/4/04
to
After fiddling around with the Bill James Postseason Prediction
System, I decided to go through postseason results, paying attention
to differences in (Pythagorean) projected won-lost records and actual
won-lost records. My stat-savvy failed me after coming up with the
raw results -- I'd love it if someone tells me where to go from here...

The following is a chart of results for the teams with the better
records, ordered by difference in record. Thus, when 'Projected' has
a 2-1 next to the 19 games difference, that means that teams whose
projected records were 19 games superior to their opponents won 2
series and lost 1. (I'm typing the whole thing in by hand here -- if
someone wants the specific results, they can email me -- take out
SPAM from my email address).

Games
Projected difference Actual
26 1-0
25.5
25 1-0
0-1 24.5
24
23.5
23
22.5 0-1
22
1-0 21.5 1-0
3-0 21
20.5
20 1-1
19.5
2-1 19
18.5
1-1 18
1-0 17.5
1-1 17 2-0
16.5 1-1
1-0 16 3-0
15.5 2-0
2-1 15 3-0
1-2 14.5
1-2 14 3-1
3-1 13.5
6-2 13 2-2
3-0 12.5 1-0
3-1 12 6-2
0-1 11.5 2-2
3-2 11 2-0
0-2 10.5 4-0
3-1 10 4-5
3-2 9.5 1-3
4-5 9 0-2
1-3 8.5 2-2
8-2 8 4-2
0-2 7.5 3-1
6-5 7 6-5
2-2 6.5 1-3
3-2 6 8-5
1-1 5.5 1-7
5-4 5 5-3 (68-47 projected, 70-48 actual)...

Through a difference of 5 games, actual record appears to be a
better predictor of success than projected record. At this point,
however, it flips. I don't have the stats wherewithal to analyze the
numbers to know whether this shows actual record is superior (with
too much noise at the lower differences), whether projected record is
superior (with too small a sample size at the larger differences),
or whether there is something else going on...

Games
Projected difference Actual
6-3 4.5 2-2
6-3 4 5-9
5-2 3.5 4-2
7-6 3 6-10
5-2 2.5 2-1
11-4 2 9-8
5-1 1.5 3-4
7-9 1 5-9
2-2 0.5 1-2 (122-79 projected, 107-95 actual)
6 pairs EVEN 5 pairs

Thoughts, comments, questions?

Catch you later.
--Robert Machemer

--
Robert Paul Aubrey Machemer | "For each time he falls, he shall
Amherst College, Math & Classics | rise again, and woe to the wicked!"
IF1, IF3, IF9: best films, cast | --Don Quixote (Man of La Mancha)
(What are YOU doing this weekend? See IF12 on May 23rd, 2004)

Bob-Nob

unread,
Oct 5, 2004, 6:36:56 PM10/5/04
to
Yes, the postseason is a crapshoot, so perhaps the idea of
predicting outcomes is uninteresting to the resident statheads.
I myself am interested -- for instance, it has been suggested
that the 2004 Yankees in-game strategy has been designed to
maximize the likelihood of winning close games (which would
likely tend to give them a much better actual record than
Pythagorean record)... does this put them at an advantage or
disadvantage for the postseason? Are the Red Sox (even before
getting into W3% and all that, one of the better Projected teams,
though one that has not done as well at maximizing wins out of
component stats) better suited for postseason play? And so
forth.
I'm more than happy to do the math myself if someone can
suggest what I should be doing. I don't know how to tackle the
numbers from here. How does one take the results and properly
weight them so that one can attack the following problem:

If team A is projected to be X games better than team B, what is
the likelihood of A's winning the series? If team C is Y actual
games better than team D, what is the likelihood of C's winning?
If X = Y, is A more likely to win than C? Would A always be
more likely to win?

As I say, any help is appreciated. Catch you later.

igor eduardo küpfer

unread,
Oct 5, 2004, 8:24:43 PM10/5/04
to
On 5 Oct 2004 18:36:56 -0400, rpmac...@nospamte.amherst.edu (Bob-Nob) wrote
in <4163...@amhnt2.amherst.edu>:

> Yes, the postseason is a crapshoot, so perhaps the idea of
>predicting outcomes is uninteresting to the resident statheads.
>I myself am interested -- for instance, it has been suggested
>that the 2004 Yankees in-game strategy has been designed to
>maximize the likelihood of winning close games (which would
>likely tend to give them a much better actual record than
>Pythagorean record)... does this put them at an advantage or
>disadvantage for the postseason? Are the Red Sox (even before
>getting into W3% and all that, one of the better Projected teams,
>though one that has not done as well at maximizing wins out of
>component stats) better suited for postseason play? And so
>forth.
> I'm more than happy to do the math myself if someone can
>suggest what I should be doing. I don't know how to tackle the
>numbers from here. How does one take the results and properly
>weight them so that one can attack the following problem:
>
>If team A is projected to be X games better than team B, what is
>the likelihood of A's winning the series? If team C is Y actual
>games better than team D, what is the likelihood of C's winning?
>If X = Y, is A more likely to win than C? Would A always be
>more likely to win?
>
> As I say, any help is appreciated. Catch you later.
>

I've had some success modeling NBA playoff series using the negative binomial
distribution, but I don't think that would work here because one of the
assumptions is violated, viz the probability of success is constant from one
game to another. The actual game-to-game probabilities are (I think) heavily
dependent on the starting pitchers.

To come up with the probability of a team winning a series, I think you would
have to start by calculating the team's probability of winning each
pitcher-to-pitcher matchup, then summing the probabilities of the 9 paths a
team can follow to win a 5-game series.

--

-------------------------------------
| best, | Sticking it to |
| ed | The Man since 1971 |
-------------------------------------
Watch the spam trap -- the domain is rogers

Danil

unread,
Oct 6, 2004, 12:20:44 AM10/6/04
to
igor eduardo küpfer says...

> I've had some success modeling NBA playoff series using the negative binomial
> distribution, but I don't think that would work here because one of the
> assumptions is violated, viz the probability of success is constant from one
> game to another. The actual game-to-game probabilities are (I think) heavily
> dependent on the starting pitchers.
>
> To come up with the probability of a team winning a series, I think you would
> have to start by calculating the team's probability of winning each
> pitcher-to-pitcher matchup, then summing the probabilities of the 9 paths a
> team can follow to win a 5-game series.

10 paths ( 5 choose 3 ):

123
124
134
234
125
135
145
235
245
345

Danil

unread,
Oct 6, 2004, 1:13:36 AM10/6/04
to
Bob-Nob says...

> I'm more than happy to do the math myself if someone can
> suggest what I should be doing. I don't know how to tackle the
> numbers from here. How does one take the results and properly
> weight them so that one can attack the following problem:
>
> If team A is projected to be X games better than team B, what is
> the likelihood of A's winning the series? If team C is Y actual
> games better than team D, what is the likelihood of C's winning?
> If X = Y, is A more likely to win than C? Would A always be
> more likely to win?

The general formula is straight forward, but actually making the right
guesses to get a reasonable answer may be tricky.

Let W(x,y) be the winning percentage of a team with true talent x vs a
team with true talent y. Think log5, or something like it.

Let I represent the information available about the two teams.

W(I) = Integral Integral W(x,y) p(x|I)p(y|I) dx dy
The integrals range over all possible levels of talent (if we use true
winning percentage here, from 0 to 1), and p(n|I) is the probability of
true talent n given the information I.

In other words, I contains information like "this team had a .680 winning
percentage and a .700 projection", and p answers the question "what is
the probability that a team with true talent of .690 would have a .680
winning percentage and a .700 projection"

How do you find p? You sim a lot of teams with a fixed level of true
talent, and plot how the distribution of wins and projections comes out,
then reverse it.


BUT, I'm not sure this quite answers the question that you intend;
because the problem, as you describe it, really doesn't address the
Redsox Yankees comparison you posited.

What I think you want is a game model where the expected scoring is a
function of the score. Ie, team A averages 5/9 runs per inning and plays
at that level constantly; team B also averages 5/9 runs per inning, but
scores more frequently when the game is close and less frequently when
the game breaks open.

Here the right answer is to start with a run scoring distribution (I'd
recommend tango's), make the simplifying assumption that scoring rate
only changes between innings. Then sim it - for each time up, figure out
the scoring rate to use for that inning, generate a number of runs, move
onto the next inning. Don't forget to calibrate that varying offense so
that it generates the average run scoring you intend.

There is an exact solution, of course, but it again involves more
integrals than I would like to think about - my guess is that it is
easier to put together a sim and come up with a reasonable approximation
than it is to set up the equations for calculating an exact result.

Danil

0 new messages