Let's consider the following example, which is a little contrived to
make it simple. However, the principles remain the same in other pools.
Suppose there are 4 1500's, A, B, C, D. They are all established
players. Their ratings are stable. To make the calculations easy, we
will assume that we will calculate rating changes under the old Elo
formula with K=32. Suppose that they are the only four players in the
rating pool. The average rating of the pool is, therefore, 1500.
Elo recognized that simply having an improving player can cause
deflation.
Let's suppose that A decides to study for awhile. As a result, his
strength increases to a degree that on average he scores 3 out of 4
against B, C, and D. These odds represent roughly a 200 rating point
spread.
What we would want the pool to do is this: Since B, C, and D are the
same strength as before, their ratings should stay at 1500. A should
see his rating go toward 1700.
That is, their performances inidcate a strength relative to where they
started of 1500 for B, C, and D and 1700 for A.
Let's suppose that they play 10 rated games against each opponent (30
total.) B, C, and D score 50% against each other, but only 25% against
A, exactly as outlined above. That means B, C, D and E win 12.5 games
each, and A wins 22.5 games.
What are their ratings (assuming for ease that we rate this as 1 event)
at the end of these encounters? (This example is simplified, but
illustrates the point, and the prinicples hold true even if we treat it
as several events.)
The rating formula is:
(W-We) x 32 + Rating old = Rating new
Since all players started at 1500, we expect them all to score 50%.
The winning expectancy, We, therefore = 15 for all the players.
For B, C, D:
(12.5-15.0) x 32 +1500 = 1420
12.5 points, for B for example, is 5 points against C, 5 against D, and
2.5 against A.
For A:
(22.5 - 15.0) x 32 + 1500 = 1740
What is the average rating of the pool? (1420 + 1420 + 1420 +1740)/4 =
1500.
Hmm...exactly the same as before.
Yet, B, C, D are all rated LOWER than their actual skill level of
1500. And even if A loses his "40 extra" points back to the pool
fairly evenly in another series of games, we would see:
A: 1700
B: 1433
C: 1433
D: 1433
That is, 75% of the players in the pool would be deflated, by 67 points
each, even though the average rating of the entire pool is unchanged.
End the myth. Average rating of the pool does not indicate deflation.
--
Kevin Bachler
Caveman
"Caveman chess is chess without finesse."
Sent via Deja.com
http://www.deja.com/
> A common misconception is that the change in the average rating of a
> pool can indicate whether it is deflated.
> That is, 75% of the players in the pool would be deflated, by 67 points
> each, even though the average rating of the entire pool is unchanged.
>
> End the myth. Average rating of the pool does not indicate deflation.
Actually what you showed was that: Reduction in the average rating of the
pool is not necessary for deflation to have occurred. It still might
"indicate" deflation, i.e. provide some evidence without being absolutely
neccessary or providing absolute proof.
- Tom Martinak
Ok, two new players of 1000 strength enter the pool of 4 1500s. They
each lose 4 games, one ot each 1500, garnering ratings of 1100 each.
The 1500's in return each gain about 1 rating point.
The average rating of the pool is now [(1501 x 4) + (1100 x 2)]/6 =
1367 as the average rating. Yet none of the 1500's is deflated, and
both of the 1100's are inflated.
End the myth.
If you reverse all these results, you get an average of 1500 still, but 75% of the players would be inflated, but there
is no inflation. But while your statement is true, that begs the real questions:
1. Is there really deflation when a bunch of young players enter a pool with a very low rating? Of course, because they
entered the pool with a very low rating, far below the average of the pool, thus creating a new pool with a much lower
average. So even if the points are thereafter equal (no bonus points, etc.), the average of the new pool is lower.
This is what happens in USCF when the percent of scholastic players grows rapidly, and
2. Have USCF ratings been deflating recently? Of course. The experienced players who remained in the pool were about
as good as before, but they had points taken away when the scholastic players started gaining points at their expense.
In our chess club it is quite rempant, but I see it almost everywhere. This also happened prominently in the Fischer
Boom of the 70's when many new players came in. This time the cause was, partly, the movie, Searching for Bobby
Fischer, which brought in many new scholastic players who had been unaware of tournament chess until the movie came out.
So you are correct; just because we measure a population to have more lower players, it may not be deflating; however,
in this case it is, and when you are dealing with a very large population, the kind of example you showed tends to even
out with some higher and some lower, as I showed at the top.
Finally,
> That is, 75% of the players in the pool would be deflated, by 67 points
> each, even though the average rating of the entire pool is unchanged.
>
> End the myth. Average rating of the pool does not indicate deflation.
You meant the median rating went down, but the mean did not. In your example, the mean stayed the same even though the
median changed. Using the word "average" here instead of "mean", "median" or "mode" is misleading.
Regards,
NM Dan Heisman
USCF Delegate
Sorry, Tom. I had missed part of what you intended. You are using
indicate in terms more like correlate. I need to think about that a
little, but I still think its probably not true. An influx of members
can easily lead to inflation (some portions of the scholastic
membership is inflated) while lowering the average rating.
I guess I would agree that it's less likely that much of the pool is
deflated if the average rating is increasing, since a relatively small
portion of new players are likely to enter the pool with an initial
rating above average.
But again, I would need to think about this a little more. Perhaps Ken
Sloan or Tom Doan can comment on this more definitively.
??
I don't understand what you were intending to say here. But let me
take a stab at something. If B, C, and D are rated above 1500 in the
example you are giving, yes, they are inflated, and yes there is
inflation.
Inflation/deflation refers to the actual rating of a player when
compared against a "targeted" performance level.
Elo's point is that a 1500 today, or a master today, should also (in
common thought) be a 1500 tomorrow. Otherwise ratings would only be
relative.
Without this "external" meaning, the fact that we want a 1500 to mean
the same thing from day to day, there could not be inflation or
deflation.
But he saw that when a player improves, as in the example I gave, he
necessarily causes the other players to deflate. Consequently, he
suggested various means to combat deflation. Among them was a lower k
and also bonus points.
> But while your statement is true, that begs the real questions:
>
> 1. Is there really deflation when a bunch of young players enter a
pool with a very low rating?
Not implicitly, no. See your next question.
> Of course, because they
> entered the pool with a very low rating, far below the average of the
pool, thus creating a new pool with a much lower
> average. So even if the points are thereafter equal (no bonus
points, etc.), the average of the new pool is lower.
> This is what happens in USCF when the percent of scholastic players
grows rapidly, and
>
> 2. Have USCF ratings been deflating recently? Of course. The
experienced players who remained in the pool were about
> as good as before, but they had points taken away when the scholastic
players started gaining points at their expense.
The key here is to note that deflation is not caused by the fact that
the players started with a low rating. It was caused by their
improvement. Had they not improved, the average rating of the pool
would have stayed lower, but deflation would not have occured.
> In our chess club it is quite rempant, but I see it almost
everywhere. This also happened prominently in the Fischer
> Boom of the 70's when many new players came in.
Again, the entry of new low rated players did not cause deflation.
Deflation occured when the players improved.
> This time the cause was, partly, the movie, Searching for Bobby
> Fischer, which brought in many new scholastic players who had been
unaware of tournament chess until the movie came out.
>
> So you are correct; just because we measure a population to have more
lower players, it may not be deflating; however,
> in this case it is,
Not really, there is a subtle extra step. The improvement of those low-
rated players is a required step.
> and when you are dealing with a very large population, the kind of
example you showed tends to even
> out with some higher and some lower, as I showed at the top.
>
Actually, I don't get this last point.
> Finally,
> > That is, 75% of the players in the pool would be deflated, by 67
points
> > each, even though the average rating of the entire pool is
unchanged.
> >
> > End the myth. Average rating of the pool does not indicate
deflation.
>
> You meant the median rating went down, but the mean did not.
Generally "average" is taken to be the arithmetic mean.
> In your example, the mean stayed the same even though the
> median changed. Using the word "average" here instead
of "mean", "median" or "mode" is misleading.
>
> Regards,
> NM Dan Heisman
> USCF Delegate
>
Ah...NOW I see why you are so confused. You see, when you lose games, and your
rating goes down, we don't call that deflation. Deflation is when your rating
is lower than it is supposed to be, meaning you are playing at a certain level,
but your rating does not reflect this. If you lose games, your rating is
SUPPOSED to go down, and in fact, player Bs strength with respect to A is
really 1433, even though he might like it to be higher. We won't even get to
the fact that your examples are "simplistic" because they give a distorted
perception of reality. Obviously most players don't gain and lose points at
the expense of the ENTIRE population of chess players, as is suggested in your
example. Hope this clears things up for you.
Let's define some terms before we disagree:
First, a "pool" in this sense can either be:
1) A group of all players "A" at time X and then just "A" at time Y (later) even though there are others, or
2) A group of all players "A" at time X and then all "B" at time Y
The USCF pool is the latter, while your example is the former. That is OK, but a difference.
Then, there are three mathematical definitions of "average":
1) "Mean" - the sum divided by the number - as the average of 1, 3, 8, 9, 9 is 30/5 = 6
2) "Median" - the number which half are larger and half are smaller 1, 3, 8, 9, 9, the median (middle) is 8
3) "Mode" - the number that occurs the most - in the above example, 9
So now, we can define "deflation" in your example as "pool" definition 1 with the median (you said "most players go
down, so that means the median goes down while the mean stays the same). This is possible, but what does it prove other
than it might be possible? - It only serves as a counterexample to other possibilities.
But in the USCF it is pool definition 2, which means new players are arriving (and maybe old players leaving). So if
new players come in with low ratings and the rest of the pool roughly stays the same, then both the mean and median go
down.
I think the USCF considers deflation if the mean (or median) of existing players go down, and this is happening as the
new players improve, which means their mean and median go up and the old players must go down to maintain the same
number of points in the pool, assuming no "bonus" or "activity" points. That is why the rating committee had to propose
a system with bonus points to "inflate" the pool.
Regards,
Dan H
>End the myth. Average rating of the pool does not indicate deflation.
_______________
That's certainly true, and not only for the reasons related to your example,
but for a lot of other reasons as well.
A real-life example occurred around 1973 because of the Fischer boom. Huge
numbers of inept players joined the rating pool, reducing the average rating
of the pool. This was mistaken for deflation. The result was the original
"fiddle points" to get the average rating back up to where it was.
Since there had not really been deflation, the original fiddle points did
not correct deflation, but rather caused inflation. By the time the
powers-that-be recognized this, everybody's rating was so inflated that
people's concept of ratings had changed. The "typical" stodgy, booked-up
player that used to be thought of as an A player was now thought of as an
Expert.
So, when the inflation caused by fiddle points was reversed, it now looked
like deflation to many. And the pool was, indeed, deflated -- compared to
the inflated pool resulting from fiddle points. Perhaps the players should
then have revised their concept of ratings back to what it had been -- but
nobody likes to be thought of as an A player when he was once thought of as
an Expert.
That's how the current mess began. Now it's difficult -- especially
politically -- to separate genuine deflation from perceived deflation.
The ratings committee is pretty good at figuring out which is which -- and
ought to be listened to, carefully, by the mathematical amateurs at the top
of the political pyramid.
Bill Smythe
> A real-life example occurred around 1973 because of the Fischer boom. Huge
> numbers of inept players joined the rating pool, reducing the average rating
> of the pool. This was mistaken for deflation.
This is not necessarily true. As the low rated players improve and gain rating points, they have to get it from
somewhere. The somewhere is the established players. For example, If you have 4 1500 players and a 500 player joins
the pool, assume he improves to 1500 by playing the others and they stay about the same. Then the pool has 6500 points
for 5 players and the four 1500 players deflate to 1300 as all become equal in strength. We can argue about what
happened in 1973, but the above example is real deflation, which I have seen a lot lately here in Philadelphia.
Regards,
NM Dan Heisman
>
> I think the USCF considers deflation if the mean (or median) of
> existing players go down, and this is happening as the new players
> improve, which means their mean and median go up and the old players
> must go down to maintain the same number of points in the pool,
> assuming no "bonus" or "activity" points. That is why the rating
> committee had to propose a system with bonus points to "inflate" the
> pool.
This is incorrect. I know of no member of the USCF Ratings Committee
who believes that "considers deflation if the mean (or median) of
existing players go down".
For the reasons you outlined, measuring the population mean, or median,
or mode is essentially worthless as a way to track inflation or
deflation.
Offhand, I can think of two methods used by the Ratings Committee to
look at deflation. One involves tracking the distribution of ratings
for a sub-population which is *believed* to be relatively immune to
"social effects" (players joining/leaving the pool - such as the huge
influx of lower rated players). For example, we have looked at players
in the age range 35-45 with established ratings. [note: population
dynamics appear to affect even *this* population - you can see the baby
boom march it's way through the group]. Another method involves
tracking individual ratings histories.
But, again - no one on the Ratings Committee (to my knowledge) advocates
watching the population mean/median/mode and tweaking the system to keep
it stable. That sort of misconception died out in the 1980's (well,
sometimes the people screwing up things in the 1980's come back into
power and start doing it again, but don't blame the Ratings Committee
for that!)
Bonus points do tend to re-inflate the pool - but that's not *really*
the primary reason to have them. The primary goal of ratings is to
predict future performance. Bonus points are used when the system
"notices" that a particular player is grossly underrated. Of course,
the proper response to an underrated player is to simply raise that
player's rating.
It is *desirable* to monitor the system and correct for systemic
deflation (OR inflation!). But, it's *essential* that individual
player's ratings are accurately positioned with respect to other players
in the pool.
Bonus points help to re-inflate - but they directly address the problem
of "obviously" underrated players. This is good. Other "inflation"
methods are not so good...
--
Kenneth Sloan sl...@uab.edu
Computer and Information Sciences (205) 934-2213
University of Alabama at Birmingham FAX (205) 934-5473
Birmingham, AL 35294-1170 http://www.cis.uab.edu/info/faculty/sloan/
Who here said it did?
A better example is SAT scores....so long as the pool of SAT takers remains the
elite college-bound group...average score will be high.
If the pool starts to include EVERY student...including those who do not take
college prep classes....the average score of the pool will plummet. By
necessity...because the pool is including more persons with a probability of
low scores than before.
Same as when USCF started seeing high growth in scholastic chess...most
scholastic/beginning players are very low chess skill...so you see some drag on
the pool average.
Back to your example...
>
>Suppose there are 4 1500's, A, B, C, D. They are all established
>players. Their ratings are stable.
I am guessing you have taken these four 1500 ratings from a larger
pool...separated them...and declared this pool to be a separate
population..yes?
>
>Elo recognized that simply having an improving player can cause
>deflation.
Yes, he did.
>
>Let's suppose that A decides to study for awhile. As a result, his
>strength increases to a degree that on average he scores 3 out of 4
>against B, C, and D. These odds represent roughly a 200 rating point
>spread.
Parenthetically, I would also note that all FOUR players might study the same
amount...raise their games to unimaginably high levels...play games of
remarkable clarity and beauty...and they would see NO CHANGE in their ratings
at all. Why? Because their rating pool only tracks relative performance.
When they all played at a low level (1500...measured by some other, larger
pool)...they scored the same against each other by playing low quality games.
When they all became super-Kasparov's...they still score the same against each
other...by playing super high quality games. Result - same 1500 rating.
And it is accurate...within their pool.
Back to your example.
>
>What we would want the pool to do is this: Since B, C, and D are the
>same strength as before, their ratings should stay at 1500.
Only if you are going to make comparisons between the four players in your new
"small" pool...and the larger world pool you plucked them from.
If you only study and care about the small pool of four...your statement is not
quite correct.
>
>Yet, B, C, D are all rated LOWER than their actual skill level of
>1500. And even if A loses his "40 extra" points back to the pool
>fairly evenly in another series of games, we would see:
>
>A: 1700
>B: 1433
>C: 1433
>D: 1433
That's because the small pool quite correctly shows the results of the study on
the part of the improving player and the sloth on the part of the other
three...the one who studies now towers over the other three by a super
margin...
The small pool performed correctly...by showing the relative differences of the
group.
You are correct, however, that the "pool administrator" should be injecting
points into the pool on a regular basis.
>
>End the myth. Average rating of the pool does not indicate deflation.
Correct as far as it goes.
But...average rating trending down CAN tell you that a large, generally low (or
high) skill segment has entered the pool.
Eric C. Johnson
This is precisely the point that RC members make when they oppose adding fixed
amounts to all records in the database...that the system is relative and they
really don't care whether it matches the outside tags of A-player, B-player,
C-player that are attached to it.
All they care about is predictive power...if the scale trends downward...but
predictive power is retained in all cases, the RC is fat and happy.
Otherwise, there would have been little opposition to a 50 pt scale
correction...either over the overall database or for just those rating segments
that seem to need it.
Elo may have defended the outside scale tags...and the need to keep 1500 today
"meaning" the same thing as before, but the RC does not seem to do so.
I point out that "means" in this case is misleading.
Most readers see 1500 and assume it means "C-player" or "has the playing level
or attributes of a C-player"...but what I think Kevin B (and Elo) mean here
would be something akin to "scores the same percentage against players rated
1400 as before"...
...which also avoids the issue of quality of play.
The pool might trend downward to the point where GMs have 1500 ratings...and so
long as the downward pressure is uniform...so that 1500-rated GMs still score
the same against 1400-rated IMs in the reduced pool....
...as 1500-rated C-players used to do against 1400-rated C-players...
..then predictive power is retained. Strength of play/quality of games is
irrelevant to the rating system.
Tagging rating classes to the objective qualities of specific player-types is
also something Elo supported (i.e., masters of yesteryear should equal masters
of today, etc....so that if there are more master-players today thene the scale
needs constant adjustments to reflect that fact)...
...but all too often, we see a view that ratings are merely percentile
rankings...and that if you happen to be the lowest master in a pool of GMs you
may earn a very low rating...and so long as predictive power is retained,
everyone is happy. Never mind that your 1200 rating in a pool of masters...is
not anything like a 1200 rating in a pool of beginners.
Etc.
>
>Without this "external" meaning, the fact that we want a 1500 to mean
>the same thing from day to day, there could not be inflation or
>deflation.
Again, what you are calling "external meaning" appears to be something akin to
predictive power...as opposed to reflecting the real-life attributes of player
types.
I take it as a given that there are real, discoverable player types such as
C-player, Expert, quasi-Expert, master, GM, etc.....separate and apart from the
rating system that seems to identify them.
If the rating system starts to fail to correctly identify these types...not
just by how they score in tourneys but by how they play...then there is
inflation/deflation....of a type OTHER THAN pure predictive power.
This is a point that Elo makes very subtly in his book....when he discusses
tags such as "strong club player = 2000", etc.
Eric C. Johnson
Indeed, Kevin confuses two very different facts about the USCF universe.
Fact 1: Wins/Losses/Draws are assumed to show strength....and so if a player
loses games his rating goes down. Often the RC will go to great lengths to say
that such a lower rating is quite appropriate, because they zoom in on the idea
that only performance equals or defines strength, therefore, reduced
performance equals lower strength equals lower legitimate rating.
The concept that a player's "strength" might remain constant despite a string
of results is quite foreign to the RC line of thinking...because in that view,
the rating system is constantly calculating better and more accurate estimates
of player strength based on temporary success.
It can be very difficult to convince the RC that a player's strength is
constant even though they may be losing some games....yet most layman see this
as self-evident.
Example...I play 10 games vs. GM Korchnoi. I lose all 10. My rating will go
down...even though my strength is exactly equal to where it started...it may
even have risen due to my experience in playing those 10 games. Thus, the
"pain" most players feel when they must play up in the first round of a Swiss!
In the RC-world, my "strength" was shown to be lower than my initial
rating...base on wins/losses/draws...so my new lower rating is more
"appropriate"...etc. I guess I "got weaker" by playing a match with Korchnoi!
This is also true for GM Korchnoi....who might lose 20 games straight..yet on
game 21...he's still Korchnoi!
Fact 2: The rating system has independent tags such as A-player, B-player,
C-player attached to it...and some of us believe that these player types do, in
fact, exist independently of the rating system.
Club players are different from experts and masters, who are different from
GMs.
The rating system needs to be able to correctly identify and segment these
independently-existing types....quite apart from pure predictive power....or it
fails in a very important respect (under this view).
Allowing a pool to deflate to the point where GMs carry 1500 ratings, masters
are 1100, and club players are 800...might not matter in terms of pure
predictive power. But it matters a great deal in terms of identifying the real
player types in the world....over a stretch of historical time.
Thus, retaining relative predictive power of the system is quite separate from
maintaining the independent validity of the attached tags (A-player, B-player,
etc.)....which spring forth from other characteristics of players (i.e., how
they play, quality of moves, how they tend to lose)...other than pure
wins/losses/draws.
Otherwise, deflation/inflation would be no problem at all for the system!
****
I *do* find it interesting that Kevin makes reference to the independent
standing of these four 1500 players (he claims they are "established" 1500s and
that their strength appears to be constant...judged how? by reference to some
independent, non-rating-system-based measure?)....and therefore seems to accept
the notion that relative predictive power is not the only point of the rating
system...
..but then he goes to great lengths in previous posts (months ago)...where he
focused almost entirely on predictive power as being the ONLY point of ratings.
That predictive power (constant over time) was the "meaning" of ratings.
There is a POV that begs to differ...including Elo, apparently, since Elo makes
reference to the idea that a strong club player should be approx. 2000...today,
yesterday, and tomorrow.
Assuming that "strong club player" is a synonym for "player with skills X, Y,
Z"....and not one for "just the top player in any club"....we can see that even
Elo acknowledged the view that there are independent qualities, apart from
wins/losses/draws...that determine the anchor points for the rating system.
Thus, this is why Elo was concerned about deflation...because pure predicitve
power issues are not, by themselves, sufficient to keep the moorings of the
rating system intact.
One must also ask "are my C-players today playing chess of the same quality as
the C-players of yesteryear"...not just "are my C-players scoring the same
against D-players as they did 10 yrs ago"...
Why?
Because in the first case, you are making sure that quality of play/skill set
is constant.
In the second, you could have pool-wide deflation to the point where 1500s were
playing GM-level chess, and predictive power would be retained
completely...thus, masking important other changes.
Eric C. Johnson
>
>Since there had not really been deflation, the original fiddle points did
>not correct deflation, but rather caused inflation.
I would note for the record that once 10 or 15 yrs passed...it was really
STUPID to try to correct such inflation.
Why? Because your player pool had changed and...even more importanly...player
perceptions about the outside tags (D-player, B-player, Expert) attached to the
system had changed.
Forcing that population to swallow the idea that they were all
"inflated"....compared to some magic average of 15 yrs prior...was a bitter
pill...
...rather akin to forcing the USA back to the gold standard because the price
of a loaf of bread is higher/lower today than it was in 1973...relative to
other goods.
In a word...such radical corrective measures...after years and years of
non-attention to rating pool issues...create their own consumer perception
problems.
Where would you want to correct the pool to?
1950?
1960?
1970?
1980?
1990?
Telling the entire USCF pool "you are overrated and we are fixing that" is a
prime example of not knowing your customer base.
Basing the decision on the idea that FIDE /USCF ratings were growing apart was
a poor rationale for such change...in large part because the change was going
to be perceive negatively by most consumers (i.e.., ratings were gonna go
down).
In the present day...there is a growing group who want to see some modest
INFLATIONARY corrections to USCF's rating pool...prior to the new formula
effectively locking everyone into place by using reduced K for higher rated
players.
Yet there is resistance to such modest inflationary corrections by the RC and
others..why?
In part...so they argue...because going back to that "bad-old-time" of 1990 or
so...is wrong...because we were all "inflated" back then.
In other words, the last 10 yrs of deflation is "good for us" because it brings
us closer to the "pure times" of the 1970s...instead of the "bad inflated times
of 1990"..etc.
Hogwash.
> By the time the
>powers-that-be recognized this, everybody's rating was so inflated that
>people's concept of ratings had changed. The "typical" stodgy, booked-up
>player that used to be thought of as an A player was now thought of
>as an
>Expert.
Yes, and you have formulated things very succinctl, Bill. It was misguided to
try to "change everyone's perceptions" back to 1973-level in 1985.
>
>So, when the inflation caused by fiddle points was reversed, it now looked
>like deflation to many. A
And it was!
See my gold standard example at the beginning. Tell me, do most consumers
react well (initially) when their governments slash the value of their
currencies in attempts to go back in time?
>Perhaps the players should
>then have revised their concept of ratings back to what it had been -- but
I note for the record that today's USCF pool is radically different...in that
it has a large deflated pool of scholastic players.
>
>That's how the current mess began. Now it's difficult -- especially
>politically -- to separate genuine deflation from perceived deflation.
A very insightful comment...and one that should give pause to those who stress
only the mathematical purity of the rating system.
>
>The ratings committee is pretty good at figuring out which is which -- and
>ought to be listened to, carefully, by the mathematical amateurs at the top
>of the political pyramid.
False...see above.
Eric C. Johnson
Ken Sloan answered part of this, but I wanted to address another part...
> Then, there are three mathematical definitions of "average":
> 1) "Mean" - the sum divided by the number - as the average of 1, 3,
8, 9, 9 is 30/5 = 6
> 2) "Median" - the number which half are larger and half are smaller
1, 3, 8, 9, 9, the median (middle) is 8
> 3) "Mode" - the number that occurs the most - in the above example, 9
Only the mean is considered a true average, and it is the arithmetic
average not the geometric average. The other terms, median and mode,
are in fact measures of central tendency, but are not an average.
> So now, we can define "deflation" in your example as "pool" definition
Deflation occurs in an individual. A pool exhibits it only as systemic
deflation, defined by Elo as a gradual downward trend of all ratings,
including those whose proficiency remains stable.
Elo then states clearly: "Deflation becomes more acute with a greater
percentage of new and improving placers in the pool..."
1 with the median (you said "most players go
> down, so that means the median goes down while the mean stays the
same).
There are several other factors that can have larger effect than
deflation that cause the same result. That is why deflation is not
measurable in this way.
This is possible, but what does it prove other
> than it might be possible? - It only serves as a counterexample to
other possibilities.
The example I provided shows the standard pattern of deflation. It is
much like a chess puzzle that illustrates a simple mating pattern. It
was not intended to show the full richness of the position, just to
convey the concept.
> But in the USCF it is pool definition 2, which means new players are
arriving (and maybe old players leaving). So if
> new players come in with low ratings and the rest of the pool roughly
stays the same, then both the mean and median go
> down.
Deflation is when proficiency exceeds rating (systemically.)
>
> I think the USCF considers deflation if the mean (or median) of
existing players go down,
No, it doesn't. Ken Sloan addressed this misconception.
> Regards,
> Dan H
> All they care about is predictive power
No...predictive power is a high concern, but not the only one. If so,
the master title would fluctuate from 2200.
I disagree, I think the RC does care about the labels.
And now I see why you are confused. The 3 players in this example were
stated a priori to be 1500's and not have their strength change. 1
player improved.
Consequently you have 3 players at 1500 strength who are rated 1433.
This is an old problem, well known, and written about by Elo. Please
don't tell me I'm confused about something so well documented.
> If you lose games, your rating is
> SUPPOSED to go down, and in fact, player Bs strength with respect to
A is
> really 1433, even though he might like it to be higher.
Actually, its not. A 200 point difference represents winning 75% of
the time. The gap is more than 200 points.
B is deflated.
> We won't even get to
> the fact that your examples are "simplistic" because they give a
distorted
> perception of reality.
No, just a simplistic one, to illustrate the issue.
> Indeed, Kevin confuses two very different facts about the USCF
universe.
>
> Fact 1: Wins/Losses/Draws are assumed to show strength....and so if
a player
> loses games his rating goes down. Often the RC will go to great
lengths to say
> that such a lower rating is quite appropriate, because they zoom in
on the idea
> that only performance equals or defines strength, therefore, reduced
> performance equals lower strength equals lower legitimate rating.
>
Yes, and since A was winning 75% of the time, and B, C, and D did not
have their performances relative to each other change, that indicates
that Ais now 200 points stronger than B, C, and D.
But you will note the delta is MORE than 200 points.
There is no confusion here at all Eric. But this simple example is
making great points at how poorly people understand the system.
> The concept that a player's "strength" might remain constant despite
a string
> of results is quite foreign to the RC line of thinking.
Agreed. The performance actually remained constant AS SHOWN by the
string of results.
Look harder.
Ablue for 1.
> .... As the low rated players improve and gain rating points, they have
to get it from
>somewhere ....
_____________
The above statement, and the details you provided to back it up, certainly
constitute a valid explanation of rating deflation caused by the learning
process.
My point, however, was that the fiddle points following the Fischer boom
were introduced specifically to raise the *average* rating among USCF
tournament players. I even remember a phrase, in Chess Life, referring to
the difference between former and current "median" ratings as a
justification for fiddle points.
Specifically, if ten thousand existing players have an average rating of
1500, and ten thousand new players of 700-strength enter the pool, the
average rating is now 1100, just as it should be. If these 700-rated
players then improve to 1300 strength, while the 1500s remain at 1500
strength, the proper average rating for the pool is now 1400, not 1500.
(Ten thousand 1300s and ten thousand 1500s average out to 1400.)
The actual average, of course, will be 1100, just as it was when the 700s
entered the pool, because one player's gain is another's loss.
Thus, the apparent 400-point deflation (from an average of 1500 to an
average of 1100) is actually only a 300-point deflation (average strength
1400, average rating 1100). To "correct" a 300-point deflation as though it
were a 400-point deflation will result in a 100-point inflation. Something
like this, I think, is what happened with the post-Fischer fiddle points.
Bill Smythe
> .... you have formulated things very succinctl, Bill. It was misguided
to
>try to "change everyone's perceptions" back to 1973-level in 1985.
______________
Perhaps you and I agree pretty much on what happened. I might even agree
with your second sentence above, if this sort of counter-inflation were
being contemplated today. But it's already done -- the ratings are back
down. So perhaps we should take advantage of the situation, bask in the
glory of our new alignment with history and with FIDE, and work now only to
correct the genuine deflation of the past 3 to 5 years -- most of which, by
the way, has been caused by the failure of the office to implement Ratings
Committee directives.
As I've said before, there will always be a lot of upward political pressure
on the rating system. If this pressure succeeds, there is a danger that our
rating system will become meaningless. Maybe my rating will again go over
2000, but if all those fish who can't play K-and-P endings also have a
rating beginning with the digit 2, why will I care?
Bill Smythe
The whole message is rather speculative. You are proposing a
hypothetical situation to "prove" your point. After reviewing so-
called "Scientific Studies", nothing surprises me anymore. Statistics
and facts can be twitsted by the Scientist to support his
suppositions.
I did a case study of the rating histories of the active players in
Iowa from 1980 to 1999. Generally speaking the majority of players who
had reached a steady state for a period of a few years were over 150
points off their previously established normal mean. I.E. a player who
was approximately 1800 for a period of years had slipped to 1650 over
the course of the last couple of years.
I also reviewed the ratings of players I were familar with. My
unscientific opinion is that was an inflationary period followed by a
deep recession. Additionally, unemployment had risen and new housing
starts were dimished. Consequently, I am suggesting a sharp decrease
in interest rates would be prudent at this juncture....
WildWeasel on ICC
Not quite. The illustration uses the same key issues as impacts the
real situation.
Note that I agree that there is deflation (I've been contending this
for some time.) I am simply pointing out that looking at the average
rating of a pool is not a reliable indicator of deflation. Other
methods must be used.
> After reviewing so-
> called "Scientific Studies", nothing surprises me anymore. Statistics
> and facts can be twitsted by the Scientist to support his
> suppositions.
>
If you feel that I have done so, considering the great simplicity of
this model, please describe how.
What we have is a baseline model. No funny age groups. No bonus or
activity points. No multiple K sizes. This points out in its simplest
form that even when deflation exists, the average rating of a pool is
not a reliable indicator of it.
The next step of course would be to show that adding age groups, bonus
points, etc. doesn't change that. On the face of it, it is hard to see
why any of those factors would suddenly make the average rating of a
pool a reliable indicator of deflation.
> I did a case study of the rating histories of the active players in
> Iowa from 1980 to 1999. Generally speaking the majority of players
who
> had reached a steady state for a period of a few years were over 150
> points off their previously established normal mean. I.E. a player
who
> was approximately 1800 for a period of years had slipped to 1650 over
> the course of the last couple of years.
>
Assuming you have eliminated other social factors, your study IS one of
the approaches used to detect deflation. Note key factors mentioned:
*You did not look at the average of the entire rating pool.
*You looked at specific players, searched for a steady state (an
indcation that proficiency was not changing) and then look at their
average ratings over intervals to see that their ratings were
decreasing even though you had indications that proficiency was
unchanged.
Excellent work.
--
Kevin Bachler
Caveman
"Caveman chess is chess without finesse."
Thanks for your helpful reply.
> But, again - no one on the Ratings Committee (to my knowledge) advocates
> watching the population mean/median/mode and tweaking the system to keep
> it stable.
I think you misunderstood what I meant. This is not it; I apologize if I was ambiguous or unclear. I meant that if you
have specific groups of players whom you think are staying at a stable playing strength, they should also, in the long
run, have a stable rating. If their rating goes down, then that is deflation. Of course, if a bunch of scholastic
players joins a rating pool the average rating goes down and this is not deflation (but when they improve and take
points from those stable players, well that is another story!)
> One involves tracking the distribution of ratings
> for a sub-population which is *believed* to be relatively immune to
> "social effects" (players joining/leaving the pool - such as the huge
> influx of lower rated players). For example, we have looked at players
> in the age range 35-45 with established ratings.
This is exactly what I meant. Sorry if there was any confusion - you said it better. I may not have a PhD in
Statistics, but I do have a BS in Math and graduated at the top of my college class, so I am not totally ignorant in
this area (not that you implied that I was, just wanted those reading my comments to realize that I do have some good
understanding of the problem).
"can indicate" is certainly different than "definitely proves". An "open" pool of rated players which has a lower
rating at some future time may or may not be deflated, so a lower rating "can indicate" the pool might be deflated.
Certainly a higher mean rating of that same pool would less likely indicate deflation!
Dan H
Here's an example of how this problem of deflation caused by improving
players is dealt with in a similar rating system (the german national chess
rating system): when a player exhibits a performance way above his rating
(200+ points above) during a tournament then, when this tournament is rated
, the player is considered to be rated at his performance level, not his
real rating, for rating calculations of all his opponents in the
tournament. Thus, a 700 player showing a 1300 performance is actually
considered to be a 1300 player for all his opponents (though his own rating
change is still calculated from his original 700). This seems to reduce the
influence of the deflation caused by such improving players quite a bit.
Any opinions?
>Regards,
>NM Dan Heisman
Joachim Vaerst
The new rating system largely does that. While not using the P.R., it uses a
first pass recomputed rating. For a 700 playing 1300 strength, this would be
over 1000 (how much depends upon the number of rounds).
Tom Doan
> > I did a case study of the rating histories of the active players in
> > Iowa from 1980 to 1999. Generally speaking the majority of players
> who
> > had reached a steady state for a period of a few years were over 150
> > points off their previously established normal mean. I.E. a player
> who
> > was approximately 1800 for a period of years had slipped to 1650
over
> > the course of the last couple of years.
> >
> Assuming you have eliminated other social factors, your study IS one
of
> the approaches used to detect deflation. Note key factors mentioned:
>
> *You did not look at the average of the entire rating pool.
> *You looked at specific players, searched for a steady state (an
> indcation that proficiency was not changing) and then look at their
> average ratings over intervals to see that their ratings were
> decreasing even though you had indications that proficiency was
> unchanged.
>
> Excellent work.
>
Actually, that wasn't the end of the story. I talked to one of the
local reps and we sent an email to the USCF officials. They responded
they we aware of the problem and had some changes in mind to rectify
the problem.
From what I have read, the changes haven't exactly been what people we
hoping....
Anyway, the USCF have admitted that there has been ratings deflation
over the past few years. I noticed it will browsing ratings of my
friends. I think the "ratings deflation" is undeniable - unless one
can prove we were inflated before.
I could go on, but I would like to say "ratings shouldn't matter".
However, we all know they do. It is more enjoyable to beat an 2200
player than a 1800 player. Maybe I am not just zen enough....
This is complete gobbledygook Current results are a noisy indicator of current
strength. Past results are a noisy indicator of current strength. The two are
combined using standard statistical
methods to obtain a new estimate. The rating system is NOT
based upon the premise that more and more accurate ratings can be obtained as
we get more data, because it is not based upon the belief that a player has an
immutable playing strength. If we believed that, we would eventually end up
with a K factor near zero, and I know how happy that would make you, Eric. For
a player with K=32, the underlying assumption is that we can
never push the standard deviation of our estimate below
approximately 75. For K=24, it's around 60 and for K=16,
around 50. Interestingly, after having made a big deal about the
new K factors, you now seem to be arguing that the K factors
are way too HIGH and that we should be making SMALLER
adjustments. Another indication that you just need to have
something to kvetch about, and really don't care whether what
you say makes any sense.
BTW, what other observable were you planning to employ in
computing updated ratings? You've already speculated that
games played might be useful, but it has no predictive ability
beyond that available in the results.
>It can be very difficult to convince the RC that a player's strength is
>constant even though they may be losing some games....yet most layman see
>this
>as self-evident.
Actually, the RC knows very well that it's possible (and in fact is
the most likely outcome) that a player's strength won't change.
The RC knows how Bayesian learning works. You're the one
who doesn't. The reason the rating changes is that the prior
estimate has limited precision. If the player's result doesn't come
up to the level expected for the prior mean, there is some
downward adjustment.
>Example...I play 10 games vs. GM Korchnoi. I lose all 10. My rating will go
>down...even though my strength is exactly equal to where it started...it may
>even have risen due to my experience in playing those 10 games. Thus, the
>"pain" most players feel when they must play up in the first round of a
>Swiss!
>
>In the RC-world, my "strength" was shown to be lower than my initial
>rating...base on wins/losses/draws...so my new lower rating is more
>"appropriate"...etc. I guess I "got weaker" by playing a match with
>Korchnoi!
Eric, you are a complete twit. This is correct only if you take the
attitude that your prior rating is a perfect estimate. It isn't. (What a
surprise. It's based upon the same type of information, just older). You didn't
get weaker, you just revealed yourself to (likely) be every so slightly not as
good as we thought based upon past results. (Your expected score in ten games
would be around 0.5).
>This is also true for GM Korchnoi....who might lose 20 games straight..yet on
>game 21...he's still Korchnoi!
His passport stays the same, but he's clearly not playing the way
he did. And his rating would reflect that.
>Fact 2: The rating system has independent tags such as A-player, B-player,
>C-player attached to it...and some of us believe that these player types do,
>in
>fact, exist independently of the rating system.
>
>Club players are different from experts and masters, who are different from
>GMs.
Are they different? Yes. Are these well-defined concepts? Of course not. Among
any "class" there is a wide range of ability types.
>The rating system needs to be able to correctly identify and segment these
>independently-existing types....quite apart from pure predictive power....or
>it
>fails in a very important respect (under this view).
>
>Allowing a pool to deflate to the point where GMs carry 1500 ratings, masters
>are 1100, and club players are 800...might not matter in terms of pure
>predictive power. But it matters a great deal in terms of identifying the
>real
>player types in the world....over a stretch of historical time.
>Thus, retaining relative predictive power of the system is quite separate
>from
>maintaining the independent validity of the attached tags (A-player,
>B-player,
>etc.)....which spring forth from other characteristics of players (i.e., how
>they play, quality of moves, how they tend to lose)...other than pure
>wins/losses/draws.
>
>Otherwise, deflation/inflation would be no problem at all for the system!
Yes, it would, because the process of inflation and deflation, because it does
not take place in a smooth and predictable fashion, would result in poorer
predictions of results.
>****
>
>I *do* find it interesting that Kevin makes reference to the independent
>standing of these four 1500 players (he claims they are "established" 1500s
>and
>that their strength appears to be constant...judged how? by reference to
>some
>independent, non-rating-system-based measure?)....and therefore seems to
>accept
>the notion that relative predictive power is not the only point of the rating
>system...
>
>..but then he goes to great lengths in previous posts (months ago)...where he
>focused almost entirely on predictive power as being the ONLY point of
>ratings.
> That predictive power (constant over time) was the "meaning" of ratings.
>
>There is a POV that begs to differ...including Elo, apparently, since Elo
>makes
>reference to the idea that a strong club player should be approx.
>2000...today,
>yesterday, and tomorrow.
>
>Assuming that "strong club player" is a synonym for "player with skills X, Y,
>Z"....and not one for "just the top player in any club"....we can see that
>even
>Elo acknowledged the view that there are independent qualities, apart from
>wins/losses/draws...that determine the anchor points for the rating system.
Given that Elo tried to compute ratings based upon results across many decades,
it's pretty safe to say that it was not HIS view that the definition of GM,
master and expert would be based upon a specific skill set.
<<snip>>
>Eric C. Johnson
Tom Doan
It may not be written the best way, due to the nature of USENET which
encourages fast, off-the-cuff replies, but it ain't "gobbledygook" by any
means.
The RC types often take the line that...if a particular player complains about
being "underrated" ...that the player is simply wrong.
Early complaints about deflation were certainly treated this way...with RC
expressing "doubt" about the matter while club players were griping.
The standard line is that results....wins/losses/draws...allow the rating
system to continually monitor the player and to get better and better estimates
of "playing strength" (presumably deducing that longterm variable from the
directly measured one of temporary success).
The RC types often explicity deny that a player could be strength X but exhibit
temporary success Y.
They often require the player to play through additional games, saying words to
the effect that if the player is really strength X then he should be able to
play through any bad stretch and regain a particular rating by showing
prolonged success Y.
I point to the recent debate over a database-wide rating correction to handle
deflation. Several posters here thought that was "unfair" and that rating
corrections had to be "earned" through "future wins/losses/draws"...something
that is odd if one is merely looking to bring rating estimates in line with all
known facts.
Nobody would say that a machine has to earn a recalibration!
>The rating system is NOT
>based upon the premise that more and more accurate ratings can be obtained as
>we get more data
Well, actually it is...in part precisely because it chooses to take in the
highly fluctuating temporary success of day-to-day event-by-event play.
Win some/rating goes up, lose some/rating goes down is precisely how it works.
More and more data = more and more changes, fluctuating around a set of points.
And when particular players complain about "having a bad day" or other
extraneous non-strength factors...the standard RC response is "oh, well, you
were overrated...and all these other factors do matter"...etc.
>because it is not based upon the belief that a player has an
>immutable playing strength.
Over sufficient time or games, it darn near does resemble a non-changing
quantity. Experts stay experts. GMs stay GMs. 1200s stay 1200s.
Only if you include the day-to-day "noise" in your estimate, and issue frequent
rating changes...do you see the wide ups and downs.
A static rating system issuing changes every 6 months (ala FIDE) is preferred
from a strength analysis perspective...over one that issues updates every two
months (USCF) or monthly or weekly...precisely because there is so much
day-to-day and event-by-event noise.
More frequent updates *do* tend to make people confuse the noise of temporary
success changes...for the stability of real strength changes. It is a bit like
confusing quarterly grades with daily quizzes!
(This is why online rating systems that change with every game are so
misleading, and why they tend to resemble token economies that lead to greater
gambling-like obsessive behaviors.)
>you now seem to be arguing that the K factors
>are way too HIGH and that we should be making SMALLER
>adjustments.
No, I have always made a *separate* point about how more and more frequent
rating "updates" really tell us nothing....that all it does is show more and
more noise and less and less about the underlying strength variable.
Chess strength...except for rapidly improving AMATEUR PLAYERS under a certain
low rating....changes rather slowly.
> Another indication that you just need to have
>something to kvetch about, and really don't care whether what
>you say makes any sense.
Hardly...but I do observe that you are the kind of fellow who doesn't much care
for debate.
If someone disagrees with you, you call them names.
IMHO you have a big "volunteer" chip on your shoulder.
>
>BTW, what other observable were you planning to employ in
>computing updated ratings?
I did not state any examples...but noted that ELO made reference that "strong
club player = 2000" is something that should be maintained.
I'd assert that "quality of play" is something that can be assessed...and that
we all have notions as to the "quality level" of games played by 1200s, 1600s,
1900s, 2100s, and 2300s.
If a closed pool of 1200s is producing games that look like they were played by
2100s, we have evidence of deflation or lack of synchronization in that pool.
Very observable...but not based on results, because the 1200s might all be
scoring equally against each other and thus maintaining their relative
predictive positions.
>You've already speculated that
>games played might be useful, but it has no predictive ability
>beyond that available in the results.
If the rating system slipped so that GMs were rated 1500, IMs rated 1300,
masters rated 1100, A-players rated 900, and club amateurs rated 700...the
predictive power would be unchanged. But the associated tags would have lost
all meaning....because it is not only predictive power based on results but
also other things such as quality of play and overall knowledge that help
ground our system.
If I told you it was 85 degrees outside...you have to know that I am talking
Fahrenheit and not Kelvin before you go without a coat.
>
>Actually, the RC knows very well that it's possible (and in fact is
>the most likely outcome) that a player's strength won't change.
Yet their rating will fluctuate...
...and moreover, there is a greater and greater emphasis in USCF-land to issue
more frequent rating updates. Why?
Not for any strength-based reasons....for all the frequent updates do is muddy
the boundary between temporary success and longterm strength.
>If the player's result doesn't come
>up to the level expected for the prior mean, there is some
>downward adjustment.
>
Yes, there is.
But the standard response from RC types is to attribute this downward
adjustment to some underlying strength of the player (i.e., your new rating is
more reflective of your play, you weren't as strong as you thought you were
since you lost 0-10, etc.) rather than to some temporary succcess.
They almost *never* say "well, your new, slightly lower estimate is -
functionally - equal to the older, slightly higher one.
Instead, they like to needle the player by pointing (erroneously, as you so
nicely point out!) to "strength estimates...rather than success ones.
My beef is with the snobbery and the snide remarks, not with the model.
>
>>Example...I play 10 games vs. GM Korchnoi. I lose all 10. My rating will
>go
>>down...even though my strength is exactly equal to where it started...it may
>>even have risen due to my experience in playing those 10
>games. Thus, the
>>"pain" most players feel when they must play up in the first round of a
>>Swiss!
>>In the RC-world, my "strength" was shown to be lower than my initial
>>rating...base on wins/losses/draws...so my new lower rating is more
>>"appropriate"...etc. I guess I "got weaker" by playing a match with
>>Korchnoi!
>
>Eric, you are a complete twit.
Mr. Doan, your debating style is at a very high level today!
> This is correct only if you take the
>attitude that your prior rating is a perfect estimate. It isn't. (What a
>surprise. It's based upon the same type of information, just older). You
>didn't
>get weaker, you just revealed yourself to (likely) be every so slightly not
>as
>good as we thought based upon past results. (Your expected score in ten games
>would be around 0.5).
IMHO you have said just exactly the same thing I did.
That my demonstrated score was such that a downward adjustment was
needed...because I did not exhibit the required score to keep the older
estimate.
In a word, my success measure needs adjusting, even if my strength measure does
not (indeed, just by playing the match, one might suppose my strength went up
ever so slightly...because most chess learning is monotonic increasing).
That by playing Korchnoi 10 games and going 0-10, the rating system gained new
info about me...i.e., that my rating needed to be adjusted downward. On what
basis?
My score, and therefore my temporary success. But everyone and their uncle
views ratings as indirect measures of underlying strength...and by playing the
match...my indirect measure of my strength went down. Connect the dots.
In a ratings-world, I would have been better off not playing the match.
Idleness pays off.
That's *precislely* why nobody likes to play way up in the first round of a
Swiss - it is a rating-losing proposition!
>
>Given that Elo tried to compute ratings based upon results across many
>decades,
>it's pretty safe to say that it was not HIS view that the definition of GM,
>master and expert would be based upon a specific skill set.
Well, we disagree.
And if you are saying..that "expert" only means "top 8-9 percent of the rating
pool"...then we also disagree.
IMHO it is also linked to certain, necessary skills.
And anytime that link is sufficiently broken...then the average player in the
street will think the rating system is inflated/deflated....regardless of
whether the RC thinks so.
Eric C. Johnson
>
>No, I have always made a *separate* point about how more and more frequent
>rating "updates" really tell us nothing....that all it does is show more and
>more noise and less and less about the underlying strength variable.
I should also clarify in advance:
I am *delighted* to have the ratings calculations crunch through a large number
of games...at high K....middle K...any K you like.
But the public, obervable ratings changes....issued by the ratings agency
(USCF...FIDE...etc.)...ought to be occurring relatively infrequently, so that
there *are* large numbers of games to crunch for the new, public, official
estimate!
Otherwise, if the public update occurs over a very small number of games, the
new "update" is just a noise change...not a real one.
A crude analogy.
I can tell if Johnny is an A student, B student or C student on a semester by
semester basis...based on lots of homework and quizzes.
I cannot tell if Johnny is A- today, B+ tomorrow, B next week, back to B+ the
week after...based on day-to-day results. And I would be chasing windmills if I
tried to grade my class that way...every day.
Eric C. Johnson
Eric, you are allowed to actually think through your replies before making
them. Don't blame USENET for fuzzy thinking.
>The RC types often take the line that...if a particular player complains
>about
>being "underrated" ...that the player is simply wrong.
Because they usually are. People have a tendency to underestimate the extent to
which ratings fluctuate.
>Early complaints about deflation were certainly treated this way...with RC
>expressing "doubt" about the matter while club players were griping.
>
>The standard line is that results....wins/losses/draws...allow the rating
>system to continually monitor the player and to get better and better
>estimates
>of "playing strength" (presumably deducing that longterm variable from the
>directly measured one of temporary success).
>
>The RC types often explicity deny that a player could be strength X but
>exhibit
>temporary success Y.
Show me where anyone says anything remotely resembling that. The whole point of
the rating system is to pull signal out of the noise from the tournament data.
>They often require the player to play through additional games, saying words
>to
>the effect that if the player is really strength X then he should be able to
>play through any bad stretch and regain a particular rating by showing
>prolonged success Y.
And the alternative is....?
>I point to the recent debate over a database-wide rating correction to handle
>deflation. Several posters here thought that was "unfair" and that rating
>corrections had to be "earned" through "future wins/losses/draws"...something
>that is odd if one is merely looking to bring rating estimates in line with
>all
>known facts.
Who said it was unfair? A number of people of various persuasions on other
issues thought that a prospective correction in some form would at least have
the side effect of encouraging activity, while a retrospective one wouldn't. I
don't remember anyone couching this as a fairness issue.
>Nobody would say that a machine has to earn a recalibration!
>
>>The rating system is NOT
>>based upon the premise that more and more accurate ratings can be obtained
>as
>>we get more data
>
>Well, actually it is...in part precisely because it chooses to take in the
>highly fluctuating temporary success of day-to-day event-by-event play.
Actually, it isn't. I don't know if you read the explanation before snipping
it. Try doing that next time.
>Win some/rating goes up, lose some/rating goes down is precisely how it
>works.
>More and more data = more and more changes, fluctuating around a set of
>points.
Unless, of course, a player is improving or worsening, in which case, the
fluctuations are around a changing function.
>And when particular players complain about "having a bad day" or other
>extraneous non-strength factors...the standard RC response is "oh, well, you
>were overrated...and all these other factors do matter"...etc.
Having a bad day, having a good day, playing an opponent who is having a off
day, being paired with an opponent who has a strength or weakness which matches
with yours, are all part of the randomness inherent in the results.
>>because it is not based upon the belief that a player has an
>>immutable playing strength.
>
>Over sufficient time or games, it darn near does resemble a non-changing
>quantity. Experts stay experts. GMs stay GMs. 1200s stay 1200s.
For some players, yes. For others, no. I would particularly quibble with your
1200s stay 1200s. A 1200 who would like not to be a 1200 any more can probably
do that with a modest amount of effort.
>Only if you include the day-to-day "noise" in your estimate, and issue
>frequent
>rating changes...do you see the wide ups and downs.
>A static rating system issuing changes every 6 months (ala FIDE) is preferred
>from a strength analysis perspective...over one that issues updates every two
>months (USCF) or monthly or weekly...precisely because there is so much
>day-to-day and event-by-event noise.
Not necessarily, particularly in the population of USCF players. Low rated
players can easily improve hundreds of points in a few months. A highly time
aggregated system would be a mess when applied to those players.
>More frequent updates *do* tend to make people confuse the noise of temporary
>success changes...for the stability of real strength changes. It is a bit
>like
>confusing quarterly grades with daily quizzes!
No, it's not. Confusing quarterly grades with daily quizzes would be like
confusing the published rating with a set of performance ratings.
>(This is why online rating systems that change with every game are so
>misleading, and why they tend to resemble token economies that lead to
>greater
>gambling-like obsessive behaviors.)
>
>>you now seem to be arguing that the K factors
>>are way too HIGH and that we should be making SMALLER
>>adjustments.
>
>No, I have always made a *separate* point about how more and more frequent
>rating "updates" really tell us nothing....that all it does is show more and
>more noise and less and less about the underlying strength variable.
>Chess strength...except for rapidly improving AMATEUR PLAYERS under a certain
>low rating....changes rather slowly.
>
>> Another indication that you just need to have
>>something to kvetch about, and really don't care whether what
>>you say makes any sense.
>
>Hardly...but I do observe that you are the kind of fellow who doesn't much
>care
>for debate.
OK, let's get this straight. You have indicated that you don't like the new
system, and apparently prefer the old to the new. (since you don't consider
implementing the new system to be a positive step). I am assuming that this is
because you dislike the lower K factors which will make it harder for players
to move up into that next class. So the alternative you propose is to aggegate
tournaments over time, thus lowering the K factors even further.
Actually, I don't mind a good debate. Maybe if you take some statistics and
critical thinking courses, we could actually have one.
>If someone disagrees with you, you call them names.
Not really. If someone disagrees with themselves, I call them names.
>IMHO you have a big "volunteer" chip on your shoulder.
A chip on my shoulder? Hardly. I find ripping you to shreds mildly
entertaining.
>>
>>BTW, what other observable were you planning to employ in
>>computing updated ratings?
>
>I did not state any examples...but noted that ELO made reference that "strong
>club player = 2000" is something that should be maintained.
>I'd assert that "quality of play" is something that can be assessed...and
>that
>we all have notions as to the "quality level" of games played by 1200s,
>1600s,
>1900s, 2100s, and 2300s.
Who would make these decisions, Eric? Oh, BTW, where would you put a player who
has never seen the idea of counting moves in a K+P endgame?
>If a closed pool of 1200s is producing games that look like they were played
>by
>2100s, we have evidence of deflation or lack of synchronization in that pool.
>Very observable...but not based on results, because the 1200s might all be
>scoring equally against each other and thus maintaining their relative
>predictive positions.
>
>>You've already speculated that
>>games played might be useful, but it has no predictive ability
>>beyond that available in the results.
>
>If the rating system slipped so that GMs were rated 1500, IMs rated 1300,
>masters rated 1100, A-players rated 900, and club amateurs rated 700...the
>predictive power would be unchanged. But the associated tags would have lost
>all meaning....because it is not only predictive power based on results but
>also other things such as quality of play and overall knowledge that help
>ground our system.
You can stop repeating that, Eric. I think we all, like, kind of understand
that.
>If I told you it was 85 degrees outside...you have to know that I am talking
>Fahrenheit and not Kelvin before you go without a coat.
Thank you for telling me that, Eric.
>>
>>Actually, the RC knows very well that it's possible (and in fact is
>>the most likely outcome) that a player's strength won't change.
>
>Yet their rating will fluctuate...
Yes, because the old rating is itself an imperfect measure of playing strength,
as is the new. It's just the best that we have at a given time.
>...and moreover, there is a greater and greater emphasis in USCF-land to
>issue
>more frequent rating updates. Why?
>
>Not for any strength-based reasons....for all the frequent updates do is
>muddy
>the boundary between temporary success and longterm strength.
>
>>If the player's result doesn't come
>>up to the level expected for the prior mean, there is some
>>downward adjustment.
>>
>
>Yes, there is.
>
>But the standard response from RC types is to attribute this downward
>adjustment to some underlying strength of the player (i.e., your new rating
>is
>more reflective of your play, you weren't as strong as you thought you were
>since you lost 0-10, etc.) rather than to some temporary succcess.
>
>They almost *never* say "well, your new, slightly lower estimate is -
>functionally - equal to the older, slightly higher one.
>
>Instead, they like to needle the player by pointing (erroneously, as you so
>nicely point out!) to "strength estimates...rather than success ones.
>
>My beef is with the snobbery and the snide remarks, not with the model.
Let's see. A computer program takes a set of results and computes that a rating
is now 2107 when last it was 2111. Publishing this is somehow considered
needling and snobbery. Oh, yes. And what happens if the 2107 goes to 2111?
We are not saying the same thing. You say that you teach statistics. How the
hell can you confuse an estimator (rating) with an underlying parameter
(strength)?
>In a word, my success measure needs adjusting, even if my strength measure
>does
>not (indeed, just by playing the match, one might suppose my strength went up
>ever so slightly...because most chess learning is monotonic increasing).
>
>That by playing Korchnoi 10 games and going 0-10, the rating system gained
>new
>info about me...i.e., that my rating needed to be adjusted downward. On what
>basis?
>
>My score, and therefore my temporary success. But everyone and their uncle
>views ratings as indirect measures of underlying strength...and by playing
>the
>match...my indirect measure of my strength went down. Connect the dots.
>
>In a ratings-world, I would have been better off not playing the match.
>Idleness pays off.
Only if you measure your manhood by your MR_CUR_RAT.
>That's *precislely* why nobody likes to play way up in the first round of a
>Swiss - it is a rating-losing proposition!
NOBODY???? NOBODY???? You must have a REALLY weird club if no one likes a
challenge. Jeez.
>>
>>Given that Elo tried to compute ratings based upon results across many
>>decades,
>>it's pretty safe to say that it was not HIS view that the definition of GM,
>>master and expert would be based upon a specific skill set.
>
>Well, we disagree.
>And if you are saying..that "expert" only means "top 8-9 percent of the
>rating
>pool"...then we also disagree.
>
>IMHO it is also linked to certain, necessary skills.
OK, I'll take Elo on my side any day. That is a debatable point, but the
relationship with virtually all other types of sporting endeavors would argue
in favor of standards changing as knowledge and technology change.
>And anytime that link is sufficiently broken...then the average player in the
>street will think the rating system is inflated/deflated....regardless of
>whether the RC thinks so.
And, of course, the conventional wisdom was that that link was broken in 1980
with fiddle points.
Notice that I specifically exempted low rated players..yet you give this
example anyway.
>
>>Chess strength...except for rapidly improving AMATEUR PLAYERS under a
>certain
>>low rating....changes rather slowly.
There is what I said (above)
>
>OK, let's get this straight. You have indicated that you don't like the new
>system,
For TWO reasons:
1. The announcement was bungled.
2. I don't like the fact that deflation of the last 10 yrs is being frozen into
place by lower K factors. I don't mind lower K as a means of increasing
stability per se...but I *do* mind it when the 10+ yrs of deflation are being
frozen into place. And the consumer base for ratings data seems not to like it
either, and they don't like the answer "just play through it"...etc.
> I am assuming that this is
>because you dislike the lower K factors which will make it harder for players
>to move up into that next class.
See above. Your assumption is wrong.
I dislike locking in the deflation before making K smaller. I don't mind
smaller K.
> So the alternative you propose is to aggegate
>tournaments over time, thus lowering the K factors even further.
Since you didn't understand my argument, no wonder you thought this was a
conflict.
>
>Actually, I don't mind a good debate. Maybe if you take some statistics and
>critical thinking courses, we could actually have one.
>
Didn't keep you from making the typical ratings committee snide remark. I have
taught statistics, and I hold a MS in social science.
>
>Not really. If someone disagrees with themselves, I call them names.
And if someone makes a comment like yours...you call them what?
>
>A chip on my shoulder? Hardly. I find ripping you to shreds mildly
>entertaining.
Yet you failed, because you seem incapable of understanding the following:
A. The implementation/launch of the new rating system was bungled. No
brochures. No mailings. No splashy announcement. No firm implementation date.
Just Tom Doan trekking up to New Windsor to work his magic...implementing the
parts of the formula he likes..and not doing the parts he doesn't like (i.e.,
activity points). The implementation is a mess.
B. Locking in 10 yrs of deflation is not good PR with the masses. They have
experienced and played through deflation for years. Yet now, the new system
will lock them in place with lower K.
The RC was slow to warm to the idea of a point correction. Debate on this forum
showed that. RC didn't think a point correction was needed....to paraphrase
Ken Sloan, it would have been a sop to the EB to do one.
C. I don't object to lower K....only to lower K without an initial correction.
Thus, your assumption about how my views on the "new system implementation" and
the idea that official ratings should be published less often, with
number-crunching over a large game set....are somehow in conflict...has been
shown to be wrong.
>
>Let's see. A computer program takes a set of results and computes that a
>rating
>is now 2107 when last it was 2111. Publishing this is somehow considered
>needling and snobbery. Oh, yes. And what happens if the 2107 goes to 2111?
Your reply is a good example of the phenomenon I was aiming at.
>
>We are not saying the same thing. You say that you teach statistics. How the
>hell can you confuse an estimator (rating) with an underlying parameter
>(strength)?
>
I do not..and I take pains not to do so...far more so than most posters here
who glibly state that ratings measure strength.
>>
>>In a ratings-world, I would have been better off not playing the match.
>>Idleness pays off.
>
>Only if you measure your manhood by your MR_CUR_RAT.
Nice reply. It shows you are out of touch with the consumers of ratings
data...because ratings are taken VERY SERIOUSLY.
>
>>That's *precislely* why nobody likes to play way up in the first round of a
>>Swiss - it is a rating-losing proposition!
>
>NOBODY???? NOBODY???? You must have a REALLY weird club if no one likes a
>challenge. Jeez.
>
Most folks don't like playing a schedule where they are *bound* to have a
ratings loss.
>
>And, of course, the conventional wisdom was that that link was broken in 1980
>with fiddle points.
That view is like trying to place the USA back on the gold standard to drive
prices down to 1955 levels.
It is 2001. We have a current crop of rating consumers who view ratings as
deflated.
Telling them that they are really the same as they were in 1975...and that the
last 25 yrs have all been a mistake...is not the way to increase participation
or player/consumer satisfaction. Duh!
Eric C. Johnson
This is impossible under the present system, which is the reason why
tournaments now must be rated in the order the reports are received by
the office, not in the order in which they were actually played.
Once the new rating system is in place we can go back in time and see
whether and when there was rating inflation or deflation and
adjustments can be made.
Puting in fiddle points now or making a one-time adjustment would
simply destroy all the work that has been done on the rating system.
Sam Sloan
OK, since you're so damned imprecise, what exactly are you proposing? That we
aggregate rate tournaments for "higher" rated players, but rate tournament by
tournament for "lower" rated players. How, pray tell, would you pull that one
off, given that lower and higher rated players can play in the same tournament?
Or are you simply proposing that we rate tournament by tournament and not
publish the ratings on the same schedule?
>>
>>>Chess strength...except for rapidly improving AMATEUR PLAYERS under a
>>certain
>>>low rating....changes rather slowly.
>
>There is what I said (above)
>
>>
>>OK, let's get this straight. You have indicated that you don't like the new
>>system,
>
>For TWO reasons:
>
>1. The announcement was bungled.
>
>2. I don't like the fact that deflation of the last 10 yrs is being frozen
>into
>place by lower K factors. I don't mind lower K as a means of increasing
>stability per se...but I *do* mind it when the 10+ yrs of deflation are being
>frozen into place. And the consumer base for ratings data seems not to like
>it
>either, and they don't like the answer "just play through it"...etc.
And in answer to my question, given the choice between the new system and the
old, given that there will be no retrospective adjustment, do you prefer the
old? Given the options that actually are on the table, which do you want?
>> I am assuming that this is
>>because you dislike the lower K factors which will make it harder for
>players
>>to move up into that next class.
>
>See above. Your assumption is wrong.
Actually, anyone who read the nonsense from last fall when you kept talking
about rescaling would know that my assumption is actually correct.
>I dislike locking in the deflation before making K smaller. I don't mind
>smaller K.
>> So the alternative you propose is to aggegate
>>tournaments over time, thus lowering the K factors even further.
>
>Since you didn't understand my argument, no wonder you thought this was a
>conflict.
Actually, your complaint (among many) was that changing the K factors locked in
relative positions for "tangible benefits." (Remember that overworked phrase).
A +C adjustment wouldn't change those relationships. Switching to time
aggregated rating (if that's what you're proposing) would lower the adjustment
rates and thus also "lock in" historical positions.
>>
>>Actually, I don't mind a good debate. Maybe if you take some statistics and
>>critical thinking courses, we could actually have one.
>>
>
>Didn't keep you from making the typical ratings committee snide remark. I
>have
>taught statistics, and I hold a MS in social science.
Goody for you. I hope you assigned a good textbook so the students could get a
second opinion, because you betray a complete lack of understanding of how
statistical analysis actually works.
>>
>>Not really. If someone disagrees with themselves, I call them names.
>
>And if someone makes a comment like yours...you call them what?
>
>>
>>A chip on my shoulder? Hardly. I find ripping you to shreds mildly
>>entertaining.
>
>Yet you failed, because you seem incapable of understanding the following:
>A. The implementation/launch of the new rating system was bungled. No
>brochures. No mailings. No splashy announcement. No firm implementation
>date.
>Just Tom Doan trekking up to New Windsor to work his magic...implementing the
>parts of the formula he likes..and not doing the parts he doesn't like (i.e.,
>activity points). The implementation is a mess.
I seem incapable of understanding this??? This has been bungled for the last
FOUR YEARS, by several EB's, by several ED's, and God knows who else. This
wasn't just put on the back burner - it wasn't even in the kitchen. You don't
think I would rather have been done with this during the summer? I'm not the
one making the decisions about where this stands on the priority list. The RC
isn't making those decisions.
>B. Locking in 10 yrs of deflation is not good PR with the masses. They have
>experienced and played through deflation for years. Yet now, the new system
>will lock them in place with lower K.
And the old system not only locks in place the deflation, but lets it continue.
>The RC was slow to warm to the idea of a point correction. Debate on this
>forum
>showed that. RC didn't think a point correction was needed....to paraphrase
>Ken Sloan, it would have been a sop to the EB to do one.
That depends upon what a "point correction" is. Fixed adjustment? Retrospective
calculation? Activity points? Sweetened bonus? Those are all point corrections
in some form. The RC is, to my knowledge, of a single mind only with regard to
its opposition to activity points. There is widespread support for a sweetened
bonus. The others have both supporters and opponents.
>C. I don't object to lower K....only to lower K without an initial
>correction.
>
>Thus, your assumption about how my views on the "new system implementation"
>and
>the idea that official ratings should be published less often, with
>number-crunching over a large game set....are somehow in conflict...has been
>shown to be wrong.
Only if your views have changed from the "tangible benefits" days.
So how, precisely, do you propose to number crunch over a large game set? Give
me an algorithm, please. Whether this would even do what you think it would
depends upon the details.
>>
>>Let's see. A computer program takes a set of results and computes that a
>>rating
>>is now 2107 when last it was 2111. Publishing this is somehow considered
>>needling and snobbery. Oh, yes. And what happens if the 2107 goes to 2111?
>
>Your reply is a good example of the phenomenon I was aiming at.
What, that 2111 to 2107 is needling and snobbery?
>>
>>We are not saying the same thing. You say that you teach statistics. How the
>>hell can you confuse an estimator (rating) with an underlying parameter
>>(strength)?
>>
>
>I do not..and I take pains not to do so...far more so than most posters here
>who glibly state that ratings measure strength.
>
>>>
>>>In a ratings-world, I would have been better off not playing the match.
>>>Idleness pays off.
>>
>>Only if you measure your manhood by your MR_CUR_RAT.
>
>Nice reply. It shows you are out of touch with the consumers of ratings
>data...because ratings are taken VERY SERIOUSLY.
Actually, I do take ratings rather seriously myself, Eric. Which is why I spend
hundreds of hours a year on ratings matters.
>>
>>>That's *precislely* why nobody likes to play way up in the first round of a
>>>Swiss - it is a rating-losing proposition!
>
>>
>>NOBODY???? NOBODY???? You must have a REALLY weird club if no one likes a
>>challenge. Jeez.
>>
>
>Most folks don't like playing a schedule where they are *bound* to have a
>ratings loss.
Except, of course, that the expected rating change from a game is zero.
>>
>>And, of course, the conventional wisdom was that that link was broken in
>1980
>>with fiddle points.
>
>That view is like trying to place the USA back on the gold standard to drive
>prices down to 1955 levels.
Great analogy. We have millions of existing contracts payable in rating points.
Stuff
Tom, I want to thank you for taking the time to make some rating system comments.
I know that I don't always do the best job, and sometimes I get too frustrated
(people want to believe what they think is true whether it is or not) but
I am glad that you are taking the time to make the responses.
It is difficult to get across the concept that the relationship between rating
and performance is "stochastic," not "deterministic," that it is "fuzzy",
not exact. I thought the stock analogy might be really useful, but people
don't want to buy it (no pun intended.) They see a rating, much as they saw
the vote in Florida, as something that was determinable. (Although my guess
would be that the difference in the vote count is within the error margin
of counting by either man or machine. Sort of a "Planck length" of voting!)
Again Tom, Thanks.
Again, thanks.
-----= Posted via Newsfeeds.Com, Uncensored Usenet News =-----
http://www.newsfeeds.com - The #1 Newsgroup Service in the World!
-----== Over 80,000 Newsgroups - 16 Different Servers! =-----
At the moment, I'd prefer the old system.
Implementation/launch of the new one has been bungled so badly...that the gain
is outweighed by the pain.
Eric C. Johnson
A closely monitored pool with frequent point corrections by the administrator
would not lock in the deflation....adjustments can and should be made to each
segment as needed...as often as needed.
>
>And the old system not only locks in place the deflation, but lets it
>continue.
>
Elo advocates an *active* ratings administrator who constantly tests for
deflation and takes active (frequent) measures to correct for it.
USCF opts to allow the system to run on autopilot for years and years...because
any point corrections for sub-populations are seen as politically dangerous.
There is a notion running around that any rating changes need to be "earned"
through future play...this is nonsense.
If a segment of the pool is inflated, that segment should get a point
adjustment immediately. If a segment is deflated, that segment should get a
point adjustment immediately.
10 yrs of deflation don't happen by chance...they happen because we have no
administrator taking annual steps to keep things running smoothly.
We let things go...because a hands-off approach fits in with this idea that
"all rating changes must be earned" nonsense.
Your original question is incomplete...instead of a choice between:
a. new system
b. old system
I offer a third alternative
c. old system with frequent corrections for deflation/inflation via an active
ratings administrator
There may be other approaches as well...but IMHO c. is better than a very
complex new algorithm that includes age of player and even more sliding K
factors...
>
>Only if your views have changed from the "tangible benefits" days.
>
They have not.
Eric C. Johnson
>USCF opts to allow the system to run on autopilot for years and years...because
>any point corrections for sub-populations are seen as politically dangerous.
No. I think that the RC IS your active administrator, they make recommendations,
and USCF is too inept to implement them.
>There is a notion running around that any rating changes need to be "earned"
>through future play...this is nonsense.
I haven't heard that notion anywhere. Where have you heard this?
>If a segment of the pool is inflated, that segment should get a point
>adjustment immediately. If a segment is deflated, that segment should get
a
>point adjustment immediately.
>
Glad you agree. The new system does that with greater surgical precision
than the administrator can.
>10 yrs of deflation don't happen by chance...they happen because we have
no
>administrator taking annual steps to keep things running smoothly.
>
Or because the administrator's actions are not implemented.
>
>How do you determine the amounts? Read what Elo said.
He advocated monitoring a group of players known (or assumed) to be stable.
USCF could do this, and make corrections on an annual basis. They do not make
annual corrections. They allow freefall.
Eric C. Johnson
MY reading of Elo is that he advocated an active administrator who used
MULTIPLE TOOLS to combat observed inflation/deflation.
Not just building in one or two tools into a complex formula and letting it
work for a decade untouched, but constantly monitoring, tweaking, and
adjusting. Point corrections to segments of the population would be an
appropriate tool for an active admin.
I will agree that his book is sketchy on these details.
>
>No. I think that the RC IS your active administrator, they make
>recommendations,
>and USCF is too inept to implement them.
No..RC is *not* the active admin...in part because they do not have any
implementation authority at all.
That's the beef, right...that for five years the new formula languished.
RC is advisory. They can be an important resource for the admin...who might
task RC with a study or other project. But the active admin, IMHO, would be a
single person (like Elo)...charged with keeping ratings on an even keel.
>>There is a notion running around that any rating changes need to be "earned"
>>through future play...this is nonsense.
>
>I haven't heard that notion anywhere. Where have you heard this?
Straight from your keyboard, Kevin.
>>If a segment of the pool is inflated, that segment should get a point
>>adjustment immediately. If a segment is deflated, that segment should get
>a
>>point adjustment immediately.
>
>Glad you agree. The new system does that with greater surgical precision
>than the administrator can.
>
Oh, we disagree entirely. And let's not let the new system run blindly for the
next 10 yrs, either...OK?
Look...instead of crafting ever-more-complicated algorithms...why not keepp a
SIMPLE algorith and then make annual adjustments up or down to the various
population segments, as analysis indicates?
Simple formulas are easy to explain to the consumers of rating services.
Complex ones are not.
Eric C. Johnson
Now read what he said about how corrections are to be made. See Chapter
3, I think 3.5 - 3.8 (don't have the book here...)
> They do not make
>annual corrections. They allow freefall.
>
>Eric C. Johnson
-----= Posted via Newsfeeds.Com, Uncensored Usenet News =-----
What mechanism are you referring to here? Certainly neither the activity points
nor the change in K factors seem to be operating with "surgical precision".
In my opinion, the least intrusive manipulation of the system is upward and
downward adjustments to the rating floor. This is the key current mechanism for
injecting rating points, and we saw a big change when we went from a 100 point
below top rating floor to a lower floor. Rating floors are best thought of not as
a permanent lower bound on your rating, but as a way of injecting points into the
system. It should be a standard annual decision of the RC to decide how/whether
to adjust the rating floor. The rating floor can be adjusted in several ways,
either by changing the maximum number of points you are allowed to lose from your
peak rating upwards or downwards (depending on whether we are seeing inflation or
deflation), and by resetting your peak value to your current value. Small annual
adjustments would likely draw less criticism than other modifications of the
system, in my opinion, and would give the RC a tool to make these adjustments
much more reasonably than is done blindly by the new formulas.
Jerry Spinrad
>Not just building in one or two tools into a complex formula and letting
it
>work for a decade untouched, but constantly monitoring, tweaking, and
>adjusting.
Yes, but all the adjusting is being adjusting of the formula.
>Point corrections to segments of the population would be an
>appropriate tool for an active admin.
I don't recall seeing Elo make that argument.
>
>I will agree that his book is sketchy on these details.
>
I think he is straightforward. He lists the tools, I don't recall adding
points to specific players as determined by the administrator as one of them.
Elo strikes me as very process oriented, and he seemed to want to the corrections
to flow naturally from formula adjustments, not on an ad hoc basis.
>>
>>No. I think that the RC IS your active administrator, they make
>>recommendations,
>>and USCF is too inept to implement them.
>
>No..RC is *not* the active admin...in part because they do not have any
>implementation authority at all.
As I said, I see the RC as the administrator, and a superior one, because
having a committee provides checks and balances.
>>>There is a notion running around that any rating changes need to be "earned"
>>>through future play...this is nonsense.
>>
>>I haven't heard that notion anywhere. Where have you heard this?
>
>Straight from your keyboard, Kevin.
Then something is misunderstood somewhere.
>Oh, we disagree entirely. And let's not let the new system run blindly
for the
>next 10 yrs, either...OK?
There's that straw man guy again.
>Look...instead of crafting ever-more-complicated algorithms...why not keepp
a
>SIMPLE algorith and then make annual adjustments up or down to the various
>population segments, as analysis indicates?
1. Because then it isn't ad hoc.
2. Because then points flow more continuously, making better self adjustments.
3. Because the best methods for determining how many points to adjust would
require these formulae anyway?
4. Because it would avoid the perception that the Rating Administrator was
treating anyone unfairly.
The estimation formula IS simple. It just hasn't been explained that way.
1. Estimate your rating change as you did before.
2. Estimate your new K as follows (describe)
3. Adjust your point gain by newK/oldK.
4. Describe bonus points.
Done.
I'd agree that making "standard" annual decisions of this type, over time,
would de-politicize such changes.
I'd prefer to see a very basic algorithm...with annual point changes...than be
forced to deal with a very complex one.
Maybe we'd see something like:
Yr 1 scholastics +23 adults +8
Yr 2 scholastics +17 adults +9
Yr 3 scholastics -8 adults -1
Yr 4 scholastics +13 adults +11
etc....small annual adjustments.
Instead, we either 1) let the system float on autopilot for years and years
until problems build up, or 2) attempt to solve rating issues by using
ever-greater complexity in our formulas...forgetting that "explainability" is a
very important factor when dealing with the consumers of ratings data.
Annual adjustments OF ANY TYPE are better than autopilot or complex formulas.
Eric C. Johnson
1. Appropriate Treatment of Unrated Players: In rating a tournament, compute
the performance ratings of the new playes ratings when calculating the
ratings of other players. (New and old do this.)
2. Modified Processing of Provisionally Rated Players: After the new
players are rated, compute the provisional ratings and use these when
calculating the ratings for other players. (New and old do this.)
3. Adjustment of K: Use K as a development coefficient, setting a high value
early in a player's career, and reducing it as he stablizes with time or
play. (Old system did this a little, reducing K for experts and masters.
New system does this robustly.)
4. Corrective additions for exceptional performances (bonus points). Use a
higher K or award bonus points for a player where an outstanding performance
is detected.
5. Feedback processes: Use the post event rating of the exceptional
performer in calculating the ratings of his opponents.
The K factor and bonus point approach are certainly more fine-tuned than an
ad hoc approach, especially given that the processes work BY PLAYER (not by
pool segment).
It is puzzling to me why ad hoc approaches, such as that described below
(i.e., the rules change annually) are seen as simpler than a process.
"Jeremy Spinrad" <sp...@vuse.vanderbilt.edu> wrote in message
news:93frre$ic0$1...@news.vanderbilt.edu...
Yet such ad hoc methods *were* used by FIDE (under Elo's guidance, if I
understand correctly) to adjust the ratings of the pool of women players.
I see no difference between the five methods articulated explicitly in Elo's
(very) short section in his book...and the general principle that runs
throughout his book...that whenever an active admin sees inflation/deflation he
should take forceful and immediate steps against.
Elo's book...while very important...is not gospel. The principles from
it...such as an active admin...are of greater importance...than the specific
tools he mentions or fails to mention.
Clearly...there are more than just FIVE tools available to an active admin.
Targeted point corrections for sub-populations would be one such additional
tool.
>
>The K factor and bonus point approach are certainly more fine-tuned than an
>ad hoc approach, especially given that the processes work BY PLAYER (not by
>pool segment).
They require FUTURE performance before the adjustment occurs, meaning that
error remains with idle or semi-idle/less active players for a LONG TIME.
A quick point correction injects points immediately without waiting for future
performance, which may be strung out over months or years.
Do you wait to recalibrate your machines until months or years pass, or do you
change them as soon as you know they are out of calibration? Do you reset your
clocks right away or do you wait hours and hours?
>
>It is puzzling to me why ad hoc approaches, such as that described below
>(i.e., the rules change annually) are seen as simpler than a process.
Jeremy Spinrad made a cogent argument as to why this is the case...and I
suggest you reread his post. I am pleased to see that others are entering this
debate.
A simpler formula PLUS annual corrections...is preferred over a complicated
formula....IMHO...on grounds of parsimony of explanation.
Eric C. Johnson
>>
>>Elo defined 5 methods to control deflation. Ad hoc additions of points to
>>pool segments was not one of them. The five he defined are:
>
>Yet such ad hoc methods *were* used by FIDE (under Elo's guidance, if I
>understand correctly) to adjust the ratings of the pool of women players.
>
This was not Elo's idea.
This was 100 free rating points for every woman in the world except
for Zsuzsa Polgar was intended to stop Zsuzsa from becoming the number
one rated woman in the world in order to ensure the re-election of
Campomanes.
Sam Sloan
No...the official (and very reasonable) explanation was that:
1. The pool of women players was a closed pool
2. The pool of male players was a much larger closed pool
3. The overlap between the pools was negligible
4. The goal was to bring the women's ratings into line with the men's ratings.
5. The Polgars happened to be the (only? part of a select few?) players who
gained their ratings almost entirely in the men's pool. Thus, on the basis for
the action, their ratings did not need correction.
Under those conditions....the action (on the surface) appears reasonably
motivated.
After all, ratings are not "awards" or "property" to be gained. They are
estimates of future performance. If the estimates need correcting, you correct
them!
The fact that other real-world perks are awarded on the basis of ratings
(tangible benefits) is a separate issue....but an important one, I'll grant
you.
And the late action certainly caused hard feelings, that's true. But it is no
different from saying that X thousand active USCF players have seen their
ratings trend downward due to 10 yrs of deflation...while INACTIVE or IDLE
players have seen their idle ratings become more valuable.
In such a case, active players would get a correction, and idle ones would
not.....otherwise, the idle ones gain higher places on the relative rating list
pecking order than they otherwise would have if they were active.
In fact, the issue (the 100 pt correction by FIDE) is a good example...because
if annual point corrections (small ones) had been going on for years, then
nobody would have complained.
But instead, the pools remained separate so long that the size of the
correction was enormous (100 pts).
THAT is why it looks bad...action after a long period of inaction always looks
bad. It always looks political when it follows a long period of inaction, and
people come to expect that they "earned" their inaccurate ratings.
It would be as if the USCF allowed its scholastic sub-pool to deflate and
deflate and deflate, and only took action after years and years of inaction.
Hey...wait a second...
Eric C. Johnson
Excuse me. Your example (and I checked the archives) was that changing the K
factors would give a 2250 an advantage relative to a 2210 because the "speed"
of changes between 2210 and 2250 slowed. That has NOTHING to do with deflation.
You apparently favor lower K factors and more stable ratings except when you
oppose them.
<<snip>>
>We let things go...because a hands-off approach fits in with this idea that
>"all rating changes must be earned" nonsense.
We let things go because several EB's and ED's didn't give a hoot about
ratings. I doubt any of them did that based upon philosophies about the rating
system. They just didn't see it as all that critical.
>Your original question is incomplete...instead of a choice between:
>
>a. new system
>b. old system
>
>I offer a third alternative
>
>c. old system with frequent corrections for deflation/inflation via an active
>ratings administrator
Yeah, right. We keep a deflationary system with excessively volatile ratings
for stronger players, and ridiculously slow adjustments for lower rated players
and, someone sits and makes frequent subjective adjustments (since there's no
hard statistic which can be used to measure deflation on a short-term basis).
That's a GREAT idea, Eric.
>There may be other approaches as well...but IMHO c. is better than a very
>complex new algorithm that includes age of player and even more sliding K
>factors...
It includes the age of the player for ONE tournament. Terribly complicated.
>>
>>Only if your views have changed from the "tangible benefits" days.
>>
>
>They have not.
>Eric C. Johnson
So yesterday I misunderstood your position, and you really think lower K
factors are a good idea, except that you prefer the old system with the higher
K factors because it's simpler, and you still feel that lower K factors are
unfair because it takes more tournaments for a player to catch up to a higher
rated player in the race for tangible benefits, but you propose a complex
system for varying the timing of calculations of ratings which also reduces the
adjustment speed.
And you can't see why it would be possible for someone to come to the
conclusion that you complain merely to complain.
Tom Doan
Floors are a very crude device for adjusting for inflation and deflation. A
single very active player who has a floor even 50 points about his true
strength can put anywhere between 1 and 2 points a game into the system on
average. A player stuck 100 points high can easily be putting in 4 or 5. That's
a lot of points. There was one player who, between 1991 and 1996 put in over
4000. And they're resulting from a player who is overrated, not from a player
who is underrated, so they often end up in precisly the wrong places.
Tom Doan
First, I feel that it is silly to think that the new formula will magically
achieve the perfect balance between inflation and deflation; therefore, if you
want to make sure neither of these are occurring, you will have to modify your
formula periodically in any case. The more variables you have running around, the
harder the formula becomes to tweak without unpredictable consequences.
My more fundamental argument is, as I stated before, that it is more thrilling to
play a chess game when more rating points are on the line. Just as we might
irrationally pay some money for the daydream of winning a lottery, I gain
enjoyment from the edge of reality fantasy that my spectacular performance will
push me up to my goal rating (making expert, making master, ...). Lowering the K
factor makes this less possible, and you have still not explained the concrete
benefit. We all know that our rating is only accurate within a 50 point swing or
so from some "real" rating, but what is the gain of reducing this to 25 points?
The bonus point system also has its drawbacks which seem to be ignored. I am old
enough to remember what happened when we had a bonus system, and you would find
pockets of very active clubs where bonus points were occurring too often. For
example, take a group of class A players and play a tournament; one of them will
win and bring bonus points into the pool. Thus, bonus points suffer from the same
problems as other inflation mechanisms, though these have some actual advantages.
My proposal (annual adjustment of rating floors) makes the smallest incremental
change in a rating system which is fundamentally fine, needing only adjustment to
deal with deflation caused by the current high ratio of new vs established
players. I do not see the benefits, and see one major drawback described above,
of the more radical changes proposed.
Jerry Spinrad
As to some players adding 4 points per game: this should be pretty rare, though it
can happen under the current system. Noone has convinced me, however, that the
flaws in previous incarnations of bonus points have been fixed; I remember being
surprised that in tournaments I played in (in which almost all players were
established) the average rating gain over the entire field was more than 2 points
per person for a 4 round tournament. That was a lot of points too!
Jerry Spinrad
I agree. There is nothing magical about it. It is mathematical and must
still be reviewed periodically.
Mathematics is not magical.
>therefore, if you
>want to make sure neither of these are occurring, you will have to modify
your
>formula periodically in any case. The more variables you have running around,
the
>harder the formula becomes to tweak without unpredictable consequences.
>
Which is more unpredictable: Tweaking a formula, which can be modeled, or
tweaking the ratings of individuals or groups of players, which is complex
to model?
>My more fundamental argument is, as I stated before, that it is more thrilling
to
>play a chess game when more rating points are on the line. Just as we might
>irrationally pay some money for the daydream of winning a lottery, I gain
>enjoyment from the edge of reality fantasy that my spectacular performance
will
>push me up to my goal rating (making expert, making master, ...). Lowering
the K
>factor makes this less possible, and you have still not explained the concrete
>benefit.
The concrete benefit has, in fact, been explained ad nauseum. Please don't
confuse a lack of understanding of the benefit with a lack of its explanation.
> We all know that our rating is only accurate within a 50 point swing or
>so from some "real" rating, but what is the gain of reducing this to 25
points?
The actual volatility in performance varies at differing degrees of being
an established player.
Without knowing the exact numbers off the top of my head, let me give an
example:
The actual volatility for a senior master who has played 400 games may be
+/-40 points. Perhaps the current formula is actually volatile +/- 60 points.
For an 1800 with 400 games, perhaps the volatility is +/- 80 points, but
the current formula makes that volatility +/- 150 points.
The current degree of volatility in the point system exceeds the current
actual volatility of the players. Thus, rating swings are broader than they
should be. Ratings "apparently" go higher and lower than they should.
>The bonus point system also has its drawbacks which seem to be ignored.
I am old
>enough to remember what happened when we had a bonus system, and you would
find
>pockets of very active clubs where bonus points were occurring too often.
The new bonus system is different. I suggest you study the new system before
commenting on it.
>My proposal (annual adjustment of rating floors) makes the smallest incremental
>change in a rating system which is fundamentally fine,
Huh?? How did you come to this conclusion? My impression would be that
a function for K is much more precise and less intrusive.
How does the manipulation of floors monitor rating volatility? It would
seem that since volatility is bi-directional, and floors are uni-directional,
that the only way to minimize volatility is to raise the floor to the point
where the player finds it impossible to leave the floor -- which is abusrd.
> needing only adjustment to
>deal with deflation caused by the current high ratio of new vs established
>players.
It would also seem that you misunderstand deflation. Deflation/inflation
is not constant across the pool. Varying parts of the pool can be deflated/inflated
to varying degrees.
A formula applies bonus points or feedback points specifically to the person
who is deflated or who is otherwise in danger of becoming deflated -- i.e.
EXACTLY the person who needs it.
Who does a floor give points to? Every random opponent of the person who
is at the floor.
Which is more precise?
> I do not see the benefits, and see one major drawback described above,
>of the more radical changes proposed.
>
>Jerry Spinrad
Quite frankly, the changes are less radical. They have been changes that
have been recognized as the standard means of combatting deflation for many,
many years. And they are more precise than what you propose.
The goal of the game of chess is mate. Scholar's mate aims directly for
that. The Ruy Lopez is much less direct, there are certainly many responses
to Ruy Lopez. Moreover, one's opponent may pick an asymmetrical defense,
requiring even more knowledge. And Scholar's mate certain seems simpler.
My favorite Einstein quote is: "Make everything as simple as possible --
never simpler." Floors do just that. They are so simplified that they fail
to address the real issues. I hope that the above analysis assists in your
analysis in helping to understand why an improved formula surpasses adjustements
to floors.
I contend that the only problem is inflation/deflation. This can be attacked in a
variety of ways.
You contend that volatility is a problem, but I have seen no reason to believe
this is a problem. If you are contending that the reason volatility is a problem
has been discussed ad nauseum, I must disagree; since your phrasing seemed to
lump the issues of volatility and deflation together, it is not clear what you
meant when you contend the concrete benefits of the change had been discussed ad
nauseum.
You say that rating swings are larger than they should be. This may be correct
mathematically, but is not accurate when you discuss our enjoyment of competing
in a chess tournament. Volatility contributes only slightly (and an inflationary
way, whereas the net problem to the system is currently deflationary) to
inflation/deflation, and thus I see no reason to decrease it.
In my opinion, bonus points have some merit, if inflation needs to be guided.
This depends on how "closed" the various pools of players are. If there is plenty
of mixing, there is no real need to direct inflation towards improving players,
since the points will rapidly distribute throughout the system.I do not know how
much mixing there is, so I will bow out of the bonus point argument.
I have played tournament chess for about 30 years, and have never heard complaints
about ratings having too much volatility. I have heard many complaints about both
inflation and deflation (at different times). I have heard complaints about low K
values for higher rated players under the current system. This has been
particularly true in Quickchess, where lower K values are already in effect, and
tournament winners are often disappointed at how few points they gain. Therefore,
I feel that there should be a very clear reason for reducing these K values, and
I (and some other readers, from various posts) agree with me. Please contain your
nausea and either explain once more what harm you see in high volatility, or
point to a post where you feel the point is made particularly clearly; I may well
have missed part of the discussion.
Jerry Spinrad
The rating committee has discussed this online.
The ajustment to K impacts both volatility and deflation.
> If you are contending that the reason volatility is a problem
>has been discussed ad nauseum, I must disagree;
OK. I think you should look up the posts instead, though.
> since your phrasing seemed to
>lump the issues of volatility and deflation together, it is not clear what
you
>meant when you contend the concrete benefits of the change had been discussed
ad
>nauseum.
>
Both have.
>You say that rating swings are larger than they should be. This may be correct
>mathematically, but is not accurate when you discuss our enjoyment of competing
>in a chess tournament.
The rating system has as its primary focus the predictive ability of ratings.
I also believe, however, that incorrect volatility does impact the enjoyment
of a tournament. Volatility is bidrectional, not uni-directional.
> Volatility contributes only slightly (and an inflationary
>way, whereas the net problem to the system is currently deflationary)
Volatility is bidirectional. The above statement is false. Decreasing k
decreases volatility and also acts to decrease deflation which is a natural
result of having improving players in the pool. Excessive volatility is
more deflationary than inflationary.
>to
>inflation/deflation, and thus I see no reason to decrease it.
Since you misunderstood volatility, I can see why you would come to the wrong
conclusion. Hopefully the above will help with that.
>In my opinion, bonus points have some merit, if inflation needs to be guided.
"If inflation needs to be guided?" What does that mean?
>This depends on how "closed" the various pools of players are. If there
is plenty
>of mixing, there is no real need to direct inflation towards improving players,
Oh, I am very sorry, but there really, really, is, and this has been recognized
for 20, 30, maybe even 40 years. Improving players CAUSE deflation.
>since the points will rapidly distribute throughout the system.
No, what happens is that without properly distributing bonus points, DEFLATION
(anti-points, if you will) get distributed through the system.
>I do not know how
>much mixing there is, so I will bow out of the bonus point argument.
>I have played tournament chess for about 30 years, and have never heard
complaints
>about ratings having too much volatility.
The average player would have a difficult time discerning this. They would
recognize it as inflation or deflation.
> I have heard many complaints about both
>inflation and deflation (at different times).
Ah, the point.
> I have heard complaints about low K
>values for higher rated players under the current system. This has been
>particularly true in Quickchess, where lower K values are already in effect,
and
>tournament winners are often disappointed at how few points they gain. Therefore,
>I feel that there should be a very clear reason for reducing these K values,
There is. See the Rating of Chessplayers, by Arpad Elo, Chapter 3.
>and
>I (and some other readers, from various posts) agree with me. Please contain
your
>nausea and either explain once more what harm you see in high volatility,
or
>point to a post where you feel the point is made particularly clearly;
Quite frankly, it is difficult to provide the level of detail here. Read
chapter 3 of The Rating of Chessplayers. You will find it helpful.
> I may well
>have missed part of the discussion.
There were some posts. One that may help in understanding deflation is
the very first post in the thread: Does Average Rating Indicate Deflation?
Ask yourself the question how in this situation, a floor would have averted
the deflation of the 3 stable 1500's in the example. It would only happen
if the 1500's were already at their floor, which is not an indicator of true
stability, since true stability means that you are at your central tendency
with variability around that.
This example shows the primary means by which deflation occurs, and a floor
does not combat it.
>Jerry Spinrad
>with deflation.
>You apparently favor lower K factors and more stable ratings except when you
>oppose them.
You apparently don't allow folks to make *separate arguments*...but here goes:
1. The launch of the new system has been bungled through lack of effective
communication.
2. The new system locks in old deflation....by using lower K factors...without
a preliminary anti-deflation adjustment. Thus it harms active players and
favors recently idle ones.
3. Lower K factors are OK, per se.
4. Lower K factors *also* make it harder for folks to travel up and down the
rating scale.
I called this a scale change...because folks realize that it takes more "chess
energy" or "chess work" to move an equivalent length along the newer low-K
model.
If it took 10 games at performance level X to move 3 spaces on the old scale,
now it takes 17 games at performance level X to move 3 spaces in the new model.
I think players have a legitimate reason to bitch, if their deflated rating is
being locked into place AND it takes more chess work to move out of that
deflated rating AND there has been little or no communication about when/if/how
the new system will be implemented.
>Excuse me. Your example (and I checked the archives) was that changing the K
>factors would give a 2250 an advantage relative to a 2210 because the "speed"
>of changes between 2210 and 2250 slowed.
Indeed it does.
If lower K is implemented today...and we have two players:
Player A who is 2210 and used to be 2250 3 yrs ago and kept playing in a
deflationary environment
vs.
Player B who is 2250 and stopped playing 3 yrs ago
I hope we might agree that Player A and Player B are functionally
equivalent...the only difference is that Player A did the "good things" that we
want which is to keep on playing...and Player B did the "bad thing" which is to
sit idle.
Player A at 2210 in the deflated environment is equal to Player B in the
non-deflated environment.
Yet...now...if Player B comes out of retirement...he/she gains a considerable
edge vs. Player A.
Not only is Player B rated higher...but he is rated higher in an environment
with lower K so that it is *that much harder* for these two, functionally
equivalent players to gain the same rating level. Player A has been "locked
into" a lower pecking order...through no fault of his own.
Idleness pays off when K is going down...if past gains can be kept "frozen"....
...which is exactly why some players (including me) were complaining about the
lack of a database-wide adjustment for idle players.
Saying that idle players with high(er) ratings do not gain when K goes down (in
terms of tangible benefits) is just wrong.
Saying "just play through it" rewards the wrong group, the idle players.
>
>>We let things go...because a hands-off approach fits in with this idea that
>>"all rating changes must be earned" nonsense.
>
>
>We let things go because several EB's and ED's didn't give a hoot about
>ratings.
No, we as an organization don't appreciate the need for frequent ratings
corrections.
>
>Yeah, right. We keep a deflationary system with excessively volatile ratings
>for stronger players, and ridiculously slow adjustments for lower rated
>players
>and, someone sits and makes
>frequent subjective adjustments (since there's no
>hard statistic which can be used to measure deflation on a short-term basis).
Yes.
Simple.
But lacking that "surgical precision" that you RC types like...which is why
membership will soar once our rating system is insanely complicated but
ever-so-accurate, yes?
>
>So yesterday I misunderstood your position, and you really think lower K
>factors are a good idea, except that you prefer the old system with the
>higher
>K factors because it's simpler, and you still feel that lower K factors are
I don't care what the K factor is, so long as one acknowledges that at the
point one changes K...one also has to recognize that one needs to make a
systemic correction for active/idle players.
Otherwise, you are allowing "error" to sit in the system (literally) for
years....and rewarding idleness.
>but you propose a complex
>system for varying the timing of calculations of ratings which also reduces
>the
>adjustment speed.
>
Who said anything about a complex system for timing calculations?
For chrissakes...we issue official ratings every two months. We don't need
them every two days! Not only are more frequent "official" ratings less
accurate, but they are very misleading.
> simpler, and you still feel that lower K factors are
>unfair because it takes more tournaments for a player to catch up to a higher
>rated player in the race for tangible benefits,
Absent a correction AT THE TIME YOU LOCK PLAYERS IN WITH LOWER K, that is
correct.
Eric C.Johnson
>
> You say that rating swings are larger than they should be. This may be
> correct mathematically, but is not accurate when you discuss our
> enjoyment of competing in a chess tournament. Volatility contributes
> only slightly (and an inflationary way, whereas the net problem to the
> system is currently deflationary) to inflation/deflation, and thus I
> see no reason to decrease it.
Well, try this point of view. Excessive volatility, combined with
"modest" (I love that word - what does it mean?) deflation *is* a
problem.
Assume for the moment that there has been modest deflation over the past
10 years. Let's not worry about whether or not this deflation was
politically determined to be a "good thing" or not. Let's just assume a
shift in the scale of some 25-30 rating points, over the past 10 years.
Now, add in some players with a natural swing of some 100 points to
either side of their "correct" and "stable" rating.
Finally, assume that the average player considers all *drops* in his
rating to be "a bad day" and all *rises* to his rating to be "finally
playing up to my potential".
The sum of all these effects leads to players seeing an "obvious" 125
point drop (or even a 225 point drop) in their rating. Add in a little
understanding of basic rating theory, and a lot of normal human
psychology, and it's "obvious" to everyone that there has been massive,
definitely "immodest", systematic and debilitating deflation.
Some volatility is inevitable, and even desirable (for example, without
some volatility, players would never reach their "true" rating!).
Excess volatility *is* a problem. Fortunately, there are ways to
determine how much volatility is "natural" and how much is "too much".
Unfortunately, these methods are not "obvious".
> particularly true in Quickchess, where lower K values are already in
> effect, and tournament winners are often disappointed at how few
> points they gain.
Different case. Here, we agree. The original argument for lower K in
Quick chess was (in my opinion) bogus. It was "obvious"- but wrong.
> Therefore, I feel that there should be a very clear reason for
> reducing these K values
Let me ask you: do you agree with the motivation for the original
"stepped-K" in the old system? You know, the one that said K changes
dramatically when your rating goes through the 2100 and 2400 barriers?
The implementation was "stepped" because it was conceived in a time when
much of the computation was done by hand. But, the argument for
"different K for different Ro" is ancient.
The difference now is that computation is essentially free - making some
of the simplifications in the original system no longer desirable.
The actual value of the "correct" K is subject to objective test. The
original system got it mostly correct, but used a 3-step piecewise
CONSTANT function to approximate the correct values. The new system
uses a better (and more complicated) computation.
Yes, there are some values of Ro for which the old system had a K which
was objectively "too high". The new system fixes that. It also fixes
the much more serious problem that under the old system there were many,
many more values of Ro for which the old system had a K which was
objectively "too low".
Ratings of 2200 *should* be relatively "locked in".
Ratings of 0200 *should not* be "locked in".
In between is, well, in between.
--
Kenneth Sloan sl...@uab.edu
Computer and Information Sciences (205) 934-2213
University of Alabama at Birmingham FAX (205) 934-5473
Birmingham, AL 35294-1170 http://www.cis.uab.edu/info/faculty/sloan/
> >You say that rating swings are larger than they should be. This may be
correct
> >mathematically, but is not accurate when you discuss our enjoyment of
competing
> >in a chess tournament.
>
> The rating system has as its primary focus the predictive ability of
ratings.
> I also believe, however, that incorrect volatility does impact the
enjoyment
> of a tournament. Volatility is bidrectional, not uni-directional.
But have you ever known a chessplayer when asked how strong he was, not to
tell you his highest rating ever attained. The mathematics may differ from
the psychology.
- Tom Martinak
I'm perfectly happy with separate arguments if they aren't in complete
contradiction with each other.
>1. The launch of the new system has been bungled through lack of effective
>communication.
>
>2. The new system locks in old deflation....by using lower K
>factors...without
>a preliminary anti-deflation adjustment. Thus it harms active players and
>favors recently idle ones.
>
>3. Lower K factors are OK, per se.
>
>4. Lower K factors *also* make it harder for folks to travel up and down the
>rating scale.
>
Yes, and that has zippo to do with past deflation.
<<snip - we don't need to read that crap for the fortieth time>>
>>Excuse me. Your example (and I checked the archives) was that changing the K
>>factors would give a 2250 an advantage relative to a 2210 because the
>"speed"
>>of changes between 2210 and 2250 slowed.
>
>Indeed it does.
>
>If lower K is implemented today...and we have two players:
>
> Player A who is 2210 and used to be 2250 3 yrs ago and kept playing in a
>deflationary environment
>
>vs.
>
>Player B who is 2250 and stopped playing 3 yrs ago
>I hope we might agree that Player A and Player B are functionally
>equivalent...the only difference is that Player A did the "good things" that
>we
>want which is to keep on playing...and Player B did the "bad thing" which is
>to
>sit idle.
>Player A at 2210 in the deflated environment is equal to Player B in the
>non-deflated environment.
>
>Yet...now...if Player B comes out of retirement...he/she gains a considerable
>edge vs. Player A.
And what about the player who is 2210 because he just had a bad tournament. Or
the 2250 who just had a good tournament. To say that it's only unfair under
certain circumstances, which we can't fully identify, is nonsense.
<<snip>>
>Not only is Player B rated higher...but he is rated higher in an environment
>with lower K so that it is *that much harder* for these two, functionally
>equivalent players to gain the same rating level. Player A has been "locked
>into" a lower pecking order...through no fault of his own.
>
>Idleness pays off when K is going down...if past gains can be kept
>"frozen"....
>
>...which is exactly why some players (including me) were complaining about
>the
>lack of a database-wide adjustment for idle players.
Name another.
Also, it was pointed out (repeatedly) that your proposed "solution" would have
been just as unfair (if not more so) to a different group of players - the
active players who had recent success. RATINGS GO BOTH WAYS, ERIC. You are so
full of #$#%$ on this that it isn't funny.
>Saying that idle players with high(er) ratings do not gain when K goes down
>(in
>terms of tangible benefits) is just wrong.
PLAYERS WITH HIGHER RATINGS GET THAT "BENEFIT" REGARDLESS OF WHETHER THEY WERE
IDLE. Period. End of sentence. End of argument.
>Saying "just play through it" rewards the wrong group, the idle players.
>
>>
>>>We let things go...because a hands-off approach fits in with this idea that
>>>"all rating changes must be earned" nonsense.
>>
>
>>
>>We let things go because several EB's and ED's didn't give a hoot about
>>ratings.
>
>No, we as an organization don't appreciate the need for frequent ratings
>corrections.
>>
>>Yeah, right. We keep a deflationary system with excessively volatile ratings
>>for stronger players, and ridiculously slow adjustments for lower rated
>>players
>>and, someone sits and makes
>
>>frequent subjective adjustments (since there's no
>>hard statistic which can be used to measure deflation on a short-term
>basis).
>
>Yes.
>
>Simple.
That's right. Really simple. A person or person unspecified makes changes using
a set of tools which you can't even list to achieve a vague and unquantifiable
goal. That's your definition of simple. Actually, I agree with you.
Simple - uneducated, mentally retarded, half-witted, naive
Sounds about right.
>But lacking that "surgical precision" that you RC types like...which is why
>membership will soar once our rating system is insanely complicated but
>ever-so-accurate, yes?
No one has ever claimed that it would do that.
>>
>>So yesterday I misunderstood your position, and you really think lower K
>>factors are a good idea, except that you prefer the old system with the
>>higher
>>K factors because it's simpler, and you still feel that lower K factors are
>
>I don't care what the K factor is, so long as one acknowledges that at the
>point one changes K...one also has to recognize that one needs to make a
>systemic correction for active/idle players.
>Otherwise, you are allowing "error" to sit in the system (literally) for
>years....and rewarding idleness.
>
>
>
>>but you propose a complex
>>system for varying the timing of calculations of ratings which also reduces
>>the
>>adjustment speed.
>>
>
>Who said anything about a complex system for timing calculations?
You did. Rating some players across a larger time interval and others
tournament by tournament. That's not complex? Of course it's not in your world,
where you toss out stupid ideas and refuse to ever describe how you would
actually accomplish them.
>For chrissakes...we issue official ratings every two months. We don't need
>them every two days! Not only are more frequent "official" ratings less
>accurate, but they are very misleading.
OK. Chew on this one. Who said the following in October:
>[The new rating system will]
>Hurt us because 1) players don't understand it, and 2) the >longer rating
periods [referring to the change from weekly to >biweekly rates] make players
fear that events are being >missed/not sent in/not processed.
>We'll take a modest reduction in predictive accuracy for a
>simple/transparent system.
>Eric C. Johnson
So three months ago, when we went from weekly to biweekly (not on the RC's
recommendation, BTW) that was REALLY bad because people couldn't check on their
ratings. But what we really need is a rating system which holds ratings for two
months and then aggregates a number of tournaments making it damned near
impossible for anyone to check their ratings.
And, BTW, what is your "solution" for the "scale change" that would be imposed
by implementing the type of rating system that you are not proposing?
You're a real piece of work, Eric. As is typical, this grows wearisome after a
while.
Tom Doan
Not as rare as one would like. That's about 100 points overrated. There were a
not insignificant number of player who went floor to floor pretty quickly when
the floors went down in 1996.
Noone has convinced me, however, that
>the
>flaws in previous incarnations of bonus points have been fixed; I remember
>being
>surprised that in tournaments I played in (in which almost all players were
>established) the average rating gain over the entire field was more than 2
>points
>per person for a 4 round tournament. That was a lot of points too!
>
>Jerry Spinrad
Actually, with the change in the K factors, it will end up being pretty close
to that (1/2 point per game on average), though lower for higher rated players.
But that's for a mix which includes improving players, not stable ones.
Tom Doan
>
>Yes, and that has zippo to do with past deflation.
>
><<snip - we don't need to read that crap for the fortieth time>>
You do if you are too dense to understand that:
LOCKING IN THE CURRENT DEFLATED RATINGS via lower K is what is being objected
to...not lower K per se.
If there is an adjustment prior to the change in K, all is well.
>
>>...which is exactly why some players (including me) were complaining about
>>the
>>lack of a database-wide adjustment for idle players.
>
>Name another.
For chrissakes, Tom...there were independent posts by other readers complaining
about this too.
>
>Also, it was pointed out (repeatedly) that your proposed "solution" would
>have
>been just as unfair (if not more so) to a different group of players - the
>active players who had recent success.
>RATINGS GO BOTH WAYS, ERIC. You are so
>full of #$#%$ on this that it isn't funny.
I fully expect to see Liam chime in with "ratings are not the only things that
go both ways", so I'll beat him to the punch.
Look...active players are deflated. Idle players are inflated, due to playing
less under deflation.
No, you cannot help each individual case. But the *groups* have these
attributes...and we are talking about corrections to the groups, not individual
players.
Your "chips fall where they may" defense is equally applicable to my position,
except that I assert that by recognizing the fact that deflation hurts active
players more than idle ones...and that this recognition places us closer to an
understanding of the real situation.
All I sense from you is bitterness...and a big attitude.
>
>PLAYERS WITH HIGHER RATINGS GET THAT "BENEFIT" REGARDLESS OF WHETHER THEY
>WERE
>IDLE. Period. End of sentence. End of argument.
Wrongo. Idle players might get a point correction to reduce their higher
ratings.
Or active players might get a targeted point correction to raise their ratings.
>
>PLAYERS WITH HIGHER RATINGS GET THAT "BENEFIT" REGARDLESS OF WHETHER THEY
>WERE
>IDLE. Period. End of sentence. End of argument.
>
Also, you are simply not willing to admit that part of the reason these "idle"
higher rated players ARE higher rated...is due to an artifact.
The artifact is that they sat out, and so experienced less deflation...NOT that
they are deserving of higher ratings.
>
>Simple - uneducated, mentally retarded, half-witted, naive
A person who speaks to delegates in this fashion has no business being on any
committee.
>
>You did. Rating some players across a larger time interval and others
>tournament by tournament.
Never said it...strawman.
>
>>Hurt us because 1) players don't understand it, and 2) the >longer rating
>periods [referring to the change from weekly to >biweekly rates] make players
>fear that events are being >missed/not sent in/not processed.
Yep...cuz we already implemented weekly rating updates years ago. I was against
them then. I'm against them now.
But the genie is out...and the service expectation is there.
Last update was 12/12 or thereabouts...new one is due 1/10...that's a month.
It's not by design...but because of sloth and lack of resources...in the FACE
of a service expectation.
I didn't want to create that service expectation..but it is there. The longer
period is now perceived as POOR SERVICE.
THAT is a fact. My opposition to frequent official updates is a separate
matter.
Are you really this dense? That you cannot see that a person might recognize a
perceived customer service issue in one post...AND...make a comment about how
more frequent official updates actually don't buy us any greater power...in
another?
>
>So three months ago, when we went from weekly to biweekly (not on the RC's
>recommendation, BTW) that was REALLY bad because people couldn't check on
>their
>ratings.
It was bad because USCF has, through faulty policy, built up an expectation of
service that is not being met right now. Not by a long shot.
That is very different from my separate argument that ever-more-frequent rating
updates (the extreme being the daily game-by-game updates of online services)
do not buy you anything other than more noise.
>But what we really need is a rating system which holds ratings for two
>months and then aggregates a number of tournaments making it damned near
>impossible for anyone to check
>heir ratings.
>
Hardly impossible..and I don't see how anyone could argue for "official" rating
changes based on only a small set of games.
A larger set of games is needed, else any changes are just noise.
>
>You're a real piece of work, Eric. As is typical, this grows wearisome after
>a
>while.
>
>Tom Doan
Yes it does, because you appear to have zero comprehension skills for context.
Eric C. Johnson
Indeed that is true, which why the EB makes a gross error by stocking the
Ratings Committee with just mathematicians.
A few marketing or psychology types in the mix would avoid these kind of
"mathematically good, consumer psychology bad" types of actions.
NOBODY complains about volatility...or "lack of predictive power"...but they do
complain when stable Experts at their clubs are now 1850....and they can see it
happen...and are told "just play through it"
Eric C. Johnson
>>
>>Simple - uneducated, mentally retarded, half-witted, naive
>A person who speaks to delegates in this fashion has no business being on any
>committee.
I was referring to the idea. That's an accurate characterization of it.
>>
>>You did. Rating some players across a larger time interval and others
>>tournament by tournament.
>
>Never said it...strawman.
OK, Eric. You tell me PRECISELY what you were recommending. I already asked you
to do that, and you snipped that out. I think that's an accurate restatement of
what you said.
<<snip>>
Tom Doan
I don't like to hear these blanket assumptions. I am not at all like this.
There was a stretch of a year or so, in the early to mid'90s, when I was
probably playing at low-master strength, but I did not stay active long
enough to get the title.
Since then, I have been too busy with work and family to play often or study
at all, and my last few performance ratings were probably below 2000.
I would be the first to tell you that my current strength is probably Class
A, if that.
I would also be the first to tell you that I could certainly earn the master
title if I tried, despite "Soltis's Law." And so, I believe, could most
A-players. The main reasons that players decline in practical strength as
they age, at amateur levels anyway, have more to do with motivation and
priorities than physical or mental decline.
Timothy Hanke
In my experience, I have usually gained more points than I expected, or lost
fewer. I don't know why. Any theory?
Tim
> > But have you ever known a chessplayer when asked how strong he was, not
to
> > tell you his highest rating ever attained. The mathematics may differ
> > from the psychology.
>
> I don't like to hear these blanket assumptions. I am not at all like this.
> There was a stretch of a year or so, in the early to mid'90s, when I was
> probably playing at low-master strength, but I did not stay active long
> enough to get the title.
You more than prove my point. You told us you were even stronger than the
highest rating that you ever attained.
- Tom Martinak
> >But have you ever known a chessplayer when asked how strong he was, not
> >to tell you his highest rating ever attained. The mathematics may differ
from
> >the psychology.
> >
> > - Tom Martinak
> >
> Again, volatility is bi-directional. Have you ever known a player who
felt
> he lost more points in a tournament than he should have?
You seem to have missed my point. More volatility involves a wider range of
ratings for a player (both higher & lower). But the players tend to view
their strength by their highest rating. So they are happiest with more
volatility because their highest rating will tend to be higher in that case.
- Tom Martinak
My counterpoint is that when a player has that random bad event, they
also lose too many points.
I can use myself for an example. I had several good events in a row,
and peaked around 2320. Then I had two horrible events, and dropped
to 2200.
Should EITHER event have happened? Probably not. My real strength was
likely in the 2250-2270 range. A narrower K would more actively
reflect that.
I'm sorry Tom. I prefer to go with what is real and measurable.
--
Kevin Bachler
Caveman
"Caveman chess is chess without finesse."
Sent via Deja.com
http://www.deja.com/
> No I understood your point. I was making the counterpoint that it is
> incomplete. Of course players think of their peak. The system is
> concerned with predictability, not with peaks. If we wanted to make
> the players happy with high ratings, make K=1000. Or 10,000.
Actually where our confusion lies is in the purpose of the rating system.
You state that "system is concerned with predictability" as if everyone
accepts that as the purpose. I would instead claim that the purpose is to
stimulate chess activity. Certainly predictability is one feature that does
this, but not the only one. There needs to be a balance, but the balance is
one of psychology not math.
- Tom Martinak
The problem is that the one system cannot serve both masters. It's been
recognized (for a long time) that it's necessary to complement a rating system
geared for predicting performance with some type of relentlessly non-decreasing
measure of peak or consistent performance. The original title system was
designed to do that, but suffered from a number of serious flaws in practice.
The Lifetime Achievement awards have been passed by the delegates, but, like
most such things, await someone to decide to implement them.
Tom Doan
Jerry Spinrad
PS: I invite you to a unique inflationary event. We would love to see you
at the Northern Tennessee winter open this Saturday/Sunday. This event is
structured in such a way (slow time controls and separate sections for high rated
players) that I will no doubt be offering many points into the rating pool. As a
former beneficiary of some of my donated points in a similar event, please feel
free to take advantage of this unique opportunity!
I don't disagree that the rating system can spur activity. I am opposed
to the types of changes that permanently harm the primary purpose of the
system, which is predictability.
Yes, predictability is the primary purpose, Tom. If it weren't, we'd still
be using the old Harkness system that it replaced, or K would be 10,000.
I prefer to better educate people about what ratings mean, rather than delude
them with false formula. In essence, a larger K is sizzle, no steak and
given that we now have an alternative (i.e. that we can do better) it is
a falsehood, fair and square.
Meanwhile, you are still looking at volatility uni-directionally, as Ablue
and many others have done during this discussion. Decreasing the losses
(yes, people really do lose points from time to time) helps the psychological
impact just like bigger gains do.
Suppose a player peaks at 1850. Which is more important: more volatility
so that they will some day "accidentally" hit 1900, or less volatility so
that some day they won't accidentally hit 1799?
Which one impacts psychology more? (I offer a draw on this.) Which one
impacts prize distribution more (you can resign that position.)
Lower volatility wins the match.
Here, among other things, Tim is making the useful point that it's good
for your chess to be objective about your current playing strength.
Some posters in this thread seem to be saying that predictive accuracy
is (from a psychological point of view) fairly low on the list of
players' priorities. This is an interesting point, but it's not the
same as saying that nobody cares about predictive accuracy, so long as
we get the appropriate psychic reinforcements out of the rating system.
For example, the predictive accuracy of ratings is one of the
presuppositions that make fair Swiss pairings possible. This is one of
the many problems with "activity points". About a year down the road
(if activity points are actually implemented) we'll see two kinds of
2000 players: the ones who have been active recently, with ratings
boosted artificially by activity points, and the ones who haven't
played much and are therefore still roughly 2000 players by today's
standards.
Players who are looking to maximize their scores (and their ratings)
will obviously prefer to be paired with players in the first category
rather than players in the second. This will give Swiss pairings a
noticeably more lottery-like character than they have now. I'm sure
that this will all be common knowledge after a while. Someone will
think of a catchy and derisive nickname for the nouveau A players and
experts who have bloated their ratings just by playing a lot. And the
whole situation will promote an attitude of renewed cynicism about the
integrity of the USCF rating system.
Larry
In article <FS976.42810$1M.94...@typhoon.ne.mediaone.net>,
Sent via Deja.com
http://www.deja.com/
> The reason for both the past and current "different K values"
> formula is > logical. On average, a 2000 player is changing
> strength at a slower pace than a 1200 player.
Excellent start - but let's finish the point. You have given the reason
or a *difference* in K for 1200 players and 200 players - but you
haven't said anything about the absolute value of K. How do you think
these values should be chosen:
a) by tradition (we did it this way last year...)
b) by political means (Resolved: the EB sets K to be 100 for 6
months)
c) ???
> I do not think that the K factor is
> too high anywhere along the current spectrum, if player enjoyment
> is the dominant concern.
What values would be "too high"? What values would be "too low"? On
what basis do you make such a judgement?
Let's pick 2000 players as an example. What do you think the value of K
should be? Why?
Sorry I won't make it to your tournament this weekend. I have fond
memories of our last meeting, and often use it as an object lesson.
> I don't disagree that the rating system can spur activity. I am opposed
> to the types of changes that permanently harm the primary purpose of the
> system, which is predictability.
But again, you are assuming that everyone agrees that the primary purpose of
the rating system is predictability. I disagree & I expect that many other
people whom you have been disagreeing with do also.
> Yes, predictability is the primary purpose, Tom. If it weren't, we'd
still
> be using the old Harkness system that it replaced, or K would be 10,000.
It is a purpose, but I don't think it is the primary purpose. The rating
system is used to spur chess activity. That is why deflation is important.
If predictability is all that matters, then relative rating difference is
all that matters, not the traditional "titles" of Master, Expert, A, B, etc.
> I prefer to better educate people about what ratings mean, rather than
delude
> them with false formula. In essence, a larger K is sizzle, no steak and
> given that we now have an alternative (i.e. that we can do better) it is
> a falsehood, fair and square.
But often it is the sizzle that wins. What kind of VCR do you have? What
kind of computer - Apple or PC.
> Meanwhile, you are still looking at volatility uni-directionally, as Ablue
> and many others have done during this discussion. Decreasing the losses
> (yes, people really do lose points from time to time) helps the
psychological
> impact just like bigger gains do.
But from the psychological standpoint, the wins are much more important than
the losses. Why do people go to Las Vegas? They will on average lose.
> Suppose a player peaks at 1850. Which is more important: more volatility
> so that they will some day "accidentally" hit 1900, or less volatility so
> that some day they won't accidentally hit 1799?
And suppose he peaks at 1950? Which is more important, more volatility so
that he will someday peak at 2000 or that some day he won't accidently hit
1899?
> Which one impacts psychology more? (I offer a draw on this.) Which one
> impacts prize distribution more (you can resign that position.)
I claim a win on psyxhology and a win on increased chess activity. I make
2-1 my way.
- Tom Martinak
No, it is not an assumption Tom. It is an observation. It is easy to see
that USCF previously rejected a rating system that had little predictive
ability for one that had significant ability. It is easy to see that refining
that system, to make it more difficult for sandbaggers or to improve pairings,
has been a frequent subject of discussion. It is recognizing what is in
place. That is not an assumption. An assumption supposes "What if...?"
I am not saying "What if..." we had a highly predictive rating system instead
of an unpredictive one, I am regognizing the fact that we ACTUALLY HAVE THAT.
Now, you may not like that. That's fine. you may disagree that that should
be the primary purpose. That's fine. But denying that it is actually the
primary puspose is denying what actually exists -- and that's lunacy.
>I disagree & I expect that many other
>people whom you have been disagreeing with do also.
This is like disagreeing with the idea that gravity pulls things down. You
can disagree all you want, but it doesn't change that gravity pulls things
down. You can disagree with whether WE SHOULD HAVE a rating system that
is highly predictive -- but disagreeiing with what we do have is just denying
reality. I don't cotton to that.
>> Yes, predictability is the primary purpose, Tom. If it weren't, we'd
>still
>> be using the old Harkness system that it replaced, or K would be 10,000.
>
>It is a purpose, but I don't think it is the primary purpose.
I understand that you don't think that. Then why do we have a rating system
with the primary attribute that it is predictive?
> The rating
>system is used to spur chess activity.
The Harkness system spurred activity -- that was its primary purpose. Why
was it rejected if that is USCF's primary purpose in a rating system?
> That is why deflation is important.
No, deflation is important because it destroys predictive ability.
>If predictability is all that matters, then relative rating difference is
>all that matters, not the traditional "titles" of Master, Expert, A, B,
etc.
>
Then you misunderstand deflation, number 1, and you misunderstand predictive
ability, number 2.
Deflation is not constant across the pool. The impact of deflation is that
different levels of the pool will deflate differently depending on the factors
causing deflation. Consequently, predictive ability is harmed.
The sectors for this may also be geographic, or other social factors.
Deflation is also important because it changes MEANING, which may or may
not impact activity. People understand the system better if a stable 1500
yesterday is a stable 1500 today is a stable 1500 tomorrow, whether or not
that impacts their activity.
Deflation consequently impacts pairings and prize distribution (i.e. class
prizes) and again unevenly. And these are all items based on predictability.
Meaningful odds in a single game really only occur at 3:1. Those are the
odds defined by 200 pionts, a class interval.
>> I prefer to better educate people about what ratings mean, rather than
>delude
>> them with false formula. In essence, a larger K is sizzle, no steak and
>> given that we now have an alternative (i.e. that we can do better) it
is
>> a falsehood, fair and square.
>
>But often it is the sizzle that wins.
Not in my world Tom, or in the world of business I deal in. Sizzle may win
for a while, but the good steak gets rewarded. Sizzle is needed to get attention
to the steak, but you go to a meal to eat the steak, not to listen to the
sizzle.
>What kind of VCR do you have? What
>kind of computer - Apple or PC.
PC -- So that I could build my own.
>> Meanwhile, you are still looking at volatility uni-directionally, as Ablue
>> and many others have done during this discussion. Decreasing the losses
>> (yes, people really do lose points from time to time) helps the
>psychological
>> impact just like bigger gains do.
>
>But from the psychological standpoint, the wins are much more important
than
>the losses.
I disagree.
A while ago we heard people argue in this forum that people quit playing
because their ratings dropped. That statement doesn't jive with the argument
above.
I believe they are equally important, Tom. Fear is known to be a great motivator,
for many people a much better motivator than reward.
Do you associate fear with winning or losing. Do you associate reward with
winning or losing.
I think your thesis that psychologically winning is more important is just
bogus. Will people TALK ABOUT their peak rating more than their nadir?
Of course. But do they care more about hitting their nadir or their peak
-- I think to say anything other than it's unclear is hogwash, Tom.
>Why do people go to Las Vegas? They will on average lose.
>
Yes -- so? They expect to -- its a vacation. Any vacation will cost money,
and they have some chance to win, so why not? A rationalization -? Sure.
But psychologically which means more to an average person: Winning $10K or
losing $10K?
Tom -- which do you better remember -- good things or bad? (Most people
will say good -- in fact, there are arguments you have no real memory of
pain, but you can remember warmth, love)
Does that make the pain less important -- OR MORE important?
>> Suppose a player peaks at 1850. Which is more important: more volatility
>> so that they will some day "accidentally" hit 1900, or less volatility
so
>> that some day they won't accidentally hit 1799?
>
>And suppose he peaks at 1950? Which is more important, more volatility
so
>that he will someday peak at 2000 or that some day he won't accidently hit
>1899?
>
EXACTLY!! Thank you for agreeing! You finally see that its equal and dependent
on ever so slight circumstances.
>> Which one impacts psychology more? (I offer a draw on this.) Which one
>> impacts prize distribution more (you can resign that position.)
>
>I claim a win on psyxhology and a win on increased chess activity. I make
>2-1 my way.
>
Sorry, your claim of a win is denied. Your down 3 pieces with no mating
material left.
> - Tom Martinak
> No, it is not an assumption Tom. It is an observation. It is easy to see
> that USCF previously rejected a rating system that had little predictive
> ability for one that had significant ability. It is easy to see that
refining
> that system, to make it more difficult for sandbaggers or to improve
pairings,
> has been a frequent subject of discussion. It is recognizing what is in
> place. That is not an assumption. An assumption supposes "What if...?"
> I am not saying "What if..." we had a highly predictive rating system
instead
> of an unpredictive one, I am regognizing the fact that we ACTUALLY HAVE
THAT.
And the USCF EB just voted last year to add "fiddle points". It looks to me
like predictive ability is just one feature of the USCF rating system and in
the current situation isn't considered the most important by the governance
structure. You are obviously assuming too much importance for predictive
ability. It seems clear to me that K value has much less effect upon
predictive ability than fiddle points. So I'm not down 3 pieces, I'm up a
queen.
- Tom Martinak
My first impulse is to say that we would like a players rating to be within 100
points of what it should be at least 95% of the time; others can suggest
different values. That is, if you assume a player is gaining/losing points not
through changes in strength, but through random fluctuation. Probably the easiest
way to check on this is via a simulation, by randomly generating imaginary
opponents within 300 points of the original rating, and letting the player
win/lose with probability given by the usual formula.
It would be interesting to see how high you can make the K go and satisfy this
requirement, and what happens as you tweak the deviation allowed and
proabability. It seems to me that if you meet this requirement, then the ratings
are doing a sufficiently good job of prediction to meet the requirements of
tournament pairings and so forth, so that we are free to vary within this
guideline in order to maximize player enjoyment.
I take issue (not yours, but Mr. Doan's) that the rating cannot serve the "double
masters" of predicting accurately and satisfying players; in fact, it has served
both of these masters quite well, without having to choose fanatically between
one or the other.
Asking me what the K should be for a 2000 player is to the point, since my rating
is currently 2016 (though this is essentially my peak, so I feel I am overrated).
I see no problem at all with the traditional formula for players like me, i.e.
gain/lose about 16 points for a win vs. an equal opponent. I feel no differently
about how much I would expect to gain for a win than I did when I was a 1400
player, or an 1800 player.
Jerry Spinrad
By the way, if you ever want a fond object lesson for the Alabama Chess
Federation newsletter, the unusually noble actions of (David?) Presley, a
Chattanooga player, were responsible for my single win of a Grand Prix
tournament. He had taken in advance a bye for the last round, but went on a huge
upset streak (he was rated only 1700) and was tied with me for the lead with one
round to go. He felt that it would be inappropriate to win in this way without
facing me or Todd Andrews, who was 1/2 point behind, and thus insisted on taking
a 0 point bye rather than a half point bye. I was thus able to win the tournament
by drawing Todd; this was a few years ago, I would draw him much less frequently
now. He sacrificed the money and title voluntarily; I was quite impressed.
Agreed. It is one factor, and I've said that.
> and in
>the current situation isn't considered the most important by the governance
>structure.
Again, deal with reality: Did that REMOVE the current predictive system,
or did they AUGMENT it for 1 year for SOME (not all) of the players -- specifically
eliminating the players where it was felt predictive ability was most important?
Deal with what's real, Tom.
> You are obviously assuming too much importance for predictive
>ability.
See above. What part was assumed, and what part was observed.
I think you are obviously assuming too much importance for psychology. I
have the facts above to support that.
> It seems clear to me that K value has much less effect upon
>predictive ability than fiddle points. So I'm not down 3 pieces, I'm up
a
>queen.
>
Better look at the board again. I just Queened all 8 pawns and you only
have a King left.
> - Tom Martinak
> Again, deal with reality: Did that REMOVE the current predictive system,
> or did they AUGMENT it for 1 year for SOME (not all) of the players --
specifically
> eliminating the players where it was felt predictive ability was most
important?
I'm not sure what you are trying to say there. However, I'm pretty sure
that if you asked people on the Ratings Committee to choose between having
the uniform K currently used (even extending it up over 2100) or having
"fiddle points", then they would tell you that the "fiddle points" hurt
predictive ability more. So it seems clear to me that the current USCF
govenance values predictive ability much less than the K value would hurt
it. Dealing with the curent reality, predictive ability for the USCF isn't
the holy grail that you are making it out to be. You may not agree with
that - but that is the reality.
- Tom Martinak
I certainly do not want to see a rating system in which I was just going up. I
want it to basically measure my ability, and would not value a master title
obtained simply by playing enough games. However, bouncing up and down by 25%
more does not bother me at all, and I still contend that Kevin has not clearly
made his point about why this "problem" which is not perceived as an issue by
current players needs to be fixed.
Jerry Spinrad
I understand. Try this.
Did they permanently replace the current predictive system with an unpredictive
system, like the Harkness system, or did they augment the current system
for a limited amount of time with a methodology intended to increase activity
-- but in doing so maintained the current predictive system, put the augmentation
in place for only 1 year, AND did not augment the players for whom a predictive
system is most important (those above 2000)?
That is, dealing with reality : YES, there were changes made to the rating
system that highlight the psychological and activity aspects of the system,
AND NO those aspects were not stressed more than the predictive nature of
the system (as noted by the temporary augmentation of a portion of the pool.)
> However, I'm pretty sure
>that if you asked people on the Ratings Committee to choose between having
>the uniform K currently used (even extending it up over 2100) or having
>"fiddle points", then they would tell you that the "fiddle points" hurt
>predictive ability more.
That may well be true, Tom, but what that means is that the old system is
more predictive for 1 year than activity points are -- it does not mean that
the predictive ability is viewed as unimportant or viewed as anything other
than the top priority.
> So it seems clear to me that the current USCF
>govenance values predictive ability much less than the K value would hurt
>it.
I don't understand that sentence. If it means what I think you mean, I disagree,
because the needed changes were ALSO made to the system, AND the augmentation
has a 1 year sunset clause.
> Dealing with the curent reality, predictive ability for the USCF isn't
>the holy grail that you are making it out to be.
Then I have to doubt whether you are thinking about how it impacts pairings,
prizes, and other factors.
> You may not agree with
>that - but that is the reality.
>
No, the reality is what I described through observation. Again, the predictive
system stayed in place, the necessary changes were made, and the unpredictive
augmentation for activity points is limited both in time and scope. In other
words, everything that is predictive has been prioritized, and everything
that is not predictive is temporary.
How you interpret this to mean otherwise I find confusing.
It reminds me of an NFL player who intercepts a pass, runs it in 30 yards
for a touchdown, flips head over heals and dances around -- for closing the
score to 77-6. Yeah, its a touchdown, and it's immediate, and we pay attention
to it for the moment, but in the bigger picture what did it mean?
Kevin L. Bachler
The explanation was very clear, and is a very common occurance when people
observe probablistic events.
I suggest you look at Doan's sample post of 3 tournaments and only then try
to contend that people will be severely disappointed by the number of points
they can gain.
> No, the reality is what I described through observation. Again, the
predictive
> system stayed in place, the necessary changes were made, and the
unpredictive
> augmentation for activity points is limited both in time and scope. In
other
> words, everything that is predictive has been prioritized, and everything
> that is not predictive is temporary.
>
> How you interpret this to mean otherwise I find confusing.
Well, you have a much different interpretation of the effects of "fiddle
points" than I do. And from the postings here, a much different
interpretation than many of the people on the Ratings Committee. What
choice do you think the Ratings Committee would have made if they were told
to pick: (A) Leave the system as it is; (B) Make the changes to the system &
add "fiddle points" for a year; (C) Implement the new system with K doubled.
I'm pretty sure that (B) would be the last choice. And of course, you are
assuming that once a year is up, "fiddle points" will go away. What is the
USCF track record at making software changes. I remember many years ago,
complaining when they dropped the rating floor 100 points because it was the
only anti-deflation part of the system (though admittedly not a very good
one). I was told not to worry. They had designed this new system & it
would be implemented in less than a year. Well, that's a lot longer than a
year ago.
- Tom Martinak
Note that my only objections are to those places where K values have been
lowered, so I am not trying to quarrel about changes among the other groups of
players. For these players, both the winners and losers, the thrills of the
tournament have become somewhat smaller.
Jerry Spinrad
No, I actually agree with most of what the rating committee wrote. I've
spoken to committee members about it, and they understand that I agree.
> What
>choice do you think the Ratings Committee would have made if they were told
>to pick: (A) Leave the system as it is; (B) Make the changes to the system
&
>add "fiddle points" for a year; (C) Implement the new system with K doubled.
>I'm pretty sure that (B) would be the last choice.
With those choices, I ma not at all certain that your statement is correct.
More problematic though, is that you think the statement is relevant.
The question is what is the higher priority for the rating system, psychology
or predictive ability -- not the relative psychology or relative predictive
ability for variants.
ALL of the choices you give are more focused on predictive ability than on
psychology. Consequently, Tom, I fail to even see the relevance of the comparison.
> And of course, you are
>assuming that once a year is up, "fiddle points" will go away.
No, I am taking as literal what was passed by the EB. That is what is real.
If, and only if, it changes, then it will not be real and my opinion may
need to change.
Are you telling me that you are assuming that the non-current reality should
take priority over the current reality in terms of determining what is real?
If so, on what basis do you make such a choice?
> What is the
>USCF track record at making software changes. I remember many years ago,
>complaining when they dropped the rating floor 100 points because it was
the
>only anti-deflation part of the system (though admittedly not a very good
>one). I was told not to worry. They had designed this new system & it
>would be implemented in less than a year. Well, that's a lot longer than
a
>year ago.
>
> - Tom Martinak
The track record is horrible Tom. And based on the current track record,
the sunset provision may expire before activity points are implemented.
And actually I agree with Jerry, too, although I did not make this
explicit in previous posts. In other words, both of us are comfortable
with --- indeed prefer --- high volatility, because even old dogs like
us have dreams of glory. This is not inconsistent with wanting a
reasonably well-behaved rating system, for example one that's free of
the kind of distortion that activity points will inevitably introduce.
Best Regards,
Eric Mark
(Legitimate 1975-ish patzer, and proud of it)
In article <93kq4k$css$1...@nnrp1.deja.com>,
> More problematic though, is that you think the statement is relevant.
>
> The question is what is the higher priority for the rating system,
psychology
> or predictive ability -- not the relative psychology or relative
predictive
> ability for variants.
>
> ALL of the choices you give are more focused on predictive ability than on
> psychology. Consequently, Tom, I fail to even see the relevance of the
comparison.
But of course, predictive ability is an important part of the rating system
and an important part of the psychology of the players. The question is
what the focus of the rating system is. At the two extremes are a perfectly
predictive system in which nobody participates and a less-predictively
accurate system with maximal participation. You are trying to make the
other extreme be a totally unpredictive system, but that won't maximize
participation and I have never claimed to want that. You are simply
creating a straw man.
Here is my question to you: Tell me what multiplier of K in the new system
you think will be more destructive to the predictive ability of the rating
system if implemented for 5 years than fiddle points if implemented for 1
year. Do you think any multiplier bigger than 1, even 1.0001 will be worse?
I doubt that. That multiplier would then give us an upper limit which is
equivalent to level of predictive ability which USCF governance deems
necessary.
- Tom Martinak
Bogus extreme argument.
K = 32 has been around a long time (with K = 24 for the experts/masters).
There is consumer resistance to changing this K...even if it is for a
"mathematically good" purpose like reducing volatility.
When Tom M. or others make this consumer psychology point, your response is to
say "fine, make K = 1,000 then", which is non-responsive to their point.
Eric C. Johnson
Wrong...it has done so very well for decades.
The "problem" comes when one side (math) or the other (consumer psychology)
tries to dominate.
Right now, the math side has their panties in an uproar.
Eric C. Johnson
As long as you're bronzing predictions, don't forget to toss in the
likely scenario for Year 2, when the trickle-up effects start to become
noticeable. Eventually we'll see New York regulars like Jay Bonin (a
fine player, but not a strong GM) sporting ratings like 2700.
This is the "they'll laugh at us in Sweden" scenario, much discussed in
previous threads. Proponents of activity points said let them laugh,
we're trying to save the Federation here. Let's see how this argument
holds up a year or two from now.
In article <93l6k4$phk$1...@nnrp1.deja.com>,
Yes, but there is also a bit of fictional "accuracy" going on.
If my players are rated 1861, 1876, 1836, 1901, and 1873....then any "pairing
differences" are mostly phantoms.
Yet the rating system and Swiss rules tell everyone what the pairings ought to
be, as if the numbers were hyper-accurate.
"Fair" pairings work at a fairly lumpy/crude level....far lower than the
"expressed but false" accuracy of reporting ratings in terms of single digits.
For the players above, a better "predictive" rating would be something like
1870, 1870, 1870, 1870, and 1870. Have fun pairing without generating
complaints!
So...we see another non-math-purity reason for ratings...namely, to allow
pairings without undue bitching.
Eric C. Johnson
Partly due to power politics within USCF at the time.
Just like now.
Eric C. Johnson
Well...it doesn't.
Gravity does not exist...but gravitational behavior does. There is no separate
thing "gravity" apart from the behavior of the objects that act
"gravitationally."
Gravity itself does not exist, nor does it "pull"...so I am surprised that you
speak so loosely about such things.
Eric C. Johnson
Only because "volunteers" like Mr. Doan refused to implement them, because he
doesn't like them.
He works on the parts he likes. IMHO that is just not right.
Eric C. Johnson
> At the two extremes are a perfectly
>predictive system in which nobody participates and a less-predictively
>accurate system with maximal participation.
I completely disagree. Why are these the extreme? Why do you even think
there is a relationship, such that as predictiablity decreases, participation
increases?
The Elo system was more predictive than the Harkness, yet participantion
increased after the move from the Harkness. To even assume your thesis is
an error.
You are trying to make the
>other extreme be a totally unpredictive system, but that won't maximize
>participation and I have never claimed to want that. You are simply
>creating a straw man.
>
No, I didn't say anything about the above rather strange thesis, and I don't
know how you derived that. I reject that thesis entirely.
>Here is my question to you: Tell me what multiplier of K in the new system
>you think will be more destructive to the predictive ability of the rating
>system if implemented for 5 years than fiddle points if implemented for
1
>year.
Why is it relevant? Under either system, predictiability is more important
than psychology. I pick C, where C is any number, because it doesn't matter.
Your argument is like this. Suppose we said: What is more important to the
US Economy, technology or books. I say technology, you say books.
So now you ask me, well yeah, but which kind of technology would have more
impact on the economy, computers or cars?
And my response is, it doesn't matter Tom, they are still both impacting
the economy more than books.
I don't see your point.
> Do you think any multiplier bigger than 1, even 1.0001 will be worse?
>I doubt that. That multiplier would then give us an upper limit which is
>equivalent to level of predictive ability which USCF governance deems
>necessary.
>
No, it does no such thing. No such comparison was made or presented to the
Board, they had no opportunity to consider it, etc.
Again, these things that you are suggesting right now -- these are hypotheticals,
these are assumptions. They didn't happen.
What happened is we kept a predictive system. We augmented it temporarily
for a subset of the entire pool. ALL the evidence from those real actions
is that it was a higher priority to keep a predictive system and to limit
the unpredictable Activity Points.
Did we choose something like life master points? No. Did we keep Activity
Points indefinitely and reject the Elo system after 1 year? Nope.
Do I agree that both aspects are important? Certainly. Based on the above,
how must I prioritize the recognition of those aspects: 1. Predictive Nature
of the system and 2. Activity.
Notice my arument: All the things that DID happen point to the above, all
the hypotheticals that did not happen point otherwise.
If you wish to argue that the order is really reversed, you need to show
actual events, not hypotheticals, that show that Activity is of higher priority.
If you want to argue that for a 1 year period Activity is of a higher priority
I would say "Yes, that's real" I agree. But as a general statement you have
failed to show that, and you have failed to show that prdictive ability and
activity need to be mutually exclusive (I don't believe they are.)
Kevin