> ... How are you going to
> explain the effects of a single temporal conjunction of response and event
> or event and event, and at the same time explain why response-independent
> reinforcement does not generally maintain responding? Do you intend to
> make time part of the context in P(response/context)?
I'm thinking along these lines. If two events A and B occur at times Ta and
Tb that are "close", then there are 4 hypotheses:
1) A has a causal influence over B
2) B has a causal influence over A
3) A and B have a common cause (there is a variable C, or a set of variables
S, with causal influence over A and B)
4) The contiguity is coincidental.
I say causal "influence" to allow for the influence to be probabilistic.
The basic idea is simple - "suspicious coincidences" are probably not just
coincidence. How to quantify that? Still conceptually simple: if events A
and B are "rare" then even a single co-occurrence boosts the likelihood of
a causal connection. But if the events are "common" then co-occurrence is
not "surprising" (or "suspicious"), though there still may or may not be a
causal relationship.
Of course, between the basic idea and probability models there is much work.
For laboratory settings that might not be hard - I'll want to read more on
the experimental work. The little I glanced at involved establishing stable
responding rates under, say, a VI schedule, and then changing the schedule
so that some or all reinforcer delivery events were independent of response
events. So at the start of the second phase the prior over the 4
possibilities would strongly favor the hypothesis that response events have
causal influence over reinforcer delivery events. Looking at the whole
experiment including phase 1 and phase 2, the distribution
P(reinforcer|time,response) is a "piecewise stationary stochastic process"
with two states and a single jump transition. Ilya Nemenman points out that
during the transient phase after a state transition the different classes
of Bayesian probability estimators undergo distinct fluctuations
characteristic of the class ("Fluctuation-Dissipation Theorem and Models of
Learning").
So if the behavior of an organism approximates a Baysian probability
estimator then the fluctuations in response rate after the start of phase 2
are diagnostic of what kind of Bayesian estimator it approximates. Of
course the behavior would be a function not only of the probability
estimate but of cost and benefits as well. But, assuming that these are
relatively constant over the course of the experiment, response rate should
track probability estimate.
This sort of piecewise stationary probability distribution occurs in natural
environments, for instance when a fly flies from a meadow into the woods,
or when it makes an abrupt transition from "cruising flight" to "chasing
flight", although here the distribution is over the statistics of the
stimuli that, uh, "guide" flight behavior (e.g. variance and mean of the
horizontal angular velocity of the wide-field background, variance and mean
of luminance contrast, etc.).
So yes, I intend to make time part of the context. It's not clear to me just
how to do that, but I am thinking in terms of "multi-resolution analysis"
to track patterns at multiple scales.
The case of conditioned taste aversion poses, at least for me, difficulties
since the context (including taste, smell, color. etc) and response (eating
or drinking), and the consequence of interest (nausea) are not "close" in
time. One obvious thought is that there is a "built-in strong prior". Does
that have to be assumed ad hoc, or is there a more principled way to derive
it?
> ... Or will temporal
> issues enter in, if at all, in other ways? For example, maybe the issues
> are handled by your added value and response cost notions - issues that
> are outside, it seems to me, of the strictly Bayesian calculations. For
> example, if a response is regarded as an attempt to evaluate a hypothesis
> concerning models of P(event/R1) and P(event/ no R1) and P(event/R2) etc.,
> you can say that a single temporal conjunction changes the estimate of the
> probability that the model is correct by only a little, but since the
> response is not very costly, you might as well run with the notion that
> the response makes the event fairly likely. But even here, your model
> would have to somehow incorporate the fact that increases in rate of
> response following a "pairing" is a probably a function of the precise
>temporal relation (i.e., arranging an FR 1 and a Tand FR 1 FT 3.0 s would
>likely have different effects).
That's correct - value and cost issues are outside the strictly Bayesian
calculations. The full problem - what to do, not just what are the odds
over outcomes given context and actions - is "Bayesian Decision Theory".
The formal statement is to find the response that maximizes the expression:
Gain(consequence) * P(consequence|context,response) - Cost(response)
For a simple model, with consequences limited to reinforcer events RE, with
some constant gain g, and with zero cost or gain per no-reinforcer "events"
(so we can leave them out, and just focus on the probability of RE), and a
constant cost c per response, then maximize:
g*P(RE|context,response) - c
The "break even point" is:
P(RE|context,response) = c/g
So yes, if the ratio of costs to gains is very low, then even a very low
estimated probability of a response resulting in a reinforcer event
justifies a response. This will depend on the precise temporal relations
because they will affect the probability distribution as function of the
time component of context.
But we are not looking for a break even point, we are looking for an optimal
strategy of responding over the course of an experiment. What makes this
difficult to analyze (i.e. to answer the question what would Bayes do) is
that there is no single optimal "response vector" because the optimal next
response (even under steady state conditions) will depend on the outcomes
of previous reponses. This is a problem in the domain of "sequential
decision theory" (a special case within decision theory), so there is some
literature on the abstract theory available.
Promising Papers
================
Yesterday I came across three papers that look promising:
http://monkeybiz.stanford.edu/pubs.html
Sugrue, LP, Corrado, GS and Newsome, WT (2004). Matching behavior and the
representation of value in the parietal cortex. Science 304:1782-1787
Corrado, G.S., Sugrue, L.P., Seung, H.S. and Newsome, W.T. (2005).
Linear-nonlinear-poisson models of primate choice dynamics. J. Experimental
Analysis of Behavior 84:581-617.
Sugrue, LP, Corrado, GS and Newsome, WT (2005). Choosing the greater of two
goods: neural currencies for valuation and decision making. Nature Reviews
Neuroscience, pgs. 1-13.
I've only glanced at them so far, but it is the middle paper, obviously,
that goes into the most depth. It also makes key improvements over the
model in the first paper. The authors are interested in neural correlates
of choice behavior, but what they do is in the spirit of what I had in mind
with "what Bayes can do for the EAB". There is a difference - I have am
thinking in terms of deriving an optimal Bayesian agent from first
principles and comparing the behavior of the agent to the behavior of
organisms. Since I take the main goal of AI to be general principles this
remains my approach to thinking about the problem.
But the authors, neuroscientists, quite naturaly have a different goal and
take a different approach:
1) Based on extensive previous work within the EAB on matching law-like
choice behavior they design a behavioral study. Since they are interested
in a mathematical model of choice behavior, and since transitions are
better than stable states at discriminating mechanisms, their experiment
involves frequent changes in probability distributions - dynamic concurrent
VI schedules.
2) They specify a model space - a parametric family of mathematical models
of choice behavior. This becomes a model selection problem over a model
space. They use both time-series analysis and maximum likelihood estimation
to arrive at a "best fit" model - i.e. a most probable explanation of the
behavioral data from the space of explanations they explore.
3) They test the validity of the model by both its ability to predict the
choices the monkeys make, given past choices, and its ability to reproduce
the choice behavior by acting as an agent making choices that have
outcomes. Obviously they got a good match.
4) They compare the specific model that describes the monkeys' behavior to
the model from within their chosen space that optimizes harvesting rewards.
The discriptive model of monkey behavior is very close to optimal (the
normative model within the space).
5) They do a sensitivity analysis of the dependency of optimality on the
parameters of their model and find that while some parameters are critical
(and are tightly matched by the descriptive model) others have little
impact on optimality (and these are where the descriptive models vary
somewhat from the optimal model).
6) They go on (though they only mention this in passing in the JEAB paper)
to examine neural correlates of the parameters of the mathematical model:
"While a literal implementation of our model is one way that the brain could
produce matching behavior, there are alternative models that might describe
the behavioral data well but have very diferent internal mechanisms. Thus
the success of our behavioral model *does* *not* necessarily imply that its
components will have direct nerual correlates. Therefore, we must be
cautious in making the leap from behavioral description to neural
mechanism. [but] ... the model presented here provides a better
account of both the behavioral *and* the neural data."
So does the EAB have anything to offer AI? Well, if one subgoal of AI is to
duplicate primate behavior ("foraging behavior" in this case) then given
the mathematical model, derived exclusively from careful analysis of
behavior, (without "opening up the box") this is all you need - you can
write an algorithm that implements the model. In fact, this was how the
behavior of the model was compared to the behavior of the monkeys.
Need more? Consider this:
"Thus, in the context of matching behavior, experimental data from several
laboratories seem at odds with the prevailing paradigm in the field of
reinforcement learning in which return is the dominant driver of choice
behavior (Sutton & Barto, 1998)".
The reference is to Sutton and Barto's introductory text on "reinforcement
learning", one of the major paradigms of (and a subdiscipline within) the
field of "machine learning".
A note on "dynamic concurrent VI schedules"
===========================================
Two responses "harvested rewards" according to two variable interval
schedules - that is, a reinforcer (juice) became available independently at
each of two "locations" after an independent random variable amount of time
had elasped. So a reward could be harvested at location x if and only if
location x was "baited", which happened randomly around some mean elapsed
time after the last harvest event at that location. In general the mean time
to next reinforcer availability at the two locations was different (though
it could be the same). The ratio of mean time one to mean time two changed
abruptly and randomly (uniform distribution) between 9 discrete cases {1:8,
1:6, 1:3, 1:2, 1:1, 2:1, 3:1, 6:1, 8:1}, though not all cases occurred on
every session, so on any given session some subset of the cases, including
the full set, occurred. This is a 9-state piecewise stationary process. The
transition between states occurred anywhere from after 50 to 300 trials.
I'm assuming, though the authors do not state (as far as I know, I've only
glanced at the paper) that state transitions were drawn uniformly over the
integers in [50, 300].
This probability distribution of reinforcer availability is well
characterized by a 9-state Hidden markov model (HMM). In general that would
mean on the order of 81 or so independent parameters. But there is much
symmetry here that vastly reduces the parameter space. First, the state
transitions reduce to two cases - same state, other state, since given that
a state transition has occurred all "other states" are equally likely. And
since the probability of same state plus the probability of other state
must sum to one, the 72 independent state transition probabilities collapse
to one number, the probability of staying in the same state. We need a
probability such that state residence time varies uniformly between 50 and
300 trials. That is not possible. Any number we choose will result in a
geometric rather than uniform distribution over state residence times. So a
9-state HMM only approximates the distribution. We can make an HMM that
models the distribution exactly if we let it have 300 * 9 = 2700 states.
Still, there would be much redundancy among states. They are tracking two
things: 1 of 9 mean time ratios, and the number of responses since last
state change, which determines the probability of the transition from
current state to other state.
That's one way to model it. Another, more economical way is to relax the
strict state-space approach, and let each of the 9 states of interest have
a counter that systematically and algorithmically alters state transition
probabilities. Now each state has a transition probability *function*
rather than a transition probability *number*. So this function is just:
P(transition back to self) = 1 - Max(0, (nTrials - 50))*1/250
I think that does the trick. The state transition probabilities of the
stochastic process have been reduced to zero degrees of freedom. That
leaves the parameters for each of the 9 states.
In this experiment the time from last harvest at location x to next bait
availability follows a "Poisson distribution". Well, that's not quite
right. The number of bait availability events at each location within any
time interval of length T is Poisson distributed. That means that the time
between last harvest and next bait avalability event at each location is
exponentialy distributed. The exponential distribution is characterized by
a single number - its mean inter-event time m:
P(bait available) = (1/m)*e^-((1/m)*(time since last harvest))
There are two such exponential distributions for each of the 9 states, but
since the ratio of the means is fixed for each state, each state is fully
characterized by a single parameter.
But there is a further constraint. The authors state that the sum of the
reward baiting probabilities was held constant at 0.12 rewards per second.
So for each state we have not only m1/m2 = aKnownConstant, but also that
1/m1 + 1/m2 = 1/0.12
(The sum of two exponential distributions of "rates" r1 and r2 is an
exponential distribution of reate r1 + r2. The rate of an exopnential
distribution is the inverse of its mean 1/m.)
So with these two equations we can solve for the values of m1 and m2 for
each state. The stochastic process, the response of the environment, is
fully specified.
A note on the descriptive model space.
=====================================
The model space, "Linear-nonlinear-poisson models" consists of 3 serial
feedforward computational stages:
1) a linear stage - a linear filter of recent reward history
2) a nonlinear function that maps the output of the first stage onto a
probability
3) a poisson process that draws a binary outcome with this probability
So it is the parameters of this model that are estimated from the behavioral
data. For the rational behind the model and how it is parameterized see the
paper. Its ultimate justification, of course, is its ability to predict and
reproduce the behavior it is supposed to model.
-- Michael
> But there is a further constraint. The authors state that the sum of the
> reward baiting probabilities was held constant at 0.12 rewards per second.
> So for each state we have not only m1/m2 = aKnownConstant, but also that
> 1/m1 + 1/m2 = 1/0.12
Oops. if the average 0.12 rewards per second then the mean time between
rewards is about 8.333 seconds. So the equation should have been:
1/m1 + 1/m2 = 1/8.333
> (The sum of two exponential distributions of "rates" r1 and r2 is an
> exponential distribution of reate r1 + r2. The rate of an exopnential
> distribution is the inverse of its mean 1/m.)
-- Michael
>
>Glen M. Sizemore wrote (among other things):
>
>> ... How are you going to
>> explain the effects of a single temporal conjunction of response and event
>> or event and event, and at the same time explain why response-independent
>> reinforcement does not generally maintain responding? Do you intend to
>> make time part of the context in P(response/context)?
>
>I'm thinking along these lines. If two events A and B occur at times Ta and
>Tb that are "close", then there are 4 hypotheses:
>
>1) A has a causal influence over B
>2) B has a causal influence over A
>3) A and B have a common cause (there is a variable C, or a set of variables
>S, with causal influence over A and B)
>4) The contiguity is coincidental.
>
>I say causal "influence" to allow for the influence to be probabilistic.
>
>The basic idea is simple - "suspicious coincidences" are probably not just
>coincidence. How to quantify that? Still conceptually simple: if events A
>and B are "rare" then even a single co-occurrence boosts the likelihood of
>a causal connection. But if the events are "common" then co-occurrence is
>not "surprising" (or "suspicious"), though there still may or may not be a
>causal relationship.
I would not use the term 'coincidental'. 'Temporally correlated' is a
more accurate choice, IMO. A temporal correlation may or may not be
causal but who cares? What is important in learning are that temporal
correlations can be found because they are necessarily probabilistic.
There can be only two types of temporal correlations: signals are
either concurrent or sequential. Whether or not a sequential
correlation is causal makes no difference to the neural learning
mechanism. There is another type of temporal correlation which is,
strictly speaking, not another: Different temporal intervals may also
be correlated such that halving one interval results in one or more
other intervals being halved as well. This is essential to memory
formation and the phenomenon known as pattern completion.
Louis Savain
Why Software Is Bad and What We Can Do to Fix It:
http://www.rebelscience.org/Cosas/Reliability.htm
>
> So it is the parameters of this model that are estimated from the behavioral
> data. For the rational behind the model and how it is parameterized see the
> paper. Its ultimate justification, of course, is its ability to predict and
> reproduce the behavior it is supposed to model.
>
I think you've answered your own question. Most people interested in AI
aren't interested in "predicting behavior", nor I doubt wish to
"duplicate
primate behavior". They want to build intelligent machines, not
duplicate
what monkeys do. All you've really done is run around the same old
block,
and arrived at where you started from, or rather, where GS+DL are
standing.
However, good to see, as GS pointed out to you recently on another
thread, that you now realize that your bayesian models are really just
behaviorist formulations.
Parroting your advice to Ray in a recent thread, maybe you should take
all this to some psychology forum.
"Michael Olea" <ol...@sbcglobal.net> wrote in message
news:1RXhg.47334$Lm5....@newssvr12.news.prodigy.com...
>
Still don't you think it is nice that someone apart from
Curt and Wolf are taking GS seriously? And in this case
has the mathematical background to analyze what EAB has
to offer AI? As far as I can tell MO has been able to make
money writing GOFAI programs using his knowledge of
statistics and its application to input data and thus
I think has the creds? And he has managed to give Glen
something to chew on for a while without having an
immediate come back which I think is a good effort.
Maybe after duplicating the behavior of a monkey he might
tackle the problem of duplicating the behavior of Microsoft
Windows after a careful analysis of its behavior. I think
there was a project called ReactOS that tried to do this.
Might be easier than the monkey project?
MO's reference to taste aversive conditioning reminded me of
the experiments where they tried to associate a mild electric
shock to the chick's feet with pecking. The result was the
chicks pecked even more vigorously and somewhat aggressively.
I wonder if this effect can be found in humans when someone
continually gives them a hard time?
--
JC
> feedbackdroid wrote:
>> Parroting your advice to Ray in a recent thread, maybe
>> you should take all this to some psychology forum.
Just for the record, in case anyone takes this at face value, I gave Ray no
such advice. Quote:
(any guesses why someone supposedly so interested in neuroscience never
posts on bionet.neuroscience?).
Granted that it is an unfortunate remark I would have done well to leave
out, but the leap from that comment to "take all this to some other forum"
is ... impressive.
> Still don't you think it is nice that someone apart from
> Curt and Wolf are taking GS seriously?
Glen has technical expertise in a field that cannot help but be germain to
at least some aspects of AI. You, John, used a phrase I have often used
myself:
"... making machines do things that if done by a human would be considered
"intelligent" behavior."
> And in this case
> has the mathematical background to analyze what EAB has
> to offer AI?
And vice versa, e.g: "what Bayes can do for the EAB".
> As far as I can tell MO has been able to make
> money writing GOFAI programs using his knowledge of
> statistics and its application to input data and thus
> I think has the creds?
Thanks.
This may shed some light on a remark you made earlier that I found strange -
something about the most "intelligent" machines/apps so far are products of
GOFAI. Most uses of the term GOFAI that I have seen restrict it to forms of
symbol manipulation, in direct and explicit contrast to, for example,
"connectionism" (e.g ANNs). So, for example, the world-class artificial
backgammon player TDGammon is not a symbol manipulator, but a product of
RL; and the winner of the DARPA grand Challenge is a Bayesian Inference
Engine.
The money making scene analysis code I've had a big role in (but no bigger
than that of two biz partners) is eclectic in approach. But not much of it
could be called GOFAI by the usage with which I am familiar.
> And he has managed to give Glen
> something to chew on for a while without having an
> immediate come back which I think is a good effort.
We started this discussion, I think, about 3 years ago. It didn't get very
far at the time.
> Maybe after duplicating the behavior of a monkey he might
> tackle the problem of duplicating the behavior of Microsoft
> Windows after a careful analysis of its behavior. I think
> there was a project called ReactOS that tried to do this.
> Might be easier than the monkey project?
I am not at all fond of Windows - war stories galore. But my buddies and I
did have to do something sort of like that. We had to make freeBSD look
like Windows to the operators..
> MO's reference to taste aversive conditioning reminded me of
> the experiments where they tried to associate a mild electric
> shock to the chick's feet with pecking. The result was the
> chicks pecked even more vigorously and somewhat aggressively.
> I wonder if this effect can be found in humans when someone
> continually gives them a hard time?
Well, there are some posters who seem to act like compulsive peckers.
-- Michael
> I have to chew on some of this for a while, Michael.
A job for a CPG? I'll be interested in what you think.
-- Michael
> ... Most uses of the term GOFAI that I have seen
> restrict it to forms of symbol manipulation, in direct
> and explicit contrast to, for example, "connectionism"
> (e.g ANNs). So, for example, the world-class artificial
> backgammon player TDGammon is not a symbol manipulator,
> but a product of RL; and the winner of the DARPA grand
> Challenge is a Bayesian Inference Engine.
Temporal difference learning has not lead to the same
performance in other applications or games that it did
with TD-Gammon.
Jordan B. Pollack & Alan D. Blair had some interesting
comments as to why TD_Gammon worked.
www.cse.unsw.edu.au/~blair/nips_hcgam.pdf
--
JC
"Michael Olea" <ol...@sbcglobal.net> wrote in message
news:1RXhg.47334$Lm5....@newssvr12.news.prodigy.com...
>
> Do you want some data, Michael?
Sure. I won't be able to do much with it for a while - I'll certainly look
it over and give it at least some thought, but it would be an "after hours"
thing, and I have a couple of those already. Still, it would give me a
better idea how to formulate a model.
-- Michael