[ad_1]
The buzz of a notification or the jingle of an email could inspire excitement or dread. In a famous experiment, Ivan Pavlov (pictured) demonstrated that dogs can be taught to salivate to the ticking of a metronome or the sound of a harmonium. This cause-and-effect connection, known as associative or reinforcement learning, is central to how most animals deal with the world.
Since the early 1970s the dominant theory of what is happening has been that animals learn by trial and error. The association of a signal (a metronome) with a reward (food) takes place as follows. When a cue comes, the animal predicts when the reward will occur. So, wait to see what comes along. Next, calculate the difference between the prediction and the error result. Finally, he uses that error estimate to update things to make better predictions in the future.
Belief in this approach was itself strengthened in the late 20th century by two things. One of them was the discovery that he is also good at solving artificial intelligence (AI) related engineering problems. Deep neural networks learn by minimizing the error in their predictions.
The other reinforcing observation was an article published in Science in 1997. He noted that fluctuations in the brain’s levels of dopamine, a chemical that carries signals between certain nerve cells and was known to be associated with the experience of reward, they looked like prediction-error signals. Dopamine-generating cells are most active when the reward arrives earlier than expected or not expected at all, and are inhibited when the reward arrives later or not at all, exactly what would happen if they were indeed such signals.
A good story, then, of how science works. But if a new paper, also published in Science, turns out to be correct, it is wrong.
Researchers have long known that some aspects of dopamine activity are inconsistent with the prediction error model. But, in part because it works so well for training artificial agents, these problems have been swept under the rug. Until now. The new study, by Huijeong Jeong and Vijay Namboodiri of the University of California, San Francisco, and a team of collaborators, has turned the world of neuroscience upside down. He proposes an associative learning model that suggests that researchers have things backwards. Furthermore, their suggestion is supported by a series of experiments.
The old model looks forward, associating cause with effect. The new one does the opposite. Match the effect to the cause. They think that when an animal receives a reward (or punishment), it looks back through its memory to figure out what might have triggered this event. The role of dopamine in the model is to flag events that are significant enough to serve as causes for possible future rewards or punishments.
Looking at it that way these are two things that have always bugged the older model. One is time scale sensitivity. The other is computational tractability.
The problem with the time scale is that cause and effect can be separated by milliseconds (turning on a light bulb and experiencing enlightenment), minutes (drinking and feeling high), or even hours (eating something bad and getting food poisoning). Looking back, explains Dr. Namboodiri, an arbitrarily long list of possible causes can be investigated. Looking ahead, without always knowing in advance how far to look, is much more complicated.
This leads to the second problem. Sensory experience is rich, and anything in it could potentially predict an outcome. Making predictions based on every single possible clue would be somewhere between the difficult and the impossible. It’s much easier, when a significant event occurs, to look back through other potentially significant events to a cause.
In practice, however, it is difficult to distinguish experimentally between the two models. And that’s especially true if you don’t even bother looking at which, until now, people haven’t. Dr Jeong and Dr Namboodiri did it. They devised and conducted 11 experiments involving mice, buzzers and drops of sugar solution specially designed for the purpose. During these they measured, in real time, the amount of dopamine released by the nucleus accumbens, a brain region where dopamine is involved in learning and addiction. All experiments have gone in favor of the new model.
The turnaround in prospective to retrospective thinking implied by these experiments is causing quite a stir in the neuroscience world. It’s inspiring and represents an exciting new direction, says Ilana Witten, a neuroscientist at Princeton University not involved in the paper.
More experiments will be needed to confirm the new findings. But if confirmation comes, it will have ramifications beyond neuroscience. It will suggest that the way AI works has, as currently claimed, not even a tenuous link to the way brains work, but it was actually a lucky guess.
But it could also suggest better ways to do AI. Dr. Namboodiri thinks so and is exploring the possibilities. Evolution has had hundreds of millions of years to optimize the learning process. So learning from nature is rarely a bad idea.
|
Sources 2/ https://www.economist.com/science-and-technology/2023/01/18/a-decades-old-model-of-animal-and-human-learning-is-under-fire The mention sources can contact us to remove/changing this article |
[ad_2]