Comment
Nicole Junkermann on Reinforcement Learning and the Value of Patience
Learning by trying, and being told afterwards how it went. Nicole Junkermann on reinforcement learning, and the patience the method quietly requires.
Back to the blog
Reinforcement learning is one of those phrases that sounds far more forbidding than the idea underneath it. Take the vocabulary away and what is left is simply this: learning by trying something, receiving a signal about how it went, and trying again with that signal in hand.
Anyone who has taught a child to catch a ball already knows the shape of it. No instruction works. There is a throw, a miss, a small adjustment that nobody involved could put into words, and another throw. The correction does not arrive as an explanation. It arrives as an outcome. Do that a few hundred times and something that cannot be written down gets learned anyway.
Machines can be set up to learn in roughly that way. Rather than being shown the right answer, a system is put in a situation, allowed to act, and then given a score. A better score suggests the behaviour is worth repeating. A worse one suggests it is not. Run that loop often enough and behaviour nobody wrote down begins to appear. Nicole Junkermann finds it the most human sounding corner of the whole field, which is probably why it is the one most often described badly.
What gets lost in the retelling is how much of it is failure. The overwhelming majority of attempts are wrong, and they have to be, because a system that only ever repeats what already works learns nothing new. Early on the results look worse than useless. There is no way to skip that stretch and no reliable way to tell, from inside it, whether the approach is sound. The only honest answer is more attempts.
Then there is the problem of delay. The score often arrives long after the decisions that earned it, and the credit has to be spread backwards across a great many small choices, most of which were neither good nor bad on their own. Working out which parts of a long sequence deserve the credit is genuinely hard. It is hard for people too. Anyone who has tried to decide which of a year's decisions produced a good year has met exactly the same difficulty.
That is where patience stops being a decorative word and starts being a requirement. Learning by trial needs a tolerance for a long run of poor results, an honest scorecard, and a willingness to keep the scorecard rather than quietly rewrite it once the answer is known. Nicole Junkermann has written before about why the best ideas are worth the wait, and the resemblance is not a coincidence. Both describe the same bargain: accept a poor near term in return for something that could not have been reached any other way.
None of which says what such systems will be able to do, or when. It is a description of a method, not a prediction, and it belongs on the shelf next to the other things worth understanding slowly in the Reading Room. What the method teaches is older than any of it. A run of failures is sometimes the only available route, and the people who can sit calmly with that are rare.
This article is a personal reflection and general commentary. It is not financial, investment or professional advice.
More from Nicole Junkermann
For more from Nicole Junkermann, read the full profile of Nicole Junkermann, or go back to what India's AI moment signals for long term investors.