(1) An Introduction to Reinforcement Learning
This is the first review post in my series on Reinforcement Learning with Python and Keras.
http://www.yes24.com/Product/Goods/44136413
Almost a month of the vacation has already passed... 😥
To put an end to a month of doing nothing, I visited a bookstore and picked up a book on reinforcement learning. I originally intended to try CS234, but the uncomfortable truth is that my limited English and well-recharged laziness kept me from even finishing the first lecture. So I decided to read this book carefully, code and all. Apparently, I am the type who cannot find it interesting without code. Anyway, I sincerely hope I make it to the end of this one. I paid the full 28,000 won at Kyobo Bookstore!
Chapter 1. An Introduction to Reinforcement Learning
The book has the following goals:
- Understand reinforcement learning through minimal equations and intuitive illustrations.
- Implement reinforcement learning theory in simple games.
Reinforcement learning has roots in behavioral psychology and machine learning. To study it, we must first define the problems it seeks to solve. Unlike other areas of machine learning, it deals with problems requiring sequential decisions. For a computer to solve these problems, we must define them mathematically.
Skinner's research on reinforcement
Reinforcement is one way animals learn through trial and error. In behavioral psychology, trial-and-error learning means trying different things and learning from their consequences.
- A hungry rat is placed in a box.
- While exploring, it happens to press a lever inside the box.
- Pressing the lever releases food.
- Not yet understanding the connection between the lever and the food, the rat continues exploring.
- When it accidentally presses the lever again, it begins to recognize the relationship and gradually presses it more often.
- Through repetition, the rat learns that pressing the lever gives it food.
Reinforcement therefore means learning, through direct experience rather than prior instruction, the relationship between an action and a desirable reward.
Machine learning and reinforcement learning
Defining reinforcement learning requires understanding not only reinforcement in behavioral psychology, but also machine learning, which falls broadly into three categories.

지도학습(Supervised Learning): Regression and classification
비지도학습(Unsupervised Learning): Clustering
강화학습(Reinforcement Learning)
Reinforcement learning differs from supervised and unsupervised learning: it is given neither explicit answers nor merely a fixed dataset to learn from. It learns through rewards. A reward is the environment's response to the computer's chosen action. It is not a direct answer, but serves as an indirect one.
Like reinforcement in behavioral psychology, the computer learns to perform reward-producing actions more often.
The agent: A computer that learns for itself
We will call a computer that learns through reinforcement learning an 에이전트(Agent). It learns without prior knowledge of the environment. After observing its state, it acts. The environment then provides a reward and the next state, as illustrated below. By learning from the rewards resulting from its actions, the agent discovers which actions lead to good outcomes.

The goal is to learn a 최적의 행동양식, 또는 정책 that maximizes the sum of rewards obtained while exploring the environment.
Advantages of reinforcement learning
What are its advantages?
'Learning without prior knowledge of the environment.'
In the real world, an agent might normally need information about many situations to learn a skill. Reinforcement learning lets it learn through trial and error without that information. Viewed this way, AlphaGo learns by playing Go rather than starting with knowledge about how to play well. It initially places stones randomly, occasionally wins, receives a reward, and tries to repeat actions that helped it win. (In reality, AlphaGo also has a supervised-learning stage using human game records; the book leaves that part out.)
Sequential decision-making problems
Reinforcement learning resembles people learning through interaction with their surroundings. But like other machine learning methods, it can produce unintended results if we do not understand the problem. What kinds of problems should it address?
'Reinforcement learning applies to problems requiring sequential decisions.'
A sequential decision-making problem requires repeated choices rather than a single action from the current position, as in the game below.

For the agent to learn and improve, we must represent the problem mathematically. The framework used to define sequential decision-making problems is MDP(Markov Decision Process). An MDP gives the agent a mathematical formulation it can work with.
Components of a sequential decision-making problem
The mathematical formulation contains the following components. Together they form an MDP, which Chapter 2 discusses in detail.
State
The agent's state includes dynamic information, such as movement speed, as well as static information. For a table-tennis agent, for example, it should include the ball's position, velocity, and acceleration.
Action
Actions are choices available in a state, such as up, down, left, and right. In a game, they correspond to controller inputs. An untrained agent has no information about which actions are good. Through learning, it increases the probability of selecting certain actions. After an action, the environment supplies a reward and the next state.
Reward
Reward is the central element distinguishing reinforcement learning from other machine learning methods. It is effectively the learning signal available to the agent. The goal is to find a policy maximizing total reward over time. Rewards belong to the environment, not the agent; the agent does not know in advance how much reward a situation will produce.
Policy
The answer to a sequential decision-making problem is a 정책. To obtain rewards, the agent must know which actions to take not just in one state, but in every state. A policy specifies those choices.
Solving the problem means obtaining the best policy, called the 최적정책(optimal policy). Following it allows the agent to maximize its total reward.
An example: Breakout
The book teaches an agent to play several simple games. Let us look at how reinforcement learning applies to the final one, Breakout.

How do we apply reinforcement learning to Breakout? This is equivalent to asking how to formulate Breakout's MDP. We must also consider how the agent should learn.
- MDP
상태: The game screen. Four frames, like those above, are supplied as the state. Since the frames are grayscale, each consists of two-dimensional pixel data.
행동: Stay still, move left, move right, or launch. Launching is available only at the start.
보상: Breaking a brick yields +1, with larger rewards for bricks higher up. Breaking nothing gives 0; missing the ball and losing a life gives -1.
- Learning
- The agent receives four consecutive game frames as input.
- Initially, it knows nothing and acts randomly.
- When it receives a reward, it learns from that reward.
- Eventually, it plays as well as—or better than—a human.

What reinforcement learning trains here is an artificial neural network, discussed in Chapter 5. It takes four consecutive game frames as input and outputs how good each available action is in that state. This action value is called the 큐함수(Q Function).
The neural network used here is a DQN(Deep Q-Network). Given a state, DQN outputs Q-values for staying still, moving left, and moving right. The agent selects an action with a high value. The environment then supplies a reward and the next state. Through these interactions, the agent gradually adjusts DQN to obtain more rewards.
This is how the agent learns Breakout. The detailed theory and code appear later. For now, the goal is to understand the overall flow of learning.
Differences between humans and reinforcement learning agents
An agent learning Breakout resembles a human learner in some respects: both observe the game screen and learn from it.
But there are differences. The agent knows nothing about the game's rules. A person playing Breakout for the first time would probably learn the rules before starting, then deliberately try to raise their score. Learning without knowing the rules is both an advantage of reinforcement learning and a reason its early learning is slow.
Would learning be faster if a skilled friend taught us? People also transfer learning between domains: studying mathematics, for example, makes science easier to learn. Current reinforcement learning agents, however, treat each task separately and must always learn from scratch.
These differences between humans and reinforcement learning agents remain challenges for current and future research.
Chapter 1 in one sentence
Writing blog posts is seriously hard... I need practice.