Reinforcement Learning
- In Turkish
- Pekiştirmeli Öğrenme
In short
Reinforcement learning is a type of machine learning in which an agent learns to make decisions by trial and error, earning rewards for good actions.
What is reinforcement learning?
Reinforcement learning, or RL, is a way of training software to make a sequence of decisions. Instead of learning from a fixed set of correct answers, an agent interacts with an environment, takes actions, and receives rewards or penalties. Over many attempts, it learns a policy, a strategy that tells it which action to take in each situation to earn the most reward over time.
Each step follows the same loop: the agent observes the current state, picks an action, and the environment returns a new state and a reward. A key challenge is balancing exploration, trying new actions to discover better options, with exploitation, repeating actions already known to work. Classic algorithms such as Q-learning estimate how valuable each action is in each state, while deep reinforcement learning uses neural networks to handle huge state spaces like the pixels of a video game screen.
Training a dog with treats is the classic analogy: nobody explains the rules, but good behavior earns a treat, and the dog gradually repeats it. RL has been used to master board games and video games, control robots, and fine-tune large language models through RLHF, reinforcement learning from human feedback, where human ratings of answers act as the reward.
Reinforcement learning is often confused with supervised learning. In supervised learning, the model is shown the correct answer for every example, such as an image labeled 'cat', while in RL the agent is only told how good an outcome was, often long after the action that caused it. This makes RL powerful for decision-making problems but also harder to train, because a badly designed reward can teach the agent to exploit loopholes instead of solving the real task.
Key takeaways
- An RL agent learns by trial and error from rewards and penalties.
- The core loop is state, action, reward, and next state.
- Agents must balance exploring new actions with exploiting known good ones.
- Unlike supervised learning, RL has no labeled correct answers, only feedback on outcomes.
- RLHF uses human feedback as the reward to fine-tune language models.
Example
import random
# Two buttons with hidden payout rates; the agent must discover the better one
payout_rates = [0.3, 0.7]
value = [0.0, 0.0] # the agent's estimate of each button's reward
count = [0, 0]
for step in range(1000):
# Explore 10% of the time, otherwise exploit the best-known button
action = random.randrange(2) if random.random() < 0.1 else value.index(max(value))
reward = 1 if random.random() < payout_rates[action] else 0
count[action] += 1
value[action] += (reward - value[action]) / count[action] # running average
print(value) # roughly [0.3, 0.7]Readers ask
What is the difference between reinforcement learning and supervised learning?
Supervised learning trains on examples that come with the correct answer, while reinforcement learning trains an agent through rewards for its actions, without telling it the right move. RL suits decision-making tasks where the best action is not known in advance.
What is RLHF?
RLHF stands for reinforcement learning from human feedback. People rate or rank a model's answers, those preferences train a reward model, and the language model is then fine-tuned with reinforcement learning to produce answers that score higher.
What is a reward function?
A reward function is the rule that gives the agent a score after each action or episode, defining what success means. Designing it carefully is critical, because agents optimize exactly what is rewarded, even in unintended ways.
See also
- Machine LearningAI & Machine Learning, p. 27Machine learning is a branch of artificial intelligence in which computers learn patterns from data to make predictions instead of following hand-written rules.
- Neural NetworkAI & Machine Learning, p. 33A neural network is a machine learning model made of layers of connected artificial neurons that learn patterns from data by adjusting numeric weights.
- Deep LearningAI & Machine Learning, p. 14Deep learning is a subset of machine learning that uses neural networks with many layers to learn complex patterns from raw data such as images and text.
- Fine-tuningAI & Machine Learning, p. 19Fine-tuning is the process of taking a pretrained machine learning model and training it further on a smaller, specific dataset to adapt it to one task.
- AI AgentAI & Machine Learning, p. 2An AI agent is a system that uses an LLM to plan and carry out multi-step tasks by deciding which tools to call, observing the results, and acting again.
- LLMAI & Machine Learning, p. 25An LLM is a machine learning model trained on huge amounts of text that generates language by repeatedly predicting the next most likely piece of text.
- RLHFAI & Machine Learning, p. 41RLHF (reinforcement learning from human feedback) trains a language model to be more helpful and safe using people's judgments of which answers are better.
Spotted a mistake or something missing on this page?Suggest an edit