Guided series
Reinforcement Learning Fundamentals
Start from zero and work up to modern policy-gradient methods: environments, DQN, PPO and SAC, with playable game projects at every step.
12 articles · in reading order
-
Reinforcement Learning for Games and Robotics: Complete Beginner's Guide
Learn reinforcement learning for games and robotics from scratch. Understand RL fundamentals, key algorithms like PPO and DQN, and how to start training agents today.
September 3, 2026 -
Introduction to Gymnasium: OpenAI's RL Environment Toolkit (2026)
Learn Gymnasium's core API, environment spaces, wrappers, and vectorized training. The complete guide to the standard RL environment toolkit maintained by the Farama Foundation.
September 4, 2026 -
Train an AI to Play CartPole with OpenAI Gymnasium (2026): Step-by-Step Tutorial
Build a DQN agent from scratch that learns to balance a pole using OpenAI Gymnasium and PyTorch. Complete CartPole tutorial with full code and explanations.
September 3, 2026 -
Build a Snake Game AI with Deep Q-Learning (2026): Complete Tutorial
Build a Snake-playing AI from scratch using DQN and PyTorch. Learn state design, reward shaping, and training loops with full code and explanations.
September 3, 2026 -
Train an AI to Play Atari Games with Deep RL (2026): Complete Guide
Learn to train a DQN agent on Atari games using raw pixels, convolutional networks, and frame stacking. Complete guide with Stable-Baselines3 and from-scratch implementations.
September 3, 2026 -
Deep Q-Networks vs PPO: Which Algorithm for Which Game?
DQN or PPO for your game? A practical comparison of action spaces, sample efficiency, and stability, with clear rules for arcade, racing, and physics games.
September 25, 2026 -
PPO Explained: Train Robust Game AI with Proximal Policy Optimization
PPO explained for game AI: how proximal policy optimization clipping works, how PPO compares to DQN and SAC, plus hyperparameters, rewards, and training tools.
September 25, 2026 -
Soft Actor-Critic (SAC) Tutorial for Continuous Control
Soft Actor-Critic (SAC) tutorial for continuous control: the maximum-entropy objective, twin critics, entropy tuning, hyperparameters, and SAC vs TD3 vs PPO.
September 25, 2026 -
Train an AI to Play Flappy Bird with Reinforcement Learning
Train an AI to play Flappy Bird with DQN: environment setup, state representation, reward shaping, epsilon-greedy exploration, tuning, and evaluation.
September 25, 2026 -
Train an AI to Play Pong from Pixels
Train an AI to play Pong from raw pixels with DQN: preprocessing, frame stacking, a CNN architecture, experience replay, epsilon scheduling, and tuning.
September 25, 2026 -
Building a Custom Gymnasium Environment from Scratch
Build a custom Gymnasium environment from scratch: action and observation spaces, reset and step, termination vs truncation, rewards, registration, and testing.
September 25, 2026 -
Evaluating Reinforcement Learning Algorithms: Metrics and Benchmarks
Introduction to RL Evaluation Evaluating reinforcement learning (RL) algorithms is not as straightforward as supervised learning where we can simply calculate accuracy or mean squared error. In RL, we’re dealing with agents learning…
January 28, 2026
Want more paths? Browse all series or pick a topic.