-
Training a Quadruped Robot to Walk with RL
Train a quadruped to walk with RL in Isaac Lab: 4,096 environments, terrain curriculum, domain randomization, teacher-student distillation, sim-to-real.
-
Behavior Cloning vs Reinforcement Learning: Robotics Approaches Compared
Behavior cloning vs reinforcement learning for robotics: covariate shift, DAgger, diffusion policies, AMP/GAIL, and today's practical BC-then-RL workflow.
-
Hindsight Experience Replay: Learning from Failed Attempts
Hindsight Experience Replay explained: relabel failed trajectories with achieved goals to learn from sparse rewards, HER strategies, SAC + HER training.
-
Isaac Gym Tutorial: NVIDIA's GPU-Accelerated Robot Simulator
Isaac Gym tutorial: GPU-accelerated robot simulation with thousands of parallel PyTorch environments, vectorized rewards, PPO, and domain randomization.
-
PyBullet Tutorial: Physics Simulation for Robot Learning
PyBullet tutorial for robot learning: GUI vs DIRECT modes, URDF loading, joint control, a Gymnasium reaching environment, PPO training, cameras, sim-to-real.
-
Building a Custom Gymnasium Environment from Scratch
Build a custom Gymnasium environment from scratch: action and observation spaces, reset and step, termination vs truncation, rewards, registration, and testing.
-
Train an AI to Play Pong from Pixels
Train an AI to play Pong from raw pixels with DQN: preprocessing, frame stacking, a CNN architecture, experience replay, epsilon scheduling, and tuning.
-
Train an AI to Play Flappy Bird with Reinforcement Learning
Train an AI to play Flappy Bird with DQN: environment setup, state representation, reward shaping, epsilon-greedy exploration, tuning, and evaluation.
-
Deep Q-Networks vs PPO: Which Algorithm for Which Game?
DQN or PPO for your game? A practical comparison of action spaces, sample efficiency, and stability, with clear rules for arcade, racing, and physics games.
-
Soft Actor-Critic (SAC) Tutorial for Continuous Control
Soft Actor-Critic (SAC) tutorial for continuous control: the maximum-entropy objective, twin critics, entropy tuning, hyperparameters, and SAC vs TD3 vs PPO.