Samuel Jarnet
Home
Projects
Project Interface
Trading Interface
Self Landing Rocket
Boids Simulator
Motion Detector
MP3 Player
RL Lab
About
◈ REINFORCEMENT LEARNING LAB
Multi-Armed Bandits
Gridworld
ALGORITHM
ε-Greedy (constant α)
UCB1
Thompson Sampling
Contextual (Thompson)
PARAMETERS
Episodes
Trials (avg)
▶ RUN
Estimated Q-value vs. True Reward Probability
Context: Sunny
Context: Cloudy
Context: Rainy
VARIANT
Q-Learning (5x5 Maze)
Value Iteration (5x5 Maze)
Continuous (no walls)
Continuous + Wall
PARAMETERS
Episodes
▶ RUN
Start
Goal
Wall
Configure and run a variant above