🎢 RL vs LQR
← Portfolio
The same pendulum, two controllers, one disturbance
LQR — optimal, from the Riccati equation
Learned policy — 113 weights, no equations
Policy search — cost per generation
Angle vs time
Setup
Task
Balance (start near upright)
Swing-up (start hanging)
Initial angle (°)
10
Learned side runs
Policy swings up, LQR catches
Policy alone (watch it drop)
Train the policy
Disturb
Restart
Pause
Live