Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
sergiopaniego 
posted an update 10 days ago
Post
647
Something I really like when I study a subject is understanding its history, how it reached the point where it is today

I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words

This is the companion piece to Class 3 of our Training Agents series with @burtenshaw . The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale

https://huggingface.co/blog/sergiopaniego/agentic-rl-2026

I like this historical approach—it makes the evolution from RLHF/PPO to GRPO and agent training much easier to understand. Seeing how the ideas developed across different labs also gives useful context for why these methods matter today. It’s a great companion to the hands-on material, much like how Waco Tribune-Herald phone number resources can provide practical context alongside local news.