What Pretraining Must Supply for effective RL

This is the second post I am writing to discuss about a series of experiments. The first post tries to answer what happens to a model after RL. This post focuses on another question - How much can RL extract from a model, and is that amount already fixed by what pretraining supplied? The question comes from an observation(mentioned in the first post) that in the last 20 updates of my runs, around 53% of GRPO groups returned identical rewards - all G completions correct, or all G wrong. Identical rewards mean zero advantage, zero gradient, no learning. The all-correct case indicates nothing left to improve, and the all-wrong case indicates that the model is not sampling anything to reinforce on(neither positive nor negative), for learning correct behaviour, there should at least be one correct completion in the G(8 in my case). If that is the mechanism, then RL can only ever reinforce what the base model already reaches within G tries. ...

September 13, 2026 · 10 min

What RL Actually Changes in a Model

I have been following the discourse around RL, hearing a lot of terminology (like async RL, sample efficiency etc) and how people are using it to post-train models, and I did RL on smaller models like Qwen by following the docs of trl, unsloth. I could see the rewards going up, sometimes wiggling, sometimes 0, at times answer still being wrong at the end of training. I knew the math, the code, but I wanted to go deeper and had a lot of questions, like: what exactly changes in a model’s behaviour after RL? does it depend on how the pretraining was done? how stale can the rollouts be? - and many more. So I did some experiments and now writing this post to present those. ...

August 30, 2026 · 17 min