<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>X2 on Aakash Thatte</title><link>https://sky-2002.github.io/tags/x2/</link><description>Recent content in X2 on Aakash Thatte</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 13 Sep 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://sky-2002.github.io/tags/x2/index.xml" rel="self" type="application/rss+xml"/><item><title>What Pretraining Must Supply for effective RL</title><link>https://sky-2002.github.io/posts/2026-09-13-tiny-rlvr-02/</link><pubDate>Sun, 13 Sep 2026 00:00:00 +0000</pubDate><guid>https://sky-2002.github.io/posts/2026-09-13-tiny-rlvr-02/</guid><description>&lt;p&gt;This is the second post I am writing to discuss about a series of experiments. The first post tries to answer what happens to a model after RL. This post focuses on another question - &lt;strong&gt;How much can RL extract from a model, and is that amount already fixed by what pretraining supplied?&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;The question comes from an observation(mentioned in the &lt;a href="https://sky-2002.github.io/posts/2026-08-30-tiny-rlvr-01/"&gt;first post&lt;/a&gt;) that in the last 20 updates of my runs, around 53% of GRPO groups returned identical rewards - all G completions correct, or all G wrong. Identical rewards mean zero advantage, zero gradient, no learning. The all-correct case indicates nothing left to improve, and the all-wrong case indicates that the model is not sampling anything to reinforce on(neither positive nor negative), for learning correct behaviour, there should at least be one correct completion in the G(8 in my case). If that is the mechanism, then RL can only ever reinforce what the base model already reaches within G tries.&lt;/p&gt;</description></item></channel></rss>