Building enterprise AI solutions at Needl - everything agents, search and workflows. I am exploring and playing around with agents, RL. Once a while I also try coding model architectures.
What Pretraining Must Supply for effective RL
This is the second post I am writing to discuss about a series of experiments. The first post tries to answer what happens to a model after RL. This post focuses on another question - How much can RL extract from a model, and is that amount already fixed by what pretraining supplied? The question comes from an observation(mentioned in the first post) that in the last 20 updates of my runs, around 53% of GRPO groups returned identical rewards - all G completions correct, or all G wrong. Identical rewards mean zero advantage, zero gradient, no learning. The all-correct case indicates nothing left to improve, and the all-wrong case indicates that the model is not sampling anything to reinforce on(neither positive nor negative), for learning correct behaviour, there should at least be one correct completion in the G(8 in my case). If that is the mechanism, then RL can only ever reinforce what the base model already reaches within G tries. ...
What RL Actually Changes in a Model
I have been following the discourse around RL, hearing a lot of terminology (like async RL, sample efficiency etc) and how people are using it to post-train models, and I did RL on smaller models like Qwen by following the docs of trl, unsloth. I could see the rewards going up, sometimes wiggling, sometimes 0, at times answer still being wrong at the end of training. I knew the math, the code, but I wanted to go deeper and had a lot of questions, like: what exactly changes in a model’s behaviour after RL? does it depend on how the pretraining was done? how stale can the rollouts be? - and many more. So I did some experiments and now writing this post to present those. ...
RL for LLMs - from softmax to GRPO
After learning about RL for LLMs in research papers, some nice youtube playlists and reading some amazing blogs, I also wanted to write one, which can introduce RL for LLMs for someone who already knows LLMs, and one which I can refer anytime as a refresher. I have tried to build this blog right from softmax all the way to GRPO, and also tried to show calculations on a five-token toy example(which I will skip in one of the sections). ...
(GSoC) Summary of Neural Extraction Framework
Hello. This is the last blog related to GSoC as my project nears to an end and the final evaluation approaches. In this blog, I will try to summarize the project, write about what could be added further and future scope etc. Summary of Neural Extraction Framework GSoC project page - Link Github Repo - Link Neural Extraction Framework has been a wonderful experience for me. This project was in its third iteration this year in the Google Summer of Code and I am delighted to share that we have achieved to have an end-2-end framework to automatically extract relational triples from wikipedia articles and link them to entities and predicates in DBpedia. ...
(GSoC) Week-14 Recent Updates
Hello. This blog will be about a few updates I made in recent weeks. Also, I am excited to share that I will be giving a lightning talk in the Google Summer of Code Virtual Contributor Lightning Talks 2023 on my project. Its going to be a short talk(3 minutes max), I will be summarizing my project and I am really excited for it! In think blog, I will give some info on the recent updates, like adding coreference resolution, some small speedups we got in the code, etc. ...
(GSoC) Week-11, 12 and 13 Recent Updates
Hello. Its been a while since I have put an update on the project(travelling and network issues). I had been working on some polishing, enhancing and re-experimenting, with the models that we are using, the time it takes to process the inputs etc. In this blog, I will give a review of these. Restating the Goal Just to make sure we are clear about the end goal, given a wikipedia page(an entity), we need to disambiguate the predicates/relationships between the page and all other pages(other entities) that are linked from that page. More precisely, we want to find the relationships between those pages which are only connected by the dbo:wikiPageWikiLink predicate and no other predicate, or mine the text to see even a relationship exists or not. There are cases where some page is linked in references, but is not related to the wikipedia page. The goal is to find all those hidden relationships by processing the wiki page text. ...
(GSoC) Week-9 and 10 Mid Evaluations and remaining TODOs
Hello. I recently passed my mid term evaluation for GSoC 2023. It has been a great experience so far. DBpedia finally has an end-2-end neural extraction framework. Got insightful feedback from the mentors. In this blog, I will discuss about some tasks that still need to be done in order to make this end-2-end relation extraction reliable and correct. Background To give a quick overview - our end-2-end framework takes in raw text as input and gives a list of (<subject>, <predicate>, <object>) triples as output. We are using REBEL relation extraction model for joint entity-relation extraction, GENRE for entity linking and text based vector similarity for mapping relations to predicates. This pipeline is explained in the [previous blog]({% post_url 2023-07-24-Week-8-End-2-End-RE %}) and works great. ...
(GSoC) Week-8 End-2-End Relation extraction
Hello. It feels amazing to start this blog as we have been able to create an end-2-end system for relation extraction from text, the ultimate goal of this Neural Extraction Framework project. Before we begin, check the sample below of how the extracted and mapped triples look like: Expectations The expectation from an end-2-end relations extraction is to take as input a wikipedia article text and return a set of relational triples. These relational triples are in the form of <subject_entity> , <predicate> , <object_entity>/<literal>. Also, initially we want these triples to have the wikipedia page itself as the subject entity. For example, if the page is about Berlin_Wall, we would like to pick the triples where the subject entity is Berlin_Wall. This way, we have same objective as extracting from infobox, just that we also make use of the text to extract relations and thus enhance the process to mine more information about that entity/page. ...
(GSoC) Week-7 Predicate mapping
Hello. In this blog, we will be looking into mapping natural language relations to predicates in DBpedia. This task is extremely important in order to achieve our ultimate goal of a neural extraction framework. Background In the previous post, we had discussed about REBEL model for relation extraction. This model extracts the head entity, relation and tail entity from texts. But they are in natural language form - for example, “lives in” - we want them to be in the form of URIs - for example “https://dbpedia.org/ontology/residence" - so that we can link them to dbpedia predicates. For mapping entities to dbpedia resource URIs, we already have a few methods for entity linking. ...
(GSoC) Week-6 Relation Extraction - REBEL
Hello. The previous blog was the last in the entity-linking part of the project. Now we will move onto relation extraction. In this blog, we will discuss about REBEL - a relation extraction model. Given a text, this model can find relational triples(head-relation-tail) in the form of natural language text. Background REBEL is a text2text model trained by BabelScape by fine-tuning BART for translating a raw input sentence containing entities and implicit relations into a set of triplets that explicitly refer to those relations. It has been trained on more than 200 different relation types. REBEL is a joint model, meaning that it extracts entities and relations simultaneously. ...