<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>reinforcement learning on Sadman Kabir Soumik</title>
    <link>https://soumik.blog/tags/reinforcement-learning/</link>
    <description>Recent content in reinforcement learning on Sadman Kabir Soumik</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>en</language>
    <copyright>Copyright © 2022, Sadman Kabir Soumik</copyright>
    <lastBuildDate>Sun, 14 Dec 2025 18:00:00 +0000</lastBuildDate><atom:link href="https://soumik.blog/tags/reinforcement-learning/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>How LLM Post Training Works: SFT, Preference Training, Reinforcement Learning, and Distillation</title>
      <link>https://soumik.blog/artificial-intelligence/how-llm-post-training-works/</link>
      <pubDate>Sun, 14 Dec 2025 18:00:00 +0000</pubDate>
      
      <guid>https://soumik.blog/artificial-intelligence/how-llm-post-training-works/</guid>
      <description>
        
          
            Training a large language model usually happens in more than one stage.
The first stage is pretraining.
This is where the model learns language, facts, patterns, coding concepts, reasoning patterns, and many other things from a very large amount of data.
But a pretrained model is not automatically a good assistant.
It might know a lot, but it may still struggle to follow instructions, answer in the right format, understand human preferences, or behave the way we want.
          
          
        
      </description>
    </item>
    
  </channel>
</rss>
