The Two Pillars of Post-training: Reinforcement Learning and Supervised Fine-Tuning

This is the second article in Sharon Zhou’s post-training series. Read part 1 here. In the first post of this series, you learned how post-training closed the fundamental gap in usability of LLMs by making them behave in a certain way. In this post, you’ll ex…


