Blogs
Using Model-as-a-Judge for Reward in Reinforcement Finetuning
In domains that are inherently challenging to quantify, such as creative writing, we demonstrate that leveraging a superior large language model (LLM) as a judge can meaningfully improve the performance of the policy model.
