Blogs

Using Model-as-a-Judge for Reward in Reinforcement Finetuning

Fireworks

In domains that are inherently challenging to quantify, such as creative writing, we demonstrate that leveraging a superior large language model (LLM) as a judge can meaningfully improve the performance of the policy model.

Visit Site

Blogs Fireworks

BlogsQwen3 Decoded: Choosing the Right Model For Your TaskFireworks BlogsFireFunction V1 - Fireworks’ GPT-4-level function calling model - 4x faster than GPT-4 and openFireworks BlogsHow Gumloop Scaled Open-Weight Model Usage 7x in 3 Weeks with FireworksFireworks BlogsReinforcement Fine Tuning: Train expert open models to surpass closed frontier modelsFireworks BlogsReinforcement learning: Why alignment of numerics and MoE routing matterFireworks BlogsThe frontier isn’t a model. It’s a router.Fireworks ResourcesVercel Functions LimitsVercel ResourcesHow to alias a preview deployment using the CLIVercel ResourcesUsing Headless WordPress with Next.js and VercelVercel ResourcesDeploy a headless Shopify storefront with VercelVercel LearnModel Types and PerformanceVercel Resourceswith Image OptimizationVercel ResourcesHow to test a Slack bot with your Vercel preview deploymentVercel BlogsNetlify pro tip: Using Split Testing to power private beta releasesNetlify BlogsThere Is Only One Key Difference Between Observability 1.0 and 2.0Honeycomb BlogsObservability: It's Every Engineer’s Job, Not Just Ops’ ProblemHoneycomb BlogsUsing Core Web Vitals in Honeycomb Frontend TelemetryHoneycomb BlogsHoneycomb Announces Availability of MCP in the New AWS Marketplace AI Agents and Tools CategoryHoneycomb BlogsWhat Is Auto-Instrumentation?Honeycomb Products & ServicesWhat we learned building a complete docs site using Claude, MCP and skill.mdGitbook