Research

Improving Reward Models with Synthetic Critiques

Cohere

Reward models (RMs) play a critical role in aligning language models through the process of reinforcement learning from human feedback.

Visit Site

Research Cohere

ResearchImproving Policy Learning via Language Dynamics DistillationCohere ResearchRewardBench 2: Advancing Reward Model EvaluationCohere ResearchHere's a Free Lunch: Sanitizing Backdoored Models with Model MergeCohere ResearchSparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following ModelsCohere ResearchLanguage Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-ThoughtCohere ResearchThe Reality of AI and BioriskCohere NewsMulticloud Paves the Way for Cloud Native Resiliency ModelsThenewstack BlogsMigrating Notion's marketing site to Next.jsNotion So BlogsDeepSeek V4 Pro: Validating Frontier Models for ProductionFireworks BlogsQwen 3.7 Plus is now live on FireworksFireworks BlogsFireLLaVA: the first commercially permissive OSS LLaVA modelFireworks BlogsLaunching Fireworks for Startups Program!Fireworks BlogsAccelerating Code Completion with Fireworks Fast LLM InferenceFireworks BlogsSimplifying Code Infilling with Code Llama and Fireworks.aiFireworks BlogsCursor Composer 2 + FireworksFireworks BlogsDocker Model Runner on DGX Station GB300Docker BlogsAnnouncing IBM Granite AI Models Now Available on Docker HubDocker BlogsvLLM 0.12, Ministral 3 & DeepSeek-V3.2Docker ResearchTracing the thoughts of a large language modelAnthropic BlogsThe AI-driven shift in vulnerability discovery: What maintainers and bug finders need to knowCncf