Research

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Cohere

As Large Language Models (LLMs) have become more advanced, they have outpaced our abilities to accurately evaluate their quality.

Visit Site

Research Cohere

ResearchHere's a Free Lunch: Sanitizing Backdoored Models with Model MergeCohere ResearchLLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable ObjectivesCohere ResearchCIRCLE: A Framework for Evaluating AI from a Real-World LensCohere ResearchSparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following ModelsCohere ResearchFrom Tools to Teammates: Evaluating LLMs in Multi-Session Coding InteractionsCohere ResearchLanguage Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-ThoughtCohere BlogsStructured memory management for AI Applications and AI Agents with DuckDBMotherduck BlogsBest Practices for Multi-Turn RLFireworks BlogsWhy do all LLMs need structured output modes?Fireworks BlogsFrontier-lab training infrastructure, now as a serviceFireworks BlogsIntroducing FireRouter with OpusFireworks BlogsGLM 5.2 Fast is live on FireworksFireworks BlogsIntroducing OpenAI gpt-oss (20b & 120b)Fireworks BlogsThe Best 8 LLM API Providers in 2026Fireworks BlogsIntroducing Fireworks on Microsoft Foundry: Bringing Best-in-Class Open Model inference to AzureFireworks BlogsCan open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on KimiFireworks BlogsFireworks Nexus: Drop-in Open Frontier Intelligence for Teams with BudgetsFireworks BlogsYour AI Performance Stack is Fireworks Models with Voyage AI embeddingsFireworks BlogsThe DeepSeek Model Lineup: V3.2, R1, and Distilled Variants Mapped to Production WorkloadsFireworks BlogsInference Providers vs. API Routers: Where Do Your Tokens Actually Come From?Fireworks