Research

BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM

Cohere

Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse use cases.

Visit Site

Research Cohere

ResearchLLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable ObjectivesCohere ResearchCountering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement LearningCohere ResearchPeriodic agent-state based Q-learning for POMDPsCohere ResearchDéjà Vu: Multilingual LLM Evaluation through the Lens of Machine Translation EvaluationCohere ResearchRewardBench 2: Advancing Reward Model EvaluationCohere ResearchNo Need for Explanations: LLMs can implicitly learn from mistakes in-contextCohere BlogsProjectional EditingMartinfowler BlogsGrep a million GitHub repositories via MCPVercel LearnViews and App HomeVercel BlogsRestrict access to your sites with role based redirectsNetlify Products & ServicesNew in GitBook: Global reusable content, auto-updating API docs, and much moreGitbook BlogsHow to Convert Google Sheets to Excel: 2 MethodsZapier ResourcesHandle structFastly ResourcesBuilder structFastly ResearchCIRCLE: A Framework for Evaluating AI from a Real-World LensCohere ResourcesThe Go Programming Language Specification - The Go Programming LanguageGo LearnBigQuery Alternative: MotherDuck Cost & Performance BreakdownMotherduck BlogsBuild sub-second data applications with MotherDuck’s Wasm SDKMotherduck EventsDuck-Flavored BI: Building an Efficient Data StackMotherduck EventsData-based: Going Beyond the DataframeMotherduck