Research

Reverse Engineering Human Preferences with Reinforcement Learning

Cohere

The capabilities of Large Language Models (LLMs) are routinely evaluated by other LLMs trained to predict human preferences.

Visit Site

Research Cohere

ResearchThe PRISM Alignment Project: What Participatory, Representative and Individualised Human FeedbackCohere ResearchConsent in Crisis: The Rapid Decline of the AI Data CommonsCohere ResearchTreasure Hunt: Real-time Targeting of the Long Tail using Training-Time MarkersCohere ResearchHow Does Quantization Affect Multilingual LLMs?Cohere ResearchGoodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented ModelsCohere ResearchBERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLMCohere BlogsHow to get engineering time back from Kubernetes upgradesCncf BlogsA new chapter: Bin Liu joins HeyGenHeygen BlogsHow to Use AI for Character GenerationHeygen ResourcesAnalyze documentation traffic - MintlifyMintlify PodcastsWhat Role Does Bias Have in Machine Learning? — AI ShowDeepgram BlogsHow to Choose an eLearning Development Company (And When to Build In-House)Synthesia BlogsHow Criteo uses AI video for newcomer onboardingSynthesia BlogsThe 7 Best Note Taking Apps in 2026Zapier BlogsAI in HR: Benefits, types, and 7 use cases worth runningZapier NewsWSO2’s Platform Widens Scope of API ManagementThenewstack NewsWhen Holt-Winters Is Better Than Machine LearningThenewstack NewsAn End to the Confusion: Jenkins or Jenkins XThenewstack NewsDynatrace Previews Next Generation Monitoring with RuxitThenewstack BlogsPreview of the Deploy Track at Cloud Engineering Summit 2021Pulumi