Research

Contrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashion

Cohere

Reinforcement Learning (RL) has been used to finetune Large Language Models (LLMs) using a reward model trained from preference data, to better align with human judgment.

Visit Site

Research Cohere

ResearchHow Does Quantization Affect Multilingual LLMs?Cohere ResearchWhen Less is More: Investigating Data Pruning for Pretraining LLMs at ScaleCohere ResearchConsent in Crisis: The Rapid Decline of the AI Data CommonsCohere ResearchTreasure Hunt: Real-time Targeting of the Long Tail using Training-Time MarkersCohere ResearchGoodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented ModelsCohere ResearchBERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLMCohere BlogsPolicy-as-Code: Flexible Kubernetes governance with KyvernoCncf BlogsFree llms.txt generator: create AI-optimized documentation filesMintlify ResourcesMermaid - MintlifyMintlify LearnEverything you need to know about Voice AI AgentsDeepgram BlogsNew Study Identifies 53 Slopsquatting Targets Across 5 Frontier LLMsSocket NewsHow Integrations Help Enterprises Level Up with KubernetesThenewstack BlogsIntroducing the pulumi policy analyze Command for Existing StacksPulumi BlogsAgent architecture: How AI decision-making drives business impactRetool BlogsPowerHell: Active Flaws in PowerShell Gallery Expose Users to AttacksAquasec BlogsFront-End ChallengesCss Tricks ResourcesCompiler options in the Kotlin Gradle pluginKotlinlang BlogsPrivate MCP Catalogs & Composable Enterprise AIDocker BlogsIs your fancy new domain hurting your performance? Benchmarking the top-level domain namesBunny BlogsGitOps policy-as-code: Securing Kubernetes with Argo CD and KyvernoCncf