Research

Sycophancy to subterfuge: Investigating reward tampering in language models

Anthropic

Empirical evidence that serious misalignment can emerge from seemingly benign reward misspecification.

Visit Site

Research Anthropic

ResearchAuditing language models for hidden objectivesAnthropic ResearchDecomposing language models into componentsAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchFrom shortcuts to sabotage: natural emergent misalignment from reward hackingAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic LearnModel Types and PerformanceVercel BlogsFast AI Feedback Loops with Honeycomb and OpenTelemetryHoneycomb Products & ServicesSpeech Understanding APIAssemblyai Products & ServicesPre-recorded Speech-to-Text API | AssemblyAIAssemblyai ResearchCritical Learning Periods: Leveraging Early training Dynamics for Efficient Data PruningCohere ResearchThe Art of Asking: Multilingual Prompt Optimization for Synthetic DataCohere Products & ServicesModel Vault | Dedicated Model Inference Platform | CohereCohere ResearchBAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsCohere ResearchAdaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?Cohere ResearchElo Uncovered: Robustness and Best Practices in Language Model EvaluationCohere ResearchNo News is Good News: A Critique of the One Billion Word BenchmarkCohere ResearchFrom One to Many: Expanding the Scope of Toxicity Mitigation in Language ModelsCohere BlogsIn-region inference, open models, and new European infrastructure for sovereign AI.Mistral BlogsAI in abundanceMistral