Research

Auditing language models for hidden objectives

Anthropic

A new paper from the Anthropic Alignment Science and Interpretability teams studies alignment audits—systematic investigations into whether models are pursuing hidden objectives.

Visit Site

Research Anthropic

ResearchDecomposing language models into componentsAnthropic ResearchSycophancy to subterfuge: Investigating reward tampering in language modelsAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic ResearchProject Swap: What happens when agents trade for us?Anthropic Products & ServicesSpeech Understanding APIAssemblyai Products & ServicesPre-recorded Speech-to-Text API | AssemblyAIAssemblyai ResearchCritical Learning Periods: Leveraging Early training Dynamics for Efficient Data PruningCohere ResearchThe Art of Asking: Multilingual Prompt Optimization for Synthetic DataCohere Products & ServicesModel Vault | Dedicated Model Inference Platform | CohereCohere ResearchBAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsCohere ResearchAdaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?Cohere ResearchElo Uncovered: Robustness and Best Practices in Language Model EvaluationCohere ResearchNo News is Good News: A Critique of the One Billion Word BenchmarkCohere ResearchFrom One to Many: Expanding the Scope of Toxicity Mitigation in Language ModelsCohere BlogsIn-region inference, open models, and new European infrastructure for sovereign AI.Mistral BlogsAI in abundanceMistral BlogsSpaces: A CLI Built for Humans and AgentsMistral BlogsUnlocking the potential of vision language models on satellite imagery through fine-tuningMistral