Research

Natural Language Autoencoders

Anthropic

AI models like Claude talk in words but think in numbers. In this study, we train Claude to translate its thoughts into human-readable text.

Visit Site

Research Anthropic

ResearchAuditing language models for hidden objectivesAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchDecomposing language models into componentsAnthropic ResearchSycophancy to subterfuge: Investigating reward tampering in language modelsAnthropic ResearchFrom shortcuts to sabotage: natural emergent misalignment from reward hackingAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic Products & ServicesPre-recorded Speech-to-Text API | AssemblyAIAssemblyai ResearchThe Art of Asking: Multilingual Prompt Optimization for Synthetic DataCohere ResearchBAM! Just Like That: Simple and Efficient Parameter Upcycling for Mixture of ExpertsCohere ResearchElo Uncovered: Robustness and Best Practices in Language Model EvaluationCohere ResearchNo News is Good News: A Critique of the One Billion Word BenchmarkCohere ResearchFrom One to Many: Expanding the Scope of Toxicity Mitigation in Language ModelsCohere BlogsUnlocking the potential of vision language models on satellite imagery through fine-tuningMistral BlogsHoneycomb Launches Integration With the Anthropic Usage and Cost APIHoneycomb ResearchUnderstanding and Mitigating Language Confusion in LLMsCohere ResearchContrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashionCohere ResearchInvestigating Continual Pretraining in Large Language Models: Insights and ImplicationsCohere ResearchEAGER: Entropy-Aware Generation for Adaptive Inference-Time ScalingCohere ResearchScalable Data Ablation Approximations for Language Models through Modular Training and MergingCohere BlogsI made a policy engine think it was in productionCncf