Research

From shortcuts to sabotage: natural emergent misalignment from reward hacking

Anthropic

We show for the first time that realistic AI training processes can accidentally produce misaligned models.

Visit Site

Research Anthropic

ResearchSHADE-Arena: Evaluating Sabotage and Monitoring in LLM AgentsAnthropic ResearchVibe physics: The AI grad studentAnthropic ResearchTowards measuring the representation of subjective global opinions in language modelsAnthropic ResearchTracing the thoughts of a large language modelAnthropic ResearchAn off switch for dual-use knowledgeAnthropic ResearchRed teaming language models to reduce harmsAnthropic BlogsRedis Hacking at the HSHacks III HackathonRedis NewsCerebras Systems Enables GPU-Impossible™ Long Sequence Lengths Improving Accuracy in NaturalCerebras LearnDetect Non-Inclusive Language with Retext and Node.js - Deepgram Blog ⚡️Deepgram Blogs​Synthesia launches talent experience program to learn from and reward the exceptional actors behindSynthesia NewsCirrascale Cloud Services® and Cerebras Systems Announce Availability of Cerebras Cloud @ CirrascaleCerebras ResourcesCalling Java from KotlinKotlinlang BlogsHow To Translate Your Videos Into Any LanguageSynthesia BlogsZapier for Alfred: Run Zaps from your Mac keyboardZapier BlogsContext is Everything: Why Maximum Sequence Length Matters - CerebrasCerebras NewsNational Energy Technology Laboratory and Pittsburgh Supercomputing Center Pioneer First EverCerebras BlogsHow an AI Agent Automates QA for the Cerebras Cloud ConsoleCerebras ResourcesKeyboard shortcuts - MintlifyMintlify BlogsMake a smooth shadow, friend.Css Tricks BlogsNine Keyboard Shortcuts for SQL Flow StateMotherduck