Research

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Anthropic

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Visit Site

Research Anthropic

ResearchProject Swap: What happens when agents trade for us?Anthropic ResearchTracing model outputs to the training dataAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchAuditing language models for hidden objectivesAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchSecurity Implementation for Responsible AI: A Practical FrameworkCybersecurity Exchange BlogsAWS Security Hub Adds Socket for Supply Chain SecuritySocket BlogsAI + a16z Podcast: Vibe Coding, Security Risks, and the Path to ProgressSocket Blogs77 Firefox Extensions Linked to Crypto Wallet and Credential TheftSocket BlogsThe Agentic Infrastructure EraPulumi BlogsBenefits of Policy as CodePulumi EventsElastic at Gartner IT Symposium/Xpo 2025 BarcelonaElastic BlogsVulnerability Management in Container ImagesAquasec BlogsA Brief Guide to Supply Chain Security Best PracticesAquasec BlogsLinguistic Lumberjack: Understanding CVE-2024-4323 in Fluent BitAquasec EventsAgents in ProdMotherduck BlogsSign in with a passkey through form autofillWeb BlogsWhat is Appliance Mode? - CerebrasCerebras BlogsStackAI × Cerebras: enabling the fastest inference for enterprise AI agentsCerebras