Research

Using dictionary learning features as classifiers

Anthropic

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Visit Site

Research Anthropic

ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchAuditing language models for hidden objectivesAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic ResearchProject Swap: What happens when agents trade for us?Anthropic ResearchForecasting rare language model behaviorsAnthropic ResearchDisempowerment patterns in real-world AI usageAnthropic BlogsAI Strategies for Software Engineering Career GrowthHoneycomb BlogsUsing GitHub Copilot to Speed Up Your Development WorkflowHoneycomb BlogsTransforming How We Run Kafka at HoneycombHoneycomb BlogsCreate Environments with Masked Production Data Using Neon Branches - NeonNeon BlogsMySQL data types: VARCHAR and CHARPlanetscale BlogsWorking with Geospatial Features in MySQLPlanetscale BlogsSelf-managed Vitess vs Managed Vitess with PlanetScalePlanetscale BlogsHow RevOps Teams Are Building Their Own ToolsReplit BlogsReplit in Review: A Recap of What We Shipped in | ReplitReplit BlogsUsing ClickHouse as a webhook endpoint with HMAC verificationClickhouse BlogsLessons learned from launching our new free planRetool BlogsThe 10 best React admin templatesRetool BlogsUsing LSM Hooks with Tracee to Overcome Gaps with Syscall TracingAquasec BlogsAutomating Configuration Auditing with Starboard OperatorAquasec