Research

Evaluating feature steering: A case study in mitigating social biases \ Anthropic

anthropic.com

A new piece of Anthropic research by Durmus et al.: "Evaluating feature steering: A case study in mitigating social biases"

Visit Site

Research anthropic.com

Listing
ResearchEvaluating and Mitigating Discrimination in Language Model Decisionsanthropic.com ResearchMitigating prompt injections in browser use \ Anthropicanthropic.com ResearchEvaluating Claude with BioMysteryBench \ Anthropicanthropic.com ResearchLLM-discovered 0 days \ Anthropicanthropic.com ResearchMoral self-correction in large language models \ Anthropicanthropic.com ResearchIntroducing the Anthropic Economic Index \ Anthropicanthropic.com Blogs10 best AI observability tools for monitoring and evaluating agents in 2026Mintlify BlogsNew Spanish and Turkish Language Models and Updated General Models - Deepgram Blog ⚡️Deepgram BlogsNoSQLNow! 2014 Case Study: Bleacher Report on Redis LabsRedis NewsGrafana Is Not Worried About AWS CommercializationThenewstack NewsIntegrate Jupyter Notebooks with GitHubThenewstack BlogsAnnouncing Restore Stacks: Recover Deleted Stacks in the Pulumi CloudPulumi BlogsDevelop, Preview, TestCss Tricks BlogsDocker Desktop 4.18: Docker Scout Updates, Container File Explorer GADocker BlogsHow to Use AI Image to Video Generator for Free to Create Dynamic VideosHeygen BlogsMeta AI Video Generator's ImpactHeygen ResourcesRolling out a new featureVercel BlogsLies, damn lies, and benchmarksDeepgram BlogsHow to make the case for LinkedIn CAPIZapier Blogsre:Invent 2020 EKS Feature ReleasesPulumi