Research

Evaluating and Mitigating Discrimination in Language Model Decisions

anthropic.com

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Visit Site

Research anthropic.com

Evaluating and Mitigating Discrimination in Language Model Decisions
ResearchEvaluating feature steering: A case study in mitigating social biases \ Anthropicanthropic.com ResearchEmotion concepts in a large language model \ Anthropicanthropic.com ResearchMoral self-correction in large language models \ Anthropicanthropic.com ResearchMitigating prompt injections in browser use \ Anthropicanthropic.com ResearchEvaluating Claude with BioMysteryBench \ Anthropicanthropic.com ResearchCommitments on model deprecation and preservation \ Anthropicanthropic.com Blogs10 best AI observability tools for monitoring and evaluating agents in 2026Mintlify BlogsTrained on 100,000+ Voices: Deepgram Unveils Next-Gen Speaker Diarization and Language DetectionDeepgram BlogsThe Language of LGBTQ Inclusion and Allyship - Deepgram Blog ⚡️Deepgram BlogsHappy National Native American Heritage Month! - Deepgram Blog ⚡️Deepgram BlogsNew Spanish and Turkish Language Models and Updated General Models - Deepgram Blog ⚡️Deepgram NewsOData or GraphQL? The Best Tech for Developing an API Is Neither or Both!Thenewstack BlogsPkg.go.dev has a new look! - The Go Programming LanguageGo BlogsInside the Go Playground - The Go Programming LanguageGo BlogsReal Go Projects: SmartTwitter and web.go - The Go Programming LanguageGo BlogsThe App Engine SDK and workspaces (GOPATH) - The Go Programming LanguageGo BlogsGo on App Engine: tools, tests, and concurrency - The Go Programming LanguageGo ResourcesThe Go Memory Model - The Go Programming LanguageGo BlogsErrors are values - The Go Programming LanguageGo BlogsUsing Subtests and Sub-benchmarks - The Go Programming LanguageGo