Research

Training a helpful and harmless assistant with RLHF

Anthropic

RLHF fine-tuning improved performance on almost all NLP evaluations, showing essentially no alignment tax for helpfulness and harmlessness.

Visit Site

Research Anthropic

ResearchTracing model outputs to the training dataAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchAuditing language models for hidden objectivesAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic ResearchProject Swap: What happens when agents trade for us?Anthropic ResearchForecasting rare language model behaviorsAnthropic NewsLinux Foundation Announces 2013 Event and Co-Located Linux Training Schedule - Linux FoundationLinuxfoundation NewsOpenChain Project Gains Facebook, Google and Uber as Platinum Members - Linux FoundationLinuxfoundation BlogsAI Video Tools for Businesses: Improve CommunicationHeygen BlogsTips on How to Create Face Swap VideosHeygen BlogsIntroducing AI Assistant: Turning docs into your product expertMintlify BlogsHow Mintlify uses Claude Code as a technical writing assistantMintlify BlogsBuilding AI That Builds With YouRetool BlogsMaking Very Small LLMs Smarter With RAGDocker EventsAWS Summit London \ AnthropicAnthropic BlogsLinux Training Scholarship Deadline this Friday - Linux FoundationLinuxfoundation BlogsThis Week at Zed Industries: #7 - Zed BlogZed BlogsZed AI: Introducing Usage-Based Billing for High-Volume Users - Zed BlogZed BlogsIntroducing the assistant panel - Zed BlogZed BlogsAdding the ESLint Tool to an AI Assistant: Improving Recommendations for JS/TS ProjectsDocker