Research

Red teaming language models to reduce harms

Anthropic

Red teaming models across scales shows RLHF models become harder to attack as they grow, with a released dataset of 38,961 attacks.

Visit Site

Research Anthropic

ResearchAuditing language models for hidden objectivesAnthropic ResearchTowards measuring the representation of subjective global opinions in language modelsAnthropic ResearchDecomposing language models into componentsAnthropic ResearchSycophancy to subterfuge: Investigating reward tampering in language modelsAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchTracing the thoughts of a large language modelAnthropic EventsWebinar On DemandWhat's new? The 1Password quarterly security spotlight and roadmap review - Q2 20261password BlogsPicture This: Open Source AI for Image DescriptionFly Products & ServicesAI Video ModelsFal BlogsDocker, JetBrains, and Zed: Building a Common Language for Agents and IDEsDocker BlogsGemma 4 Is Now Available on Docker HubDocker BlogsSentimenAnalysis and Insights on Cryptocurrencies Using Docker and Containerized AI/ML ModelsDocker NewsWhat Is Cloud Automation and How Does It Benefit IT Teams?Thenewstack NewsRed Hat OpenShift 4.8 Adds Serverless Functions, Pipelines-As-CodeThenewstack EventsWebinar On DemandThe unmanaged stack: Governing SaaS apps and AI tools outside SSO1password BlogsBalancing cost and reliability for Spark on KubernetesNotion So BlogsRunning a self-hosted LLM in Kubernetes with vLLMCncf ResourcesVercel CDN CompressionVercel BlogsNetlify Build Plugin of the Week: Hugo CacheNetlify BlogsHow to Reduce Training Costs Without Cutting What WorksSynthesia