Research

Towards understanding sycophancy in language models

Anthropic

Five state-of-the-art AI assistants consistently exhibit sycophancy, suggesting human feedback training partly drives the behavior.

Visit Site

Research Anthropic

ResearchSycophancy to subterfuge: Investigating reward tampering in language modelsAnthropic ResearchAuditing language models for hidden objectivesAnthropic ResearchDecomposing language models into componentsAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic Products & ServicesCloudflare Workers AI - Edge AI Inference PlatformCloudflare BlogsLet the Robots Generate Calculated Fields For YouHoneycomb BlogsWhat is an agent harness?Zapier BlogsGoogle's Gemini AI models now available on ZapierZapier BlogsUnderstanding the Importance of Runtime Security in Cloud Native EnvironmentsAquasec BlogsMicrobatch: how to supercharge dbt-duckdb with the right incremental modelMotherduck BlogsPowering Local AI Together: Docker Model Runner on Hugging FaceDocker BlogsUnderstanding the Difference Between Login Enforcement and SSODocker ResourcesOpenResponses Tool Calling with AI GatewayVercel BlogsAI Model Drift: How to Keep Models ReliableHoneycomb BlogsCharacter sets and collations in MySQLPlanetscale BlogsRunning unmodified Doom in the SQLite bytecode languageTurso BlogsBuilding a web app with Nix (Because why not?)Replit BlogsVibe Coding Data Apps with Replit + SnowflakeReplit