Research

Discovering behaviors with model-written evaluations

Anthropic

154 automatically generated evaluation datasets reveal sycophancy and other concerning behaviors that grow with model size and RLHF.

Visit Site

Research Anthropic

ResearchForecasting rare language model behaviorsAnthropic ResearchTracing model outputs to the training dataAnthropic ResearchDiscovering cryptographic weaknesses with ClaudeAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchAuditing language models for hidden objectivesAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic BlogsHow Much Should My Observability Stack Cost?Honeycomb Blogs15 Best AI Observability Tools for Production Teams in 2026Honeycomb BlogsInstrumenting AI Agents for the Agent Timeline: A Practical OpenTelemetry GuideHoneycomb BlogsNew in 0.15: the ChiselStrike TypeScript client APITurso BlogsMicrobatch: how to supercharge dbt-duckdb with the right incremental modelMotherduck BlogsBuild a News Roundup With Docker AgentDocker BlogsPowering Local AI Together: Docker Model Runner on Hugging FaceDocker BlogsAI Model Drift: How to Keep Models ReliableHoneycomb BlogsWide Events vs. Three Pillars: AI Observability CostsHoneycomb BlogsHow we used Quint to find over 10 bugs in SQLite while hardening TursoTurso BlogsMaking Very Small LLMs Smarter With RAGDocker BlogsOpenWebUI + Model Runner: Zero-Config Local AIDocker ResourcesOpenResponses API with AI GatewayVercel NewsLF Deep Learning Foundation Advances Open Source Artificial Intelligence With Major MembershipLinuxfoundation