Research

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

Anthropic

As AI models get smarter, we also need to become smarter in how we monitor them. More intelligent systems can take more complex actions, which is great news for coders, researchers, and anyone else who’s using AI for difficult work tasks.

Visit Site

Research Anthropic

ResearchFrom shortcuts to sabotage: natural emergent misalignment from reward hackingAnthropic ResearchDiscovering cryptographic weaknesses with ClaudeAnthropic ResearchAn alignment assessment of recent cybersecurity incidentsAnthropic ResearchDecomposing language models into componentsAnthropic ResearchPatterns and problems in multiagent systemsAnthropic ResearchTracing the thoughts of a large language modelAnthropic NewsBest Practices to Optimize Infrastructure Monitoring within DevOps TeamsThenewstack NewsNew Relic One Platform ‘Reimagines’ Full Stack ObservabilityThenewstack BlogsMigrate Datadog telemetry with the OpenTelemetry CollectorClickhouse BlogsVideo: Prometheus monitoring for Tailscale clientsTailscale EventsHex Partners and Agents Data MixerMotherduck BlogsAccelerating Code Completion with Fireworks Fast LLM InferenceFireworks BlogsBuild programmatic agents with the Cursor SDK · CursorCursor JobsForward-Deployed EngineerVercel Products & ServicesNew this month: Performance upgrades, better LLM support, new blocks and moreGitbook Products & ServicesMy first 30 days at GitBook: From customer to colleagueGitbook Products & ServicesNew this month: AI insights, smarter redirects, and improved editor controlsGitbook Products & ServicesHumans are still the real readers of your documentationGitbook Blogs8 AI agent use cases and examples in the workplaceZapier BlogsClickStack and Hud bring runtime intelligence to AI-powered developmentClickhouse