Research
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
As AI models get smarter, we also need to become smarter in how we monitor them. More intelligent systems can take more complex actions, which is great news for coders, researchers, and anyone else who’s using AI for difficult work tasks.
