Research

Measuring LLMs’ ability to develop exploits

Anthropic

Two new academic benchmarks measure AI models' ability to develop exploits, plus an updated benchmark for smart contract exploitation.

Visit Site

Research Anthropic

ResearchMeasuring faithfulness in Chain-of-Thought reasoningAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic ResearchAuditing language models for hidden objectivesAnthropic ResearchEnabling independent research on how people use ClaudeAnthropic ResearchProject Swap: What happens when agents trade for us?Anthropic ResearchForecasting rare language model behaviorsAnthropic ResearchAdaptation Odyssey in LLMs: Why Does Additional Pretraining Sometimes Fail to Improve?Cohere ResearchElo Uncovered: Robustness and Best Practices in Language Model EvaluationCohere ResearchNo News is Good News: A Critique of the One Billion Word BenchmarkCohere BlogsBuilding a News App with Replit Agent: A Step-by-Step Guide - NeonNeon BlogsSix Principles for Production AI Agents - NeonNeon ResourcesTriage IntelligenceLinear ResearchUnderstanding and Mitigating Language Confusion in LLMsCohere ResearchContrastive Policy Gradient: Aligning LLMs on sequence-level scores in a supervised-friendly fashionCohere ResearchInvestigating Continual Pretraining in Large Language Models: Insights and ImplicationsCohere ResearchTo Code, or Not To Code? Exploring Impact of Code in Pre-trainingCohere ResearchScalable Data Ablation Approximations for Language Models through Modular Training and MergingCohere BlogsDevelop and Deploy Voice AI Apps with DockerDocker BlogsObservability: Are You Measuring What Actually Matters?Honeycomb BlogsWhat Is a Multimodal LLM?Cohere