Research

Evaluating Claude with BioMysteryBench \ Anthropic

anthropic.com

BioMysteryBench tests Claude on 99 real bioinformatics problems. Opus 4.6 solved 77% of human-solvable questions and some that experts could not.

Visit Site

Research anthropic.com

Listing
ResearchFinding bugs with Claude and property-based testing \ Anthropicanthropic.com ResearchProject Fetch: Can Claude train a robot dog? \ Anthropicanthropic.com ResearchHow Canada uses Claude \ Anthropicanthropic.com ResearchHow Australia uses Claude \ Anthropicanthropic.com ResearchEvaluating feature steering: A case study in mitigating social biases \ Anthropicanthropic.com ResearchEvaluating and Mitigating Discrimination in Language Model Decisionsanthropic.com Blogs10 best AI observability tools for monitoring and evaluating agents in 2026Mintlify BlogsFrom Automation Veteran to AI Pioneer: How Evan Nison transformed his agency with Zapier MCPZapier BlogsAnthropic integration with Modal brings scalable compute to Claude ScienceModal BlogsLies, damn lies, and benchmarksDeepgram BlogsAnthropic Identifies Biased Reasoning and Recklessness as Drivers of Claude’s PyPI AttackSocket BlogsCollaborating with Anthropic on Claude Sonnet 4.5 to power intelligent coding agentsVercel BlogsClaude Connectors: How to Connect Claude to Other AppsZapier BlogsThe Claude Skills I Actually Use for DevOpsPulumi BlogsClaude Code + Dives = Any data UIMotherduck BlogsQuestions to Ask Your Hardened Image ProviderDocker Blogs7 best MCP servers for Claude Code in 2026Mintlify Resourcesllms.txt - MintlifyMintlify BlogsThe 9 best AI coding tools in 2026Zapier BlogsStop Tuning Prompts. Build a Harness.Pulumi