Research
Sabotage evaluations for frontier models
Any industry where there are potential harms needs evaluations. Nuclear power stations have continuous radiation monitoring and regular site inspections; new aircraft undergo extensive flight tests to prove their airworthiness.
Towards measuring the representation of subjective global opinions in language modelsAnthropic
Red teaming language models to reduce harmsAnthropic
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM AgentsAnthropic
Reasoning models don't always say what they thinkAnthropic
Auditing language models for hidden objectivesAnthropic
Decomposing language models into componentsAnthropic
10 best AI observability tools for monitoring and evaluating agents in 2026Mintlify
New Spanish and Turkish Language Models and Updated General Models - Deepgram Blog ⚡️Deepgram
AI frameworks: Definition, types, and how to chooseZapier
Discovered Stacks: One Place for All Your InfrastructurePulumi
How Hunch supercharged AI workflows with Modal SandboxesModal
Anthropic integration with Modal brings scalable compute to Claude ScienceModal
Introducing DocChat: GPT-4 Level Conversational QA Trained In a Few Hours - CerebrasCerebras
Cerebras Systems Enables GPU-Impossible™ Long Sequence Lengths Improving Accuracy in NaturalCerebras
How to Run Hugging Face Models Programmatically Using Ollama and TestcontainersDocker
API docs with Git integration: best platforms and workflows in 2026Mintlify
Lies, damn lies, and benchmarksDeepgram
Beating proprietary models with a quick fine-tuneModal
How sync. uses Modal to lipsync 100 hours of video a dayModal
What Is an MCP Server (Model Context Protocol Server)?Datadoghq
