Research

Reasoning models don't always say what they think

Anthropic

Since late last year, “reasoning models” have been everywhere. These are AI models—such as Claude 3.7 Sonnet—that show their working: as well as their eventual answer, you can read the (often fascinating and convoluted) way that they got there, in what’s called their “Chain-of-Thought”.

Visit Site

Research Anthropic

ResearchAuditing language models for hidden objectivesAnthropic ResearchTowards measuring the representation of subjective global opinions in language modelsAnthropic ResearchMeasuring faithfulness in Chain-of-Thought reasoningAnthropic ResearchDecomposing language models into componentsAnthropic ResearchSycophancy to subterfuge: Investigating reward tampering in language modelsAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic BlogsWhy Vulnerability Scanning Isn't Enough To Protect Your AppSocket BlogsIntroducing "safe npm", a Socket npm WrapperSocket News10 Cloud Deficiencies You Should KnowThenewstack NewsMeasuring CI/CD Adoption Rates Is a ProblemThenewstack NewsHow the Tech World Honored JuneteenthThenewstack NewsThe Hardware and Software Used in SpaceThenewstack BlogsWhat Exactly Is Cloud Engineering?Pulumi BlogsDeploy AI Models on Amazon SageMaker using Pulumi Python IaCPulumi EventsElastic at Gartner IT Symposium/Xpo 2025Elastic EventsElastic Security at Black Hat 2026Elastic EventsElastic at RSACElastic BlogsThe Origin Story of Container QueriesCss Tricks BlogsFirst Steps into a Possible CSS Masonry LayoutCss Tricks BlogsBuilding a Progress Ring, QuicklyCss Tricks