Research

Measuring progress on scalable oversight \ Anthropic

anthropic.com

Humans working with an unreliable LLM assistant outperformed both the model alone and unaided humans on challenging question answering.

Visit Site

Research anthropic.com

Listing
ResearchMeasuring AI agent autonomy in practice \ Anthropicanthropic.com ResearchLLM-discovered 0 days \ Anthropicanthropic.com ResearchMoral self-correction in large language models \ Anthropicanthropic.com ResearchIntroducing the Anthropic Economic Index \ Anthropicanthropic.com ResearchSoftmax linear units \ Anthropicanthropic.com ResearchA small number of samples can poison LLMs \ Anthropicanthropic.com BlogsAnthropic integration with Modal brings scalable compute to Claude ScienceModal BlogsThe Livecycle Docker Extension: Instantly Share Changes and Get Feedback in ContextDocker BlogsHow Beyond Retro scaled its retail operations trainingSynthesia ResourcesTime to Interactive (TTI) - GlossaryDeveloper Mozilla NewsStop Technical Debt Before It Damages Your CompanyThenewstack BlogsSmall Data is bigger (and hotter 馃敟) than ever | MotherDuckMotherduck BlogsCerebras CS-3: the world鈥檚 fastest and most scalable AI accelerator - CerebrasCerebras BlogsWhy We Need More Gender Diversity in the Cybersecurity SpaceDocker BlogsHow to build scalable AI applicationsVercel BlogsReal-time AI Is InevitableDeepgram BlogsProgress Delayed Is Progress Denied | CSS-TricksCss Tricks BlogsTurbopack updates: Moving homesVercel BlogsAI + a16z Podcast: Vibe Coding, Security Risks, and the Path to ProgressSocket NewsMeasuring CI/CD Adoption Rates Is a ProblemThenewstack