Research

Decomposing language models into components

Anthropic

Interpretability research decomposing a transformer layer into thousands of interpretable features that can steer model behavior.

Visit Site

Research Anthropic

ResearchAuditing language models for hidden objectivesAnthropic ResearchTowards measuring the representation of subjective global opinions in language modelsAnthropic ResearchSycophancy to subterfuge: Investigating reward tampering in language modelsAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchTracing the thoughts of a large language modelAnthropic ResearchConstitutional Classifiers: Defending against universal jailbreaksAnthropic BlogsLM Link: Use local models on remote devices, powered by TailscaleTailscale LearnNo More Writing SQL for Quick AnalysisMotherduck BlogsConversing with Large Language Models using DaprCncf LearnRemote Components IntroVercel BlogsWeb Components: Working With Shadow DOMSmashingmagazine NewsThe Programming Language Mathematica Marks a MilestoneThenewstack NewsMulticloud Paves the Way for Cloud Native Resiliency ModelsThenewstack BlogsRobust generic functions on slices - The Go Programming LanguageGo BlogsA Proposal for Package Versioning in Go - The Go Programming LanguageGo BlogsContributors Summit 2019 - The Go Programming LanguageGo BlogsGo Protobuf: The new Opaque API - The Go Programming LanguageGo LearnTutorial: Getting started with multi-module workspaces - The Go Programming LanguageGo BlogsWhen To Use Generics - The Go Programming LanguageGo ResourcesGo Wiki: Compiler And Runtime Optimizations - The Go Programming LanguageGo