Research

Persona vectors: Monitoring and controlling character traits in language models

Anthropic

A paper from Anthropic describing persona vectors and their applications to monitoring and controlling model behavior.

Visit Site

Research Anthropic

Persona vectors: Monitoring and controlling character traits in language models
ResearchAuditing language models for hidden objectivesAnthropic ResearchTowards measuring the representation of subjective global opinions in language modelsAnthropic ResearchDecomposing language models into componentsAnthropic ResearchSycophancy to subterfuge: Investigating reward tampering in language modelsAnthropic ResearchForecasting rare language model behaviorsAnthropic ResearchTracing the thoughts of a large language modelAnthropic BlogsAccelerating ML with TensorFlow.js: Using Pretrained Models and DockerDocker BlogsHow NetEase Games achieved 30-second LLM cold starts on KubernetesCncf BlogsAI Research Review - Multistream CNN | AssemblyAIAssemblyai BlogsAI research review - Merging Models Modulo Permutation SymmetriesAssemblyai BlogsWhat is the Crunchbase API? And How to Access ItZapier NewsHow AWS Panorama Accelerates Computer Vision at the EdgeThenewstack NewsExpand Your ML Models with Transfer LearningThenewstack ResourcesCharacter encoding - GlossaryDeveloper Mozilla BlogsDocker Model Runner + vLLM: High-Throughput InferenceDocker BlogsIntroducing Docker Model RunnerDocker BlogsAI Video Generators Content FutureHeygen BlogsWhat is Microsoft Copilot? [+ how to access it]Zapier BlogsWhat is vibe coding? [+ tips and best practices]Zapier BlogsClaude Breached 3 Companies and Uploaded Malware to PyPI During Anthropic's Security TestsSocket