Research

Constitutional Classifiers: Defending against universal jailbreaks

Anthropic

A paper from Anthropic describing a new way to guard LLMs against jailbreaking.

Visit Site

Research Anthropic

ResearchVibe physics: The AI grad studentAnthropic ResearchTowards measuring the representation of subjective global opinions in language modelsAnthropic ResearchTracing the thoughts of a large language modelAnthropic ResearchAn off switch for dual-use knowledgeAnthropic ResearchRed teaming language models to reduce harmsAnthropic ResearchAI agents find smart contract exploitsAnthropic NewsData Ethics Researcher Cautions Against Algorithmic Reordering of SocietyThenewstack BlogsProtecting Against CDN Cache PoisoningBunny BlogsIntroducing Socket Firewall: Free, Proactive Protection for Your Software Supply ChainSocket BlogsIntroducing the pulumi policy analyze Command for Existing StacksPulumi BlogsReal-world Cyber Attacks Targeting Data Science ToolsAquasec BlogsA “new direction” in the struggle against rightward scrollingCss Tricks BlogsTrivy: The Universal Scanner to Secure Your Cloud MigrationAquasec BlogsNow Available: 99 Languages, Advanced Features, One PriceAssemblyai Products & ServicesUniversal-3 Pro StreamingAssemblyai BlogsGPT-6 Astra Attempts Supply Chain Attacks Against Open Source Maintainers in TestingSocket BlogsSecure AI Agents at Runtime with DockerDocker BlogsAgora voice agent with AssemblyAI Universal-3.6 Pro RealtimeAssemblyai BlogsAI medical scribe: build vs buy against Nuance DAX and AbridgeAssemblyai BlogsUniversal model improvements: Introducing advanced contextual text formatting for Spanish and GermanAssemblyai