Research

Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Cohere

The disconnect between tokenizer creation and model training in language models allows for specific inputs, such as the infamous Solid Gold Magikarp token, to induce unwanted model behaviour.

Visit Site

Research Cohere

ResearchInvestigating Continual Pretraining in Large Language Models: Insights and ImplicationsCohere ResearchScalable Data Ablation Approximations for Language Models through Modular Training and MergingCohere ResearchBigScience: A Case Study in the Social Construction of a Multilingual Large Language ModelCohere ResearchCALIBER: Calibrating confidence before and after reasoning in language modelsCohere ResearchUnderstanding and Mitigating Language Confusion in LLMsCohere ResearchOne Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual TokenizersCohere Blogs10 best AI observability tools for monitoring and evaluating agents in 2026Mintlify BlogsTrained on 100,000+ Voices: Deepgram Unveils Next-Gen Speaker Diarization and Language DetectionDeepgram BlogsThe Language of LGBTQ Inclusion and Allyship - Deepgram Blog ⚡️Deepgram BlogsNew Spanish and Turkish Language Models and Updated General Models - Deepgram Blog ⚡️Deepgram BlogsCustomer.io Automatically Standardizes Form Submissions for ReportingZapier Blogs6 ways to use the Zapier Zoho CRM integrationZapier BlogsAI frameworks: Definition, types, and how to chooseZapier BlogsGet email alerts for Facebook Messenger messagesZapier BlogsLooping by Zapier: Repeat actions for itemsZapier NewsOData or GraphQL? The Best Tech for Developing an API Is Neither or Both!Thenewstack BlogsDiscovered Stacks: One Place for All Your InfrastructurePulumi BlogsPulumi Neo Now Supports AGENTS.mdPulumi BlogsPkg.go.dev has a new look! - The Go Programming LanguageGo BlogsInside the Go Playground - The Go Programming LanguageGo