Research

Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Cohere

The disconnect between tokenizer creation and model training in language models allows for specific inputs, such as the infamous Solid Gold Magikarp token, to induce unwanted model behaviour.

Visit Site

Research Cohere

ResearchLanguage Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-ThoughtCohere ResearchImproving Policy Learning via Language Dynamics DistillationCohere ResearchHere's a Free Lunch: Sanitizing Backdoored Models with Model MergeCohere ResearchSparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following ModelsCohere ResearchThe Reality of AI and BioriskCohere ResearchThe Culture Funnel: You can’t align what isn’t in the dataCohere NewsThe Programming Language Mathematica Marks a MilestoneThenewstack NewsAsyncAPI Looks to Unify API Workflow under Linux FoundationThenewstack NewsMulticloud Paves the Way for Cloud Native Resiliency ModelsThenewstack BlogsRobust generic functions on slices - The Go Programming LanguageGo BlogsA Proposal for Package Versioning in Go - The Go Programming LanguageGo BlogsContributors Summit 2019 - The Go Programming LanguageGo BlogsGo Protobuf: The new Opaque API - The Go Programming LanguageGo LearnTutorial: Getting started with multi-module workspaces - The Go Programming LanguageGo BlogsWhen To Use Generics - The Go Programming LanguageGo ResourcesGo Wiki: Compiler And Runtime Optimizations - The Go Programming LanguageGo ResourcesGo Wiki: Go Code Review Comments - The Go Programming LanguageGo BlogsMigrating to Go Modules - The Go Programming LanguageGo BlogsArrays, slices (and strings): The mechanics of 'append' - The Go Programming LanguageGo ResourcesGo Wiki: InstallTroubleshooting - The Go Programming LanguageGo