Research
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
The disconnect between tokenizer creation and model training in language models allows for specific inputs, such as the infamous Solid Gold Magikarp token, to induce unwanted model behaviour.
