Research

Scalable Training of Language Models using PAX pjit and TPUv4

Cohere

Modern large language models require distributed training strategies due to their size.

Visit Site

Research Cohere

ResearchLanguage Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-ThoughtCohere ResearchProcedural Knowledge in Pretraining Drives Reasoning in Large Language ModelsCohere ResearchFishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language ModelsCohere ResearchHere's a Free Lunch: Sanitizing Backdoored Models with Model MergeCohere ResearchSparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following ModelsCohere ResearchPrioritized Training on Points that are Learnable, Worth Learning, and Not Yet LearntCohere EventsQuarterly Security Spotlight & Product Updates1password LearnFree "DuckDB in Action" BookMotherduck BlogsSpecs Over Vibes: Consistent AI Results ft. Mark FreemanMotherduck EventsThe MCP Sessions - Vol 1: Sports AnalyticsMotherduck BlogsTrain past the frontier: Training API now generally availableFireworks BlogsWhy do all LLMs need structured output modes?Fireworks BlogsFrontier-lab training infrastructure, now as a serviceFireworks BlogsIntroducing FireRouter with OpusFireworks BlogsGLM 5.2 Fast is live on FireworksFireworks BlogsIntroducing OpenAI gpt-oss (20b & 120b)Fireworks BlogsThe Best 8 LLM API Providers in 2026Fireworks BlogsIntroducing Fireworks on Microsoft Foundry: Bringing Best-in-Class Open Model inference to AzureFireworks BlogsAccelerate your Vision Pipelines with the new NVIDIA Nemotron Nano 2 VL Model on FireworksFireworks BlogsCan open models carry readable silent signals before they speak? Reproducing J-Lens Readouts on KimiFireworks