Blogs

Better MoE model inference with warp decode · Cursor

Cursor

By flipping the parallelism axis we achieve 1.8x faster and more accurate MoE model inference.

Visit Site

Blogs Cursor

BlogsCursor partners with SpaceX on model training · CursorCursor BlogsHow Wayfair cut ML model costs by 90% (twice!) with Cursor · CursorCursor BlogsGrok 4.5 Model Card · CursorCursor BlogsMixture-of-Kittens: our open-source MoE megakernel for NVL72s · CursorCursor BlogsReward hacking is swamping model intelligence gains · CursorCursor BlogsHow Cursor Router chooses the right model for the task · CursorCursor BlogsNew Spanish and Turkish Language Models and Updated General Models - Deepgram Blog ⚡️Deepgram BlogsHow automation helped us get better customer reviewsZapier NewsWhen CI Meets CD: Deliver With ConfidenceThenewstack NewsDevSecOps: Why Security Shouldn’t be Sacrificed for SpeedThenewstack ResourcesThe Go Memory Model - The Go Programming LanguageGo ResourcesWhat is a CUDA Thread Block?Modal BlogsThe rise of slow personal assistantsCerebras BlogsSimulating Human Behavior with Cerebras - CerebrasCerebras Blogs100x Defect Tolerance: How Cerebras Solved the Yield Problem - CerebrasCerebras NewsCerebras Announces Six New AI Datacenters Across North America and Europe to Deliver Industry’sCerebras NewsAMD and Cerebras Announce Disaggregated AI InferenceCerebras NewsCerebras Systems, Ranovus win $45 million US military deal to speed up chip connectionsCerebras NewsCerebras Systems Raises $250M in Funding for Over $4B Valuation to Advance the Future of ArtificialCerebras BlogsHow AI Assistants Can Decode GitHub Repos for UI WritersDocker