Blogs

FireAttention

Fireworks

Mixtral has recently made waves in the AI community as the first OSS model trained on trillions of tokens to support 'mixture of experts' (MoE), which has promising features to speed up both training and inference.

Visit Site

Blogs Fireworks

BlogsDeepSeek V4 Pro: Validating Frontier Models for ProductionFireworks BlogsQwen 3.7 Plus is now live on FireworksFireworks BlogsFireLLaVA: the first commercially permissive OSS LLaVA modelFireworks BlogsLaunching Fireworks for Startups Program!Fireworks BlogsAccelerating Code Completion with Fireworks Fast LLM InferenceFireworks BlogsSimplifying Code Infilling with Code Llama and Fireworks.aiFireworks BlogsVision Model Platform Updates: Enhanced Capabilities and New FeaturesFireworks BlogsFireAttention V3: Enabling AMD as a viable alternative for GPU inferenceFireworks BlogsFireAttention V4: Industry-Leading Latency and Cost Efficiency with FP4Fireworks BlogsFireAttention V2: 12x faster to make Long Contexts practical for Online InferenceFireworks