Blogs

FireAttention V4: Industry-Leading Latency and Cost Efficiency with FP4

Fireworks

Today, we’re announcing we've achieved industry-leading speeds of >250 tokens/second on NVIDIA B200 GPUs using our latest FireAttention V4 inference engine.

Visit Site

Blogs Fireworks

BlogsFrontier AI at a fraction of the cost: open-source worker agents with a closed-source advisor.Fireworks BlogsFireworks Raises $52M Series B to Lead Industry Shift to Compound AI SystemsFireworks BlogsDeepSeek-V4.1-Flash on Fireworks: Astra-level DeepSWE at 1/15th the costFireworks BlogsDeepSeek V4 Pro: Validating Frontier Models for ProductionFireworks BlogsQwen 3.7 Plus is now live on FireworksFireworks BlogsFireLLaVA: the first commercially permissive OSS LLaVA modelFireworks NewsMicrosoft: 5 Ways to Make Open Source a Reality at Your CompanyThenewstack NewsBreak the Kubernetes Iron Triangle by Optimizing Your AppsThenewstack LearnThe Database Inside Your Lakehouse: A DuckLake Architecture Deep DiveMotherduck BlogsDuckDB Wasm : What Happens When You Put a Database in Your Browser?Motherduck BlogsTwo years of vector search at Notion: 10x scale, 1/10th costNotion So BlogsMastering Peak Software Development EfficiencyDocker BlogsSelf-hosted human and machine identities in Keycloak 26.4Cncf BlogsAI Character Voice Generator: Clone Any Voice for GamesHeygen BlogsProcess analysis: Methods and a 6-step frameworkZapier Products & ServicesSaaS Manager has 350+ Integrations for SaaS Management1password BlogsHow the leading AI companies are using AI toolsNotion So BlogsThe crucial moments and decisions leading up to the launch of Notion AINotion So BlogsLaunching Fireworks for Startups Program!Fireworks BlogsCursor Composer 2 + FireworksFireworks