Blogs
FireAttention
Mixtral has recently made waves in the AI community as the first OSS model trained on trillions of tokens to support 'mixture of experts' (MoE), which has promising features to speed up both training and inference.
BlogsFireAttention V2: 12x faster to make Long Contexts practical for Online InferenceFireworks
