Blogs
GLM-4.7 on Cerebras: Record-Speed AI Inference
See how Z.ai GLM-4.7 runs at approximately 1,000 tokens per second on Cerebras for coding, tool use, and responsive agentic AI workflows.
NewsCerebras Systems Announces Filing of Registration Statement for Proposed Initial Public OfferingCerebras
Resourcesvercel projectVercel
ResourcesDocumentation - JS Projects Utilizing TypeScriptTypescriptlang
NewsConfluent ‘Proactive Support’ Aims to Speed Resolution of Kafka Streaming Data IssuesThenewstack
BlogsBuild vs buy AI agents: why custom solutions win long-termRetool
BlogsHow we used evals and inference-time compute scaling to generate beautiful QR codes that actuallyModal
BlogsTry GLM-5.1, the new frontier of open intelligence, on Modal | Modal BlogModal
BlogsCerebras April HighlightsCerebras
NewsCerebras Systems Launches “Cerebras for Nations” -- A Global Initiative to Accelerate and ScaleCerebras
BlogsCerebras 2024 Predictions for Generative AI, LLMs, and HPCCerebras
BlogsExtending LLM context with 99% less training tokens - CerebrasCerebras
BlogsThe Economics of AI ReasoningCerebras
