Blogs

How NetEase Games achieved 30-second LLM cold starts on Kubernetes

Cncf

At NetEase Games, we learned a hard lesson about large language model (LLM) inference in production: elastic compute is only useful if data can move just as fast.

Visit Site

Blogs Cncf

BlogsPolicy-as-Code: Flexible Kubernetes governance with KyvernoCncf BlogsHow to get engineering time back from Kubernetes upgradesCncf BlogsGPU autoscaling on Kubernetes with KEDA: Building an external scalerCncf BlogsGitOps policy-as-code: Securing Kubernetes with Argo CD and KyvernoCncf BlogsKubernetes Security: 2025 Stable Features and 2026 previewCncf BlogsIntroducing Kthena: LLM inference for the cloud native eraCncf BlogsLocal Kubernetes Development Using Minikube and Redis EnterpriseRedis NewsHow Argo CD and OpenShift Enable GitOps for DevelopersThenewstack NewsIntel Unveils Next Generation Neuromorphic Computing ChipThenewstack NewsGoogle Anthos from the Eyes of a Kubernetes DeveloperThenewstack NewsSecurity Considerations for API-Driven Apps Deployed to CloudThenewstack BlogsAWS Enterprise Container Management with PulumiPulumi BlogsSupporting Kubernetes with Faster, Easier Test EnvironmentsPulumi BlogsNeo Integrations: MCP Servers and Cloud CLIsPulumi NewsKAUST and Cerebras Named Gordon Bell Award Finalist for Solving Multi-Dimensional SeismicCerebras BlogsAccelerating GPT-5.6 Sol UltrafastCerebras BlogsDocker for Windows Desktop with KubernetesDocker BlogsAGENTS.md outperforms skills in our agent evalsVercel ResourcesRolling out a new featureVercel BlogsFrom ASR to CSR: Why Conversation Changes EverythingDeepgram