Particle.news

Tencent Expands Hy4 Inference Capacity After WorkBuddy Launch Strains System

Tencent is scaling inference clusters and imposing daily free quotas to keep access to a much larger open-source model that requires scarce high-end compute.

Overview

  • The Hy4 preview drew heavy traffic when it went live on WorkBuddy, causing task queues and elevated call volume that disrupted some user requests.
  • Tencent said it performed emergency expansion and will dynamically reallocate Hy4 inference clusters to ease congestion but warned peak-period queues may still occur because high-end compute is limited.
  • The company is offering the Hy4 preview as a limited-time free trial with per-user daily quotas to ration access during the scaling period.
  • Tencent extended Hy3 free access on WorkBuddy as a fallback for users who encounter Hy4 queues and recommended switching models or using off-peak hours for non-urgent work.
  • Hy4 is a text-only large language model with 770 billion total parameters, 49 billion active parameters and a 1-million-token context window, and WorkBuddy routes image or video tasks to other multimodal models that consume credits under normal billing rules.