中国电子学会电子制造与封装技术分会会刊

中国半导体行业协会封测分会会刊

无锡市集成电路学会会刊

导航

电子与封装

• 电路与系统 •    下一篇

GPU集群并行计算系统研究及设计

魏鲜明,孙晓冬,卞玲艳,潘言心,袁盛俊   

  1. 中国电子科技集团公司第五十八研究所,江苏 无锡  214035
  • 收稿日期:2026-04-23 修回日期:2026-06-12 出版日期:2026-07-15 发布日期:2026-07-15
  • 通讯作者: 魏鲜明
  • 基金资助:
    国家重点研发计划(2024YFB4505301)

Research and Design on GPU Cluster Parallel Computing System

WEI Xianming, SUN Xiaodong, BIAN Lingyan, PAN Yanxin, YUAN Shengjun   

  1. China Electronics Technology Group Corporation No.58 Research Institute, Wuxi 214035, China
  • Received:2026-04-23 Revised:2026-06-12 Online:2026-07-15 Published:2026-07-15

摘要: 基于人工智能应用和神经网络大模型参数规模和训练数据呈指数级增长的需求,本文设计了一种新形态的通用GPU集群并行计算系统。与传统的Clos通用分布式和超节点机柜式GPU集群不同,提出了一种基于晶圆级基板的新型并行计算系统。在分析大模型并行技术的基础上,提出了GPU集群的扩展互连技术,设计了网络整体结构,完成了并行计算系统的基础模块和系统搭建,最后对比分析了该系统与当前主流GPU多卡模组之间的优缺点。评估结果表明,该并行计算系统算力密度和功率密度分别达到主流商用多卡模组的1.34倍、4.59倍,能效比提升16%,可以满足未来3~5年内人工智能应用方面不断增长的数据和算力需求。

关键词: GPU集群, 并行计算, 人工智能, 扩展互连, 网络结构

Abstract: Driven by the exponential growth of parameter scale and training data in artificial intelligence applications and large-scale neural networks, we design a novel general-purpose GPU cluster parallel computing system. Distinct from traditional Clos general distributed and supernode rack-mounted GPU clusters, we propose a new parallel computing system based on wafer-level substrate. Based on the analysis of large model parallelism technologies, an extended interconnection technique for GPU clusters is presented, the overall network architecture is designed, and the basic modules and system construction of the parallel computing system are completed. Finally, a comparative analysis is conducted on the advantages and disadvantages between the proposed system and current mainstream GPU multi-card modules. The evaluation results demonstrate that the computational density and power density of this parallel computing system reach 1.34 times and 4.59 times of those of mainstream commercial multi-GPU modules, respectively, with the energy efficiency ratio increased by 16%. The proposed system is capable of satisfying the continuously growing demands for data processing and computing power of artificial intelligence applications in the next three to five years.

Key words: GPU cluster, parallel computing, artificial intelligence, extended interconnection, network architecture