Ultra Accelerator Link (UALink): Difference between revisions

From HPCWIKI
Jump to navigation Jump to search
(Added Knowledge Graph section (Phase 7.1))
(Fix: remove --- horizontal lines (7 removed))
 
Line 18: Line 18:
* 언제 사용하는가? - 서버 구성, 성능 튜닝, 문제 해결 시
* 언제 사용하는가? - 서버 구성, 성능 튜닝, 문제 해결 시


---


== Purpose ==
== Purpose ==
Line 28: Line 27:
* Non-goals: 다른 주제로의 확장
* Non-goals: 다른 주제로의 확장


---


== Key Concepts ==
== Key Concepts ==
Line 42: Line 40:
|}
|}


---


== Detailed Explanation ==
== Detailed Explanation ==
Line 54: Line 51:
<references />
<references />


---


== Best Practices ==
== Best Practices ==
Line 62: Line 58:
* 테스트 환경에서 먼저 검증
* 테스트 환경에서 먼저 검증


---


== References ==
== References ==
Line 68: Line 63:
* [https://wiki.hpcmate.com Ultra Accelerator Link (UALink)]
* [https://wiki.hpcmate.com Ultra Accelerator Link (UALink)]


---


== Related Pages ==
== Related Pages ==
Line 77: Line 71:
* [[Network]]
* [[Network]]


---


[[Category:Server]]
[[Category:Server]]

Latest revision as of 11:31, 17 July 2026

Template:Status

Template:TOC

Overview

Ultra Accelerator Link (UALink)에 대한 기술 문서입니다.

Summary

  • 무엇인가? - Ultra Accelerator Link (UALink)
  • 왜 필요한가? - HPC 및 서버 환경에서 필수 개념
  • 언제 사용하는가? - 서버 구성, 성능 튜닝, 문제 해결 시


Purpose

이 문서가 존재하는 이유

  • Goal: Ultra Accelerator Link (UALink)에 대한 기술 정보 제공
  • Scope: Ultra Accelerator Link (UALink)의 개념, 사용법, 설정
  • Non-goals: 다른 주제로의 확장


Key Concepts

Concept Description Related
Ultra Accelerator Link (UALink) HPC/서버 환경에서 중요한 기술 개념 Linux, Server


Detailed Explanation

AMD, Broadcom, Google, Intel, Meta, and Microsoft all develop their own AI accelerators (well, Broadcom designs them for Google), Cisco produces networking chips for AI, while HPE builds servers. These companies are interested in standardizing as much infrastructure for their chips as possible, which is why they are teaming up to develop a new industry standard dedicated to advancing high-speed and low-latency communication for scale-up AI Accelerators. Called the Ultra Accelerator Link (UALink)[1]

UALink pod - Source UALink Consortium

The UALink initiative is designed to create an open standard for AI accelerators to communicate more efficiently. The first UALink specification, version 1.0, will enable the connection of up to 1,024 accelerators within an AI computing pod in a reliable, scalable, low-latency network. This specification allows for direct data transfers between the memory attached to accelerators, such as AMD's Instinct GPUs or specialized processors like Intel's Gaudi, enhancing performance and efficiency in AI compute. By standardizing the open interconnect for AI and HPC accelerators, it will be easier for system OEMs, IT professionals, and system integrators to integrate and scale AI systems in datacenters. The standard aims to promote an open ecosystem and facilitate the development of large-scale AI and HPC solutions.[2] As open industry standard, UALink should help bring it to market faster as there will be less IP to haggle over, but an optimistic 2026 release still seems rather far off, given the need for massive AI GPU matrix engines yesterday. UALink is expecting to be an industry wide open standard that to eleminate Nvidia only closed NVlink dependency. Nvidia's NVLink, a GPU-to-GPU connection, can transfer data at 1.8 terabytes per second between GPUs as of Jun'24. There is also an NVLink rack-level Switch capable of supporting up to 576 fully connected GPUs in a non-blocking compute fabric. GPUs connected via NVLink are called “pods” to indicate they have their own data and computational domain.


Best Practices

  • 최신 버전 사용 권장
  • 공식 문서 참고
  • 테스트 환경에서 먼저 검증


References


Related Pages

Knowledge Graph

Related

AMD CPUsAMD GPUsNVLinkPCIeServer