Slurm Support
Overview
Slurm Support에 대한 기술 문서입니다.
Summary
- 무엇인가? - Slurm Support
- 왜 필요한가? - HPC 및 서버 환경에서 필수 개념
- 언제 사용하는가? - 서버 구성, 성능 튜닝, 문제 해결 시
Purpose
이 문서가 존재하는 이유
- Goal: Slurm Support에 대한 기술 정보 제공
- Scope: Slurm Support의 개념, 사용법, 설정
- Non-goals: 다른 주제로의 확장
Key Concepts
| Concept | Description | Related |
|---|---|---|
| Slurm Support | HPC/서버 환경에서 중요한 기술 개념 | Linux, Server |
Detailed Explanation
- HPCMATE Co., LTD provides thermal optimized HPC, GPGPU hardware and HPC cluster solutions and covers frontline customer support as a authorized system integration partner of SchedMD.
- SchedMD is the core company behind the Slurm workload manager. The owners and employees of SchedMD have written over 90% of the Slurm code base. SchedMD also reviews and integrates contributions from others, distributes the Slurm code base and maintains the canonical version of Slurm.
HPCMATE and SchedMD offers Level 3 Slurm support, see the “Overview of Slurm Support". HPCMATE covers all level of Slurm support with direct supported by SchedMD especially for Level 3 Slurm support. When your cluster encounters a Level 3 Slurm workload manager issue or bug, then a support contract allows you to submit the issue to the HPCMATE and SchedMD team for successful resolution. When you submit the request we ask for the following information in order to ensure timely issue resolution.
- The steps or script to reproduce to issue or bug.
- A copy of your Slurm configuration files.
After we review the issue the engineering team
- Generates a fix, this can be anything from a configuration change to a patch.
- Validates the fix by conducting regression testing.
- Develops additional test cases, where applicable, as a result of discovery of root cause.
- Communicates steps of action or resolution along with any code changes and testing results to the client.
- Submits the final resolution to the client. and for inclusion in the Slurm tree.
Slurm support also includes configuration assistance for each supported cluster. This assistance is valuable when the cluster is initially being configured to use Slurm or when the cluster needs to be modified as requirements change. For each supported cluster you will be able to review your cluster requirements, operating environment and the organizational goals for the machine with a Slurm engineer and work with the engineer to optimize the configuration to your needs. Slurm support also will allow you to receive detailed answers directly from the Slurm Development team when you encounter complex Slurm questions. High performance computers typically come with expectations of high utilization from the end users and management in order to ensure a strong return on the investment. When bug or complex technical issue is encountered it can take days or even weeks to resolve the issue in house. However, when your system is covered by a support contract you can reach out to the Slurm engineering experts to get help resolving these complex issues. You should expect to make comments like this anytime you get to work with the Slurm engineering team.
- James Cuff, Assistant Dean for Research Computing at Harvard
- Phenomenal #hpc support from the @SchedMD #slurm team.
- Patch within a day (on aSaturday!) In prod this morning!
- https://twitter.com/jamesdotcuff/status/392325342351732736
- Colin McMurtrie, Head of Systems, S... [내용 계속]
Best Practices
- 최신 버전 사용 권장
- 공식 문서 참고
- 테스트 환경에서 먼저 검증
References
Related Pages
Knowledge Graph
Related