Top 10 오픈소스 AI 모델 동향: Difference between revisions

From HPCWIKI
Jump to navigation Jump to search
(정기 업데이트 (2026-09-28): 벤치마크 데이터 변동 없음 (Open LLM Leaderboard 2025-03-20 동결, LiveCodeBench 윈도우 2024-08~2025-05). Top 10 테이블 3종 모두 유지, 조사일시/last_update/다음 업데이트/데이터 신선도 날짜만 갱신.)
(Fix broken tables: {{!->{| + one cell per line, **->''', md-links->wikitext links, replace redlink Status/TOC templates)
Line 1: Line 1:
= Top 10 오픈소스 AI 모델 동향 =
= Top 10 오픈소스 AI 모델 동향 =


{{Status
'''상태:''' Draft ·  '''소유:''' Knowledge Agent ·  '''마지막 업데이트:''' 2026-09-28 ·  '''검토:''' Pending
|status=Draft
|owner=Knowledge Agent
|last_update=2026-09-28
|review=Pending
}}


{{TOC}}


== Overview ==
== Overview ==
Line 23: Line 17:


* 조사 일자: 2026-09-28 09:08 KST
* 조사 일자: 2026-09-28 09:08 KST
* 데이터 출처: [HuggingFace Open LLM Leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard), [LiveCodeBench](https://huggingface.co/spaces/livecodebench/leaderboard)
* 데이터 출처: [https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard HuggingFace Open LLM Leaderboard], [https://huggingface.co/spaces/livecodebench/leaderboard LiveCodeBench]
* 다음 업데이트 예정: 2026-10-05
* 다음 업데이트 예정: 2026-10-05


Line 38: Line 32:
== Key Concepts ==
== Key Concepts ==


{{! class="wikitable"
{| class="wikitable"
! Concept ! Description ! Related
! Concept
! Description
! Related
|-
|-
| [HuggingFace Open LLM Leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard) | 오픈소스 LLM 종합 성능 벤치마크 | [IFEval](https://github.com/google-research/google-research/tree/master/instruction_following_eval), [BBH](https://github.com/suzgunmirac/BIG-Bench-Hard), [MMLU-PRO](https://github.com/hendrycks/test)
| [https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard HuggingFace Open LLM Leaderboard]
| 오픈소스 LLM 종합 성능 벤치마크
| [https://github.com/google-research/google-research/tree/master/instruction_following_eval IFEval], [https://github.com/suzgunmirac/BIG-Bench-Hard BBH], [https://github.com/hendrycks/test MMLU-PRO]
|-
|-
| [LiveCodeBench](https://huggingface.co/spaces/livecodebench/leaderboard) | 코딩 능력 평가 벤치마크 | [Pass@1](https://github.com/LiveCodeBench/LiveCodeBench)
| [https://huggingface.co/spaces/livecodebench/leaderboard LiveCodeBench]
| 코딩 능력 평가 벤치마크
| [https://github.com/LiveCodeBench/LiveCodeBench Pass@1]
|-
|-
| [IFEval](https://github.com/google-research/google-research/tree/master/instruction_following_eval) | 지시 따르기 평가 | [GitHub](https://github.com/google-research/google-research/tree/master/instruction_following_eval)
| [https://github.com/google-research/google-research/tree/master/instruction_following_eval IFEval]
| 지시 따르기 평가
| [https://github.com/google-research/google-research/tree/master/instruction_following_eval GitHub]
|-
|-
| [BBH](https://github.com/suzgunmirac/BIG-Bench-Hard) | BIG-Bench Hard 벤치마크 | [GitHub](https://github.com/suzgunmirac/BIG-Bench-Hard)
| [https://github.com/suzgunmirac/BIG-Bench-Hard BBH]
| BIG-Bench Hard 벤치마크
| [https://github.com/suzgunmirac/BIG-Bench-Hard GitHub]
|-
|-
| [MMLU-PRO](https://github.com/hendrycks/test) | 다중 선택 지식 평가 | [GitHub](https://github.com/hendrycks/test)
| [https://github.com/hendrycks/test MMLU-PRO]
| 다중 선택 지식 평가
| [https://github.com/hendrycks/test GitHub]
|-
|-
| [MATH](https://github.com/hendrycks/math) | 수학 문제 해결 평가 | [GitHub](https://github.com/hendrycks/math)
| [https://github.com/hendrycks/math MATH]
| 수학 문제 해결 평가
| [https://github.com/hendrycks/math GitHub]
|-
|-
| [GPQA](https://github.com/idavidrein/gpqa) | 과학 문제 평가 | [GitHub](https://github.com/idavidrein/gpqa)
| [https://github.com/idavidrein/gpqa GPQA]
| 과학 문제 평가
| [https://github.com/idavidrein/gpqa GitHub]
|-
|-
| [Pass@1](https://github.com/LiveCodeBench/LiveCodeBench) | 코딩 정답률 지표 | [GitHub](https://github.com/LiveCodeBench/LiveCodeBench)
| [https://github.com/LiveCodeBench/LiveCodeBench Pass@1]
}}
| 코딩 정답률 지표
| [https://github.com/LiveCodeBench/LiveCodeBench GitHub]
|}




Line 63: Line 75:
HuggingFace Open LLM Leaderboard의 Overall Average 점수 기준 Top 10 모델.
HuggingFace Open LLM Leaderboard의 Overall Average 점수 기준 Top 10 모델.


{{! class="wikitable"
{| class="wikitable"
! Rank ! Model ! Params (B) ! Average ⬆️ ! [IFEval](https://github.com/google-research/google-research/tree/master/instruction_following_eval) ! [BBH](https://github.com/suzgunmirac/BIG-Bench-Hard) ! [MATH](https://github.com/hendrycks/math) Lvl 5 ! [GPQA](https://github.com/idavidrein/gpqa) ! [MMLU-PRO](https://github.com/hendrycks/test)
! Rank
! Model
! Params (B)
! Average ⬆️
! [https://github.com/google-research/google-research/tree/master/instruction_following_eval IFEval]
! [https://github.com/suzgunmirac/BIG-Bench-Hard BBH]
! [https://github.com/hendrycks/math MATH] Lvl 5
! [https://github.com/idavidrein/gpqa GPQA]
! [https://github.com/hendrycks/test MMLU-PRO]
|-
|-
| 1 | [https://huggingface.co/MaziyarPanahi/calme-3.2-instruct-78b MaziyarPanahi/calme-3.2-instruct-78b] | 77.965 | 52.08 | 80.6 | 62.6 | 40.3 | 20.4 | 70.0
| 1
| 2 | [https://huggingface.co/MaziyarPanahi/calme-3.1-instruct-78b MaziyarPanahi/calme-3.1-instruct-78b] | 77.965 | 51.29 | 81.4 | 62.4 | 39.3 | 19.5 | 68.7
| [https://huggingface.co/MaziyarPanahi/calme-3.2-instruct-78b MaziyarPanahi/calme-3.2-instruct-78b]
| 3 | [https://huggingface.co/dfurman/CalmeRys-78B-Orpo-v0.1 dfurman/CalmeRys-78B-Orpo-v0.1] | 77.965 | 51.23 | 81.6 | 61.9 | 40.6 | 20.0 | 66.8
| 77.965
| 4 | [https://huggingface.co/MaziyarPanahi/calme-2.4-rys-78b MaziyarPanahi/calme-2.4-rys-78b] | 77.965 | 50.77 | 80.1 | 62.2 | 40.7 | 20.4 | 66.7
| 52.08
| 5 | [https://huggingface.co/huihui-ai/Qwen2.5-72B-Instruct-abliterated huihui-ai/Qwen2.5-72B-Instruct-abliterated] | 72.706 | 48.11 | 85.9 | 60.5 | 60.1 | 19.4 | 50.4
| 80.6
| 6 | [https://huggingface.co/Qwen/Qwen2.5-72B-Instruct Qwen/Qwen2.5-72B-Instruct] | 72.706 | 47.98 | 86.4 | 61.9 | 59.8 | 16.7 | 51.4
| 62.6
| 7 | [https://huggingface.co/MaziyarPanahi/calme-2.1-qwen2.5-72b MaziyarPanahi/calme-2.1-qwen2.5-72b] | 72.7 | 47.86 | 86.6 | 61.7 | 59.1 | 15.1 | 51.3
| 40.3
| 8 | [https://huggingface.co/newsbang/Homer-v1.0-Qwen2.5-72B newsbang/Homer-v1.0-Qwen2.5-72B] | 72.706 | 47.46 | 76.3 | 62.3 | 49.0 | 22.1 | 57.2
| 20.4
| 9 | [https://huggingface.co/ehristoforu/qwen2.5-test-32b-it ehristoforu/qwen2.5-test-32b-it] | 32.764 | 47.37 | 78.9 | 58.3 | 59.7 | 15.2 | 52.9
| 70.0
| 10 | [https://huggingface.co/Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B] | 32.76 | 47.34 | 79.7 | 57.6 | 60.3 | 15.0 | 53.3
| 2
}}
| [https://huggingface.co/MaziyarPanahi/calme-3.1-instruct-78b MaziyarPanahi/calme-3.1-instruct-78b]
| 77.965
| 51.29
| 81.4
| 62.4
| 39.3
| 19.5
| 68.7
| 3
| [https://huggingface.co/dfurman/CalmeRys-78B-Orpo-v0.1 dfurman/CalmeRys-78B-Orpo-v0.1]
| 77.965
| 51.23
| 81.6
| 61.9
| 40.6
| 20.0
| 66.8
| 4
| [https://huggingface.co/MaziyarPanahi/calme-2.4-rys-78b MaziyarPanahi/calme-2.4-rys-78b]
| 77.965
| 50.77
| 80.1
| 62.2
| 40.7
| 20.4
| 66.7
| 5
| [https://huggingface.co/huihui-ai/Qwen2.5-72B-Instruct-abliterated huihui-ai/Qwen2.5-72B-Instruct-abliterated]
| 72.706
| 48.11
| 85.9
| 60.5
| 60.1
| 19.4
| 50.4
| 6
| [https://huggingface.co/Qwen/Qwen2.5-72B-Instruct Qwen/Qwen2.5-72B-Instruct]
| 72.706
| 47.98
| 86.4
| 61.9
| 59.8
| 16.7
| 51.4
| 7
| [https://huggingface.co/MaziyarPanahi/calme-2.1-qwen2.5-72b MaziyarPanahi/calme-2.1-qwen2.5-72b]
| 72.7
| 47.86
| 86.6
| 61.7
| 59.1
| 15.1
| 51.3
| 8
| [https://huggingface.co/newsbang/Homer-v1.0-Qwen2.5-72B newsbang/Homer-v1.0-Qwen2.5-72B]
| 72.706
| 47.46
| 76.3
| 62.3
| 49.0
| 22.1
| 57.2
| 9
| [https://huggingface.co/ehristoforu/qwen2.5-test-32b-it ehristoforu/qwen2.5-test-32b-it]
| 32.764
| 47.37
| 78.9
| 58.3
| 59.7
| 15.2
| 52.9
| 10
| [https://huggingface.co/Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B]
| 32.76
| 47.34
| 79.7
| 57.6
| 60.3
| 15.0
| 53.3
|}


=== General 카테고리 분석 ===
=== General 카테고리 분석 ===


* **78B 파라미터 모델 우세**: Top 4 모델이 모두 78B 파라미터 규모
* '''78B 파라미터 모델 우세''': Top 4 모델이 모두 78B 파라미터 규모
* **Qwen2.5 계열 강세**: 72B 모델 3개 포함 (abliterated, Instruct, calme 파생)
* '''Qwen2.5 계열 강세''': 72B 모델 3개 포함 (abliterated, Instruct, calme 파생)
* **32B 모델 경쟁력**: Top 10에 32B 모델 2개 포함
* '''32B 모델 경쟁력''': Top 10에 32B 모델 2개 포함
* **IFEval 점수**: abliterated 버전이 85.9로 최고, Instruct 버전 86.4
* '''IFEval 점수''': abliterated 버전이 85.9로 최고, Instruct 버전 86.4




Line 90: Line 190:
LiveCodeBench Code Generation 리더보드 기준 Top 10 모델 (평가 기간 2024-08-01 ~ 2025-05-01 기준, 최신 릴리스 모델). Pass@1과 Easy, Medium, Hard 난이도별 점수 포함.
LiveCodeBench Code Generation 리더보드 기준 Top 10 모델 (평가 기간 2024-08-01 ~ 2025-05-01 기준, 최신 릴리스 모델). Pass@1과 Easy, Medium, Hard 난이도별 점수 포함.


{{! class="wikitable"
{| class="wikitable"
! Rank ! Model ! [Pass@1](https://github.com/LiveCodeBench/LiveCodeBench) ! [Easy-Pass@1](https://github.com/LiveCodeBench/LiveCodeBench) ! [Medium-Pass@1](https://github.com/LiveCodeBench/LiveCodeBench) ! [Hard-Pass@1](https://github.com/LiveCodeBench/LiveCodeBench)
! Rank
! Model
! [https://github.com/LiveCodeBench/LiveCodeBench Pass@1]
! [https://github.com/LiveCodeBench/LiveCodeBench Easy-Pass@1]
! [https://github.com/LiveCodeBench/LiveCodeBench Medium-Pass@1]
! [https://github.com/LiveCodeBench/LiveCodeBench Hard-Pass@1]
|-
|-
| 1 | [https://platform.openai.com/docs/api-reference/chat/create O4-Mini (High)] | 80.2 | 99.1 | 89.4 | 63.5
| 1
| [https://platform.openai.com/docs/api-reference/chat/create O4-Mini (High)]
| 80.2
| 99.1
| 89.4
| 63.5
|-
|-
| 2 | [https://platform.openai.com/docs/api-reference/chat/create O3 (High)] | 75.8 | 99.1 | 84.4 | 57.1
| 2
| [https://platform.openai.com/docs/api-reference/chat/create O3 (High)]
| 75.8
| 99.1
| 84.4
| 57.1
|-
|-
| 3 | [https://platform.openai.com/docs/api-reference/chat/create O4-Mini (Medium)] | 74.2 | 98.2 | 86.5 | 52.7
| 3
| [https://platform.openai.com/docs/api-reference/chat/create O4-Mini (Medium)]
| 74.2
| 98.2
| 86.5
| 52.7
|-
|-
| 4 | [https://ai.google.dev/gemini-api/docs/models/gemini Gemini-2.5-Pro-06-05] | 73.6 | 99.1 | 87.2 | 50.2
| 4
| [https://ai.google.dev/gemini-api/docs/models/gemini Gemini-2.5-Pro-06-05]
| 73.6
| 99.1
| 87.2
| 50.2
|-
|-
| 5 | [https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 DeepSeek-R1-0528] | 73.1 | 98.7 | 85.2 | 50.7
| 5
| [https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 DeepSeek-R1-0528]
| 73.1
| 98.7
| 85.2
| 50.7
|-
|-
| 6 | [https://ai.google.dev/gemini-api/docs/models/gemini Gemini-2.5-Pro-05-06] | 71.8 | 98.2 | 82.3 | 50.2
| 6
| [https://ai.google.dev/gemini-api/docs/models/gemini Gemini-2.5-Pro-05-06]
| 71.8
| 98.2
| 82.3
| 50.2
|-
|-
| 7 | [https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-32B EXAONE-4.0-32B] | 70.0 | 98.4 | 82.3 | 46.2
| 7
| [https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-32B EXAONE-4.0-32B]
| 70.0
| 98.4
| 82.3
| 46.2
|-
|-
| 8 | [https://huggingface.co/nvidia/OpenReasoning-Nemotron-32B OpenReasoning-Nemotron-32B] | 69.8 | 98.3 | 81.4 | 46.3
| 8
| [https://huggingface.co/nvidia/OpenReasoning-Nemotron-32B OpenReasoning-Nemotron-32B]
| 69.8
| 98.3
| 81.4
| 46.3
|-
|-
| 9 | [https://platform.openai.com/docs/api-reference/chat/create O3-Mini-2025-01-31 (High)] | 67.4 | 99.1 | 84.4 | 38.4
| 9
| [https://platform.openai.com/docs/api-reference/chat/create O3-Mini-2025-01-31 (High)]
| 67.4
| 99.1
| 84.4
| 38.4
|-
|-
| 10 | [https://huggingface.co/nvidia/OpenCodeReasoning-Nemotron-1.1-32B OpenCodeReasoning-Nemotron-1.1-32B] | 66.8 | 97.9 | 79.6 | 41.1
| 10
}}
| [https://huggingface.co/nvidia/OpenCodeReasoning-Nemotron-1.1-32B OpenCodeReasoning-Nemotron-1.1-32B]
| 66.8
| 97.9
| 79.6
| 41.1
|}


=== Coding 카테고리 분석 ===
=== Coding 카테고리 분석 ===


* **상용 추론 모델 주도**: Top 6이 OpenAI O4-Mini/O3 계열과 Google Gemini 2.5 Pro 등 상용 추론 모델
* '''상용 추론 모델 주도''': Top 6이 OpenAI O4-Mini/O3 계열과 Google Gemini 2.5 Pro 등 상용 추론 모델
* **오픈소스 모델 부상**: DeepSeek-R1-0528(5위), EXAONE-4.0-32B(7위), Nemotron 계열(8위, 10위)이 Top 10에 진입
* '''오픈소스 모델 부상''': DeepSeek-R1-0528(5위), EXAONE-4.0-32B(7위), Nemotron 계열(8위, 10위)이 Top 10에 진입
* **Pass@1 최고점**: O4-Mini (High)이 80.2로 종합 1위
* '''Pass@1 최고점''': O4-Mini (High)이 80.2로 종합 1위
* **경량 추론 모델 경쟁**: 32B급 오픈소스(EXAONE, Nemotron)가 66~70 Pass@1로 상용 모델과 격차 축소
* '''경량 추론 모델 경쟁''': 32B급 오픈소스(EXAONE, Nemotron)가 66~70 Pass@1로 상용 모델과 격차 축소




Line 126: Line 281:
7B 이하 파라미터 모델을 위한 Edge Devices 카테고리 Top 10.
7B 이하 파라미터 모델을 위한 Edge Devices 카테고리 Top 10.


{{! class="wikitable"
{| class="wikitable"
! Rank ! Model ! Params (B) ! Average ⬆️ ! [IFEval](https://github.com/google-research/google-research/tree/master/instruction_following_eval) ! [BBH](https://github.com/suzgunmirac/BIG-Bench-Hard)
! Rank
! Model
! Params (B)
! Average ⬆️
! [https://github.com/google-research/google-research/tree/master/instruction_following_eval IFEval]
! [https://github.com/suzgunmirac/BIG-Bench-Hard BBH]
|-
|-
| 1 | [https://huggingface.co/JungZoona/T3Q-Qwen2.5-14B-Instruct-1M-e3 JungZoona/T3Q-Qwen2.5-14B-Instruct-1M-e3] | 0.0 | 47.09 | 73.2 | 65.5
| 1
| 2 | [https://huggingface.co/Xiaojian9992024/Qwen2.5-Dyanka-7B-Preview Xiaojian9992024/Qwen2.5-Dyanka-7B-Preview] | 7.616 | 37.30 | 76.4 | 36.6
| [https://huggingface.co/JungZoona/T3Q-Qwen2.5-14B-Instruct-1M-e3 JungZoona/T3Q-Qwen2.5-14B-Instruct-1M-e3]
| 3 | [https://huggingface.co/gz987/qwen2.5-7b-cabs-v0.3 gz987/qwen2.5-7b-cabs-v0.3] | 7.616 | 36.94 | 75.7 | 36.0
| 0.0
| 4 | [https://huggingface.co/marcuscedricridia/pre-cursa-o1-v1.2 marcuscedricridia/pre-cursa-o1-v1.2] | 7.613 | 36.89 | 75.5 | 36.1
| 47.09
| 5 | [https://huggingface.co/gz987/qwen2.5-7b-cabs-v0.4 gz987/qwen2.5-7b-cabs-v0.4] | 7.616 | 36.88 | 75.8 | 36.4
| 73.2
| 6 | [https://huggingface.co/suayptalha/Clarus-7B-v0.2 suayptalha/Clarus-7B-v0.2] | 7.613 | 36.86 | 76.8 | 36.0
| 65.5
| 7 | [https://huggingface.co/marcuscedricridia/cursa-o1-7b-v1.2-normalize-false marcuscedricridia/cursa-o1-7b-v1.2-normalize-false] | 7.613 | 36.80 | 76.2 | 36.1
| 2
| 8 | [https://huggingface.co/marcuscedricridia/pre-cursa-o1-v1.6 marcuscedricridia/pre-cursa-o1-v1.6] | 7.613 | 36.80 | 75.3 | 35.9
| [https://huggingface.co/Xiaojian9992024/Qwen2.5-Dyanka-7B-Preview Xiaojian9992024/Qwen2.5-Dyanka-7B-Preview]
| 9 | [https://huggingface.co/suayptalha/Clarus-7B-v0.3 suayptalha/Clarus-7B-v0.3] | 7.616 | 36.78 | 75.1 | 36.5
| 7.616
| 10 | [https://huggingface.co/marcuscedricridia/pre-cursa-o1-v1.3 marcuscedricridia/pre-cursa-o1-v1.3] | 7.613 | 36.71 | 75.1 | 35.5
| 37.30
}}
| 76.4
| 36.6
| 3
| [https://huggingface.co/gz987/qwen2.5-7b-cabs-v0.3 gz987/qwen2.5-7b-cabs-v0.3]
| 7.616
| 36.94
| 75.7
| 36.0
| 4
| [https://huggingface.co/marcuscedricridia/pre-cursa-o1-v1.2 marcuscedricridia/pre-cursa-o1-v1.2]
| 7.613
| 36.89
| 75.5
| 36.1
| 5
| [https://huggingface.co/gz987/qwen2.5-7b-cabs-v0.4 gz987/qwen2.5-7b-cabs-v0.4]
| 7.616
| 36.88
| 75.8
| 36.4
| 6
| [https://huggingface.co/suayptalha/Clarus-7B-v0.2 suayptalha/Clarus-7B-v0.2]
| 7.613
| 36.86
| 76.8
| 36.0
| 7
| [https://huggingface.co/marcuscedricridia/cursa-o1-7b-v1.2-normalize-false marcuscedricridia/cursa-o1-7b-v1.2-normalize-false]
| 7.613
| 36.80
| 76.2
| 36.1
| 8
| [https://huggingface.co/marcuscedricridia/pre-cursa-o1-v1.6 marcuscedricridia/pre-cursa-o1-v1.6]
| 7.613
| 36.80
| 75.3
| 35.9
| 9
| [https://huggingface.co/suayptalha/Clarus-7B-v0.3 suayptalha/Clarus-7B-v0.3]
| 7.616
| 36.78
| 75.1
| 36.5
| 10
| [https://huggingface.co/marcuscedricridia/pre-cursa-o1-v1.3 marcuscedricridia/pre-cursa-o1-v1.3]
| 7.613
| 36.71
| 75.1
| 35.5
|}


=== Edge Devices 카테고리 분석 ===
=== Edge Devices 카테고리 분석 ===


* **Qwen2.5 기반 우세**: Top 5 중 3개 Qwen2.5 7B 파생 모델
* '''Qwen2.5 기반 우세''': Top 5 중 3개 Qwen2.5 7B 파생 모델
* **7B 파라미터 표준**: 대부분의 모델이 7.6B 파라미터
* '''7B 파라미터 표준''': 대부분의 모델이 7.6B 파라미터
* **IFEval 집중**: Edge 모델들은 IFEval 75+ 점수 유지
* '''IFEval 집중''': Edge 모델들은 IFEval 75+ 점수 유지
* **T3Q 특이사항**: 0B 파라미터 표시 (계산 오류 가능성)
* '''T3Q 특이사항''': 0B 파라미터 표시 (계산 오류 가능성)




== Performance Analysis ==
== Performance Analysis ==


{{! class="wikitable"
{| class="wikitable"
! Category ! Top Model ! Score ! Key Insight
! Category
! Top Model
! Score
! Key Insight
|-
|-
| General | calme-3.2-instruct-78b | 52.08 | 78B 파라미터 기준 최고 성능
| General
| calme-3.2-instruct-78b
| 52.08
| 78B 파라미터 기준 최고 성능
|-
|-
| Coding | O4-Mini (High) | 80.2 Pass@1 | 상용 추론 모델 1위, 오픈소스(DeepSeek/EXAONE/Nemotron) 부상
| Coding
| O4-Mini (High)
| 80.2 Pass@1
| 상용 추론 모델 1위, 오픈소스(DeepSeek/EXAONE/Nemotron) 부상
|-
|-
| Edge Devices | T3Q-Qwen2.5-14B | 47.09 | 7B 이하 중 최고 (계산 오류 가능성)
| Edge Devices
}}
| T3Q-Qwen2.5-14B
| 47.09
| 7B 이하 중 최고 (계산 오류 가능성)
|}




== Limitations ==
== Limitations ==


* **General 카테고리**: 78B 이상 대형 모델 중심, 7B 이하 모델 포함 안 됨
* '''General 카테고리''': 78B 이상 대형 모델 중심, 7B 이하 모델 포함 안 됨
* **Coding 카테고리**: 상용 모델만 포함, 오픈소스 모델 부재
* '''Coding 카테고리''': 상용 모델만 포함, 오픈소스 모델 부재
* **Edge Devices**: 일부 모델 파라미터 계산 오류 가능성 (T3Q 0B)
* '''Edge Devices''': 일부 모델 파라미터 계산 오류 가능성 (T3Q 0B)
* **데이터 신선도**: 2026-09-28 재확인 — General/Coding/Edge 데이터 모두 동일 (Open LLM Leaderboard 동결: 2025-03-20, LiveCodeBench 윈도우 2024-08~2025-05)
* '''데이터 신선도''': 2026-09-28 재확인 — General/Coding/Edge 데이터 모두 동일 (Open LLM Leaderboard 동결: 2025-03-20, LiveCodeBench 윈도우 2024-08~2025-05)




== References ==
== References ==


* [HuggingFace Open LLM Leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard)
* [https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard HuggingFace Open LLM Leaderboard]
* [LiveCodeBench Leaderboard](https://huggingface.co/spaces/livecodebench/leaderboard)
* [https://huggingface.co/spaces/livecodebench/leaderboard LiveCodeBench Leaderboard]
* [IFEval GitHub](https://github.com/google-research/google-research/tree/master/instruction_following_eval)
* [https://github.com/google-research/google-research/tree/master/instruction_following_eval IFEval GitHub]
* [BBH GitHub](https://github.com/suzgunmirac/BIG-Bench-Hard)
* [https://github.com/suzgunmirac/BIG-Bench-Hard BBH GitHub]
* [MMLU-PRO GitHub](https://github.com/hendrycks/test)
* [https://github.com/hendrycks/test MMLU-PRO GitHub]
* [MATH GitHub](https://github.com/hendrycks/math)
* [https://github.com/hendrycks/math MATH GitHub]
* [GPQA GitHub](https://github.com/idavidrein/gpqa)
* [https://github.com/idavidrein/gpqa GPQA GitHub]
* [Pass@1 GitHub](https://github.com/LiveCodeBench/LiveCodeBench)
* [https://github.com/LiveCodeBench/LiveCodeBench Pass@1 GitHub]
* [MaziyarPanahi/calme-3.2-instruct-78b Hub](https://huggingface.co/MaziyarPanahi/calme-3.2-instruct-78b)
* [https://huggingface.co/MaziyarPanahi/calme-3.2-instruct-78b MaziyarPanahi/calme-3.2-instruct-78b Hub]
* [Qwen/Qwen2.5-72B-Instruct Hub](https://huggingface.co/Qwen/Qwen2.5-72B-Instruct)
* [https://huggingface.co/Qwen/Qwen2.5-72B-Instruct Qwen/Qwen2.5-72B-Instruct Hub]
* [JungZoona/T3Q-Qwen2.5-14B Hub](https://huggingface.co/JungZoona/T3Q-Qwen2.5-14B-Instruct-1M-e3)
* [https://huggingface.co/JungZoona/T3Q-Qwen2.5-14B-Instruct-1M-e3 JungZoona/T3Q-Qwen2.5-14B Hub]
* [DeepSeek-R1-0528 Hub](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528)
* [https://huggingface.co/deepseek-ai/DeepSeek-R1-0528 DeepSeek-R1-0528 Hub]
* [LGAI-EXAONE/EXAONE-4.0-32B Hub](https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-32B)
* [https://huggingface.co/LGAI-EXAONE/EXAONE-4.0-32B LGAI-EXAONE/EXAONE-4.0-32B Hub]
* [nvidia/OpenReasoning-Nemotron-32B Hub](https://huggingface.co/nvidia/OpenReasoning-Nemotron-32B)
* [https://huggingface.co/nvidia/OpenReasoning-Nemotron-32B nvidia/OpenReasoning-Nemotron-32B Hub]
* [nvidia/OpenCodeReasoning-Nemotron-1.1-32B Hub](https://huggingface.co/nvidia/OpenCodeReasoning-Nemotron-1.1-32B)
* [https://huggingface.co/nvidia/OpenCodeReasoning-Nemotron-1.1-32B nvidia/OpenCodeReasoning-Nemotron-1.1-32B Hub]
* [LiveCodeBench 공식 리더보드](https://livecodebench.github.io/leaderboard.html)
* [https://livecodebench.github.io/leaderboard.html LiveCodeBench 공식 리더보드]




== Related Pages ==
== Related Pages ==


* [HuggingFace Open LLM Leaderboard](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard)
* [https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard HuggingFace Open LLM Leaderboard]
* [LiveCodeBench](https://huggingface.co/spaces/livecodebench/leaderboard)
* [https://huggingface.co/spaces/livecodebench/leaderboard LiveCodeBench]
* [LLM 벤치마크](https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard)
* [https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard LLM 벤치마크]
* [Open Source Models](https://huggingface.co/models)
* [https://huggingface.co/models Open Source Models]


[[Category:AI]]
[[Category:AI]]
[[Category:Reference]]
[[Category:Reference]]

Revision as of 13:32, 28 September 2026

Top 10 오픈소스 AI 모델 동향

상태: Draft · 소유: Knowledge Agent · 마지막 업데이트: 2026-09-28 · 검토: Pending


Overview

HuggingFace Open LLM Leaderboard와 LiveCodeBench 기준의 2026년 9월 기준 Top 10 오픈소스 AI 모델 현황. General, Coding, Edge Devices 세 가지 카테고리로 분류하여 최신 벤치마크 성능을 비교 분석.

Summary

  • 무엇인가? HuggingFace Open LLM Leaderboard와 LiveCodeBench의 최신 Top 10 모델 목록
  • 왜 필요한가? 오픈소스 LLM의 최신 성능 트렌드 파악 및 모델 선택 참고
  • 언제 사용하는가? 모델 평가, 벤치마킹, 기술 리서치 시 참고

조사 정보


Purpose

이 문서가 존재하는 이유

  • Goal: 오픈소스 LLM의 최신 성능 트렌드를 카테고리별로 정리
  • Scope: HuggingFace Open LLM Leaderboard (General, Edge Devices), LiveCodeBench (Coding)
  • Non-goals: 상용 모델 비교, 자체 벤치마크 수행


Key Concepts

Concept Description Related
HuggingFace Open LLM Leaderboard 오픈소스 LLM 종합 성능 벤치마크 IFEval, BBH, MMLU-PRO
LiveCodeBench 코딩 능력 평가 벤치마크 Pass@1
IFEval 지시 따르기 평가 GitHub
BBH BIG-Bench Hard 벤치마크 GitHub
MMLU-PRO 다중 선택 지식 평가 GitHub
MATH 수학 문제 해결 평가 GitHub
GPQA 과학 문제 평가 GitHub
Pass@1 코딩 정답률 지표 GitHub


General 카테고리: HuggingFace Open LLM Leaderboard Overall Average

HuggingFace Open LLM Leaderboard의 Overall Average 점수 기준 Top 10 모델.

Rank Model Params (B) Average ⬆️ IFEval BBH MATH Lvl 5 GPQA MMLU-PRO
1 MaziyarPanahi/calme-3.2-instruct-78b 77.965 52.08 80.6 62.6 40.3 20.4 70.0 2 MaziyarPanahi/calme-3.1-instruct-78b 77.965 51.29 81.4 62.4 39.3 19.5 68.7 3 dfurman/CalmeRys-78B-Orpo-v0.1 77.965 51.23 81.6 61.9 40.6 20.0 66.8 4 MaziyarPanahi/calme-2.4-rys-78b 77.965 50.77 80.1 62.2 40.7 20.4 66.7 5 huihui-ai/Qwen2.5-72B-Instruct-abliterated 72.706 48.11 85.9 60.5 60.1 19.4 50.4 6 Qwen/Qwen2.5-72B-Instruct 72.706 47.98 86.4 61.9 59.8 16.7 51.4 7 MaziyarPanahi/calme-2.1-qwen2.5-72b 72.7 47.86 86.6 61.7 59.1 15.1 51.3 8 newsbang/Homer-v1.0-Qwen2.5-72B 72.706 47.46 76.3 62.3 49.0 22.1 57.2 9 ehristoforu/qwen2.5-test-32b-it 32.764 47.37 78.9 58.3 59.7 15.2 52.9 10 Saxo/Linkbricks-Horizon-AI-Avengers-V1-32B 32.76 47.34 79.7 57.6 60.3 15.0 53.3

General 카테고리 분석

  • 78B 파라미터 모델 우세: Top 4 모델이 모두 78B 파라미터 규모
  • Qwen2.5 계열 강세: 72B 모델 3개 포함 (abliterated, Instruct, calme 파생)
  • 32B 모델 경쟁력: Top 10에 32B 모델 2개 포함
  • IFEval 점수: abliterated 버전이 85.9로 최고, Instruct 버전 86.4


Coding 카테고리: LiveCodeBench Pass@1

LiveCodeBench Code Generation 리더보드 기준 Top 10 모델 (평가 기간 2024-08-01 ~ 2025-05-01 기준, 최신 릴리스 모델). Pass@1과 Easy, Medium, Hard 난이도별 점수 포함.

Rank Model Pass@1 Easy-Pass@1 Medium-Pass@1 Hard-Pass@1
1 O4-Mini (High) 80.2 99.1 89.4 63.5
2 O3 (High) 75.8 99.1 84.4 57.1
3 O4-Mini (Medium) 74.2 98.2 86.5 52.7
4 Gemini-2.5-Pro-06-05 73.6 99.1 87.2 50.2
5 DeepSeek-R1-0528 73.1 98.7 85.2 50.7
6 Gemini-2.5-Pro-05-06 71.8 98.2 82.3 50.2
7 EXAONE-4.0-32B 70.0 98.4 82.3 46.2
8 OpenReasoning-Nemotron-32B 69.8 98.3 81.4 46.3
9 O3-Mini-2025-01-31 (High) 67.4 99.1 84.4 38.4
10 OpenCodeReasoning-Nemotron-1.1-32B 66.8 97.9 79.6 41.1

Coding 카테고리 분석

  • 상용 추론 모델 주도: Top 6이 OpenAI O4-Mini/O3 계열과 Google Gemini 2.5 Pro 등 상용 추론 모델
  • 오픈소스 모델 부상: DeepSeek-R1-0528(5위), EXAONE-4.0-32B(7위), Nemotron 계열(8위, 10위)이 Top 10에 진입
  • Pass@1 최고점: O4-Mini (High)이 80.2로 종합 1위
  • 경량 추론 모델 경쟁: 32B급 오픈소스(EXAONE, Nemotron)가 66~70 Pass@1로 상용 모델과 격차 축소


Edge Devices 카테고리: HuggingFace Edge Devices

7B 이하 파라미터 모델을 위한 Edge Devices 카테고리 Top 10.

Rank Model Params (B) Average ⬆️ IFEval BBH
1 JungZoona/T3Q-Qwen2.5-14B-Instruct-1M-e3 0.0 47.09 73.2 65.5 2 Xiaojian9992024/Qwen2.5-Dyanka-7B-Preview 7.616 37.30 76.4 36.6 3 gz987/qwen2.5-7b-cabs-v0.3 7.616 36.94 75.7 36.0 4 marcuscedricridia/pre-cursa-o1-v1.2 7.613 36.89 75.5 36.1 5 gz987/qwen2.5-7b-cabs-v0.4 7.616 36.88 75.8 36.4 6 suayptalha/Clarus-7B-v0.2 7.613 36.86 76.8 36.0 7 marcuscedricridia/cursa-o1-7b-v1.2-normalize-false 7.613 36.80 76.2 36.1 8 marcuscedricridia/pre-cursa-o1-v1.6 7.613 36.80 75.3 35.9 9 suayptalha/Clarus-7B-v0.3 7.616 36.78 75.1 36.5 10 marcuscedricridia/pre-cursa-o1-v1.3 7.613 36.71 75.1 35.5

Edge Devices 카테고리 분석

  • Qwen2.5 기반 우세: Top 5 중 3개 Qwen2.5 7B 파생 모델
  • 7B 파라미터 표준: 대부분의 모델이 7.6B 파라미터
  • IFEval 집중: Edge 모델들은 IFEval 75+ 점수 유지
  • T3Q 특이사항: 0B 파라미터 표시 (계산 오류 가능성)


Performance Analysis

Category Top Model Score Key Insight
General calme-3.2-instruct-78b 52.08 78B 파라미터 기준 최고 성능
Coding O4-Mini (High) 80.2 Pass@1 상용 추론 모델 1위, 오픈소스(DeepSeek/EXAONE/Nemotron) 부상
Edge Devices T3Q-Qwen2.5-14B 47.09 7B 이하 중 최고 (계산 오류 가능성)


Limitations

  • General 카테고리: 78B 이상 대형 모델 중심, 7B 이하 모델 포함 안 됨
  • Coding 카테고리: 상용 모델만 포함, 오픈소스 모델 부재
  • Edge Devices: 일부 모델 파라미터 계산 오류 가능성 (T3Q 0B)
  • 데이터 신선도: 2026-09-28 재확인 — General/Coding/Edge 데이터 모두 동일 (Open LLM Leaderboard 동결: 2025-03-20, LiveCodeBench 윈도우 2024-08~2025-05)


References


Related Pages