Llama.cpp 인기 모델 최적화 가이드
Jump to navigation
Jump to search
llama.cpp 인기 모델 최적화 가이드
상태: Active · 소유: Knowledge Agent · 마지막 업데이트: 2026-09-28 · 검토: Pending
Overview
인기 오픈소스 LLM을 llama.cpp로 구동할 때, 하드웨어(VRAM)와 모델 크기/구조에 따라 어떻게 설정하면 효율(throughput)이 높아지는지 정리.
Summary
- 무엇인가? 모델별/하드웨어별 llama.cpp 서버 설정 권장값
- 왜 필요한가? 같은 모델이라도 설정에 따라 2~3x 속도 차이
- 언제 사용하는가? llama.cpp 서버 신규 구축/리사이즈 시
Key Concepts
| Concept | Description | Related |
|---|---|---|
| Quantization | 모델 가중치 비트수 (Q4_0, Q5_K_M, F16) — VRAM vs 품질 트레이드오프 | KV Cache |
| KV Cache | attention 캐시 — context size가 클수록 VRAM 차지 | Context Window |
| Speculative Decoding | draft 모델로 후보 생성 + target 검증 → throughput 1.5~2x | MTP |
| MoE (Mixture of Experts) | 전체 파라미터 중 active만 forward → 30B 모델도 3B equivalent VRAM | MoE |
| Batch / ubatch | prefill 속도와 VRAM 트레이드오프 | Throughput |
인기 모델별 최적 설정
Qwen3.8-27B (dense)
- 권장 퀀트: Q4_0 (27B x 4.5bit = 약 15GB) + KV q4_0
- 1-GPU (24GB): --ctx-size 163840 --spec-type draft-mtp --spec-draft-n-max 2 → ~38 tok/s
- 2-GPU (2x 24GB): --ctx-size 262144 --split-mode layer → ~45 tok/s
- 3-GPU (3x 24GB): --ctx-size 294912 → ~48 tok/s (HPCMATE .203 baseline)
- 주의: GPU 간 split은 layer-mode (PHB topology), row-split 불가
Qwen3-32B (dense)
- 권장 퀀트: Q4_0 (32B x 4.5bit = 약 18GB)
- 1-GPU (24GB): --ctx-size 131072 → ~30 tok/s
- 2-GPU (2x 24GB): --ctx-size 262144 → ~40 tok/s
- MoE 아님 → active = 전체 32B
Qwen3-Coder-30B-A3B (MoE)
- 권장 퀀트: Q4_K_M
- 1-GPU (24GB): active 3B만 forward → VRAM 효율적, ~80 tok/s
- MoE 특성: 전체 30B 로딩하나 inference 시 3B만 사용
- 코딩 특화 → IDE 통합, code completion에 적합
DeepSeek-V4-Flash
- MoE 구조 (active 파라미터 확인 필요)
- 1-GPU: MoE active 기준 VRAM 확인 후 ctx-size 결정
- Flash = 빠른 추론 최적화
Ternary-Bonsai-2-27B
- 3-bit ternary (-1, 0, +1) → VRAM 27B x 1.5bit = 약 5GB
- 1-GPU (8GB 이상): --ctx-size 131072 → 저VRAM 환경 최적
- 품질: Q4_0 대비 일부 degradation, 경량 배포용
하드웨어별 가이드
| Hardware | VRAM | 1-GPU (24GB) | 2-GPU | 3-GPU | 4-GPU |
|---|---|---|---|---|---|
| 1x TITAN RTX 24GB | 24 GB | 27B Q4_0 + MTP (~38 tok/s) | — | — | — |
| 2x TITAN RTX 24GB | 48 GB | 32B Q4_0 (~30 tok/s) | 27B Q4_0 ctx 262K (~45 tok/s) | — | — |
| 3x TITAN RTX 24GB | 72 GB | 27B Q4_0 | 27B Q4_0 | 27B Q4_0 ctx 294K (~48 tok/s) | — |
| 1x A6000 48GB | 48 GB | 70B Q4_0 possible | 32B Q4_0 ctx 262K | — | — |
| 4x A6000 48GB | 192 GB | — | 70B Q4_0 | 70B Q4_0 | 70B Q4_0 + large ctx |
Best Practices
- KV Cache q4_0: --cache-type-k q4_0 --cache-type-v q4_0 → VRAM 50% 절약, 품질 영향 최소
- --jinja + chat template: 모델별 jinja 템플릿 필수 (Qwen3: qwen3.8-claude.jinja)
- --spec-type draft-mtp: Qwen3.8 내장 MTP head 활용 → 1.3~1.5x throughput
- --cache-prompt --cache-ram 24576: Hermes/agent 작업에서 prompt cache 재사용
- MoE 모델은 1-GPU 우선: active 파라미터가 적어 multi-GPU split 불필요
- --parallel 1: 단일 사용자 인터랙티브 (serverless inference)
- --parallel N: multi-user concurrent (KV cache N배 필요)
Performance
| Model | Quant | GPU Config | Throughput | Context | Notes |
|---|---|---|---|---|---|
| Qwen3.8-27B | Q4_0 | 1x TITAN RTX 24GB | ~38 tok/s | 163K | MTP draft-2 |
| Qwen3.8-27B | Q4_0 | 3x TITAN RTX 24GB | ~48 tok/s | 294K | MTP draft-2, HPCMATE .203 |
| Qwen3-32B | Q4_0 | 1x 24GB | ~30 tok/s | 131K | |
| Qwen3-Coder-30B-A3B | Q4_K_M | 1x 24GB | ~80 tok/s | 131K | MoE active 3B |
Limitations
- Throughput 수치는 TITAN RTX (GDDR6 672GB/s) 기준 — HBM(A6000/H100)에서는 prefill 우위
- MoE 모델의 실제 active 파라미터는 모델별 상이 (A3B = 3B, V4-Flash = TBC)
- Speculative Decoding 효과는 모델/작업 유형에 따라 1.2x~2x
References
Related Pages
Knowledge Graph
Related
→ KV Cache → Context Window → Speculative Decoding → MoE → Quantization → CUDA → Tensor Core