Llama.cpp 인기 모델 최적화 가이드

From HPCWIKI
Revision as of 11:22, 29 September 2026 by Clara (talk | contribs) (Create: llama.cpp optimization guide for popular open-source models)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

llama.cpp 인기 모델 최적화 가이드

상태: Active · 소유: Knowledge Agent · 마지막 업데이트: 2026-09-28 · 검토: Pending

Overview

인기 오픈소스 LLM을 llama.cpp로 구동할 때, 하드웨어(VRAM)와 모델 크기/구조에 따라 어떻게 설정하면 효율(throughput)이 높아지는지 정리.

Summary

  • 무엇인가? 모델별/하드웨어별 llama.cpp 서버 설정 권장값
  • 왜 필요한가? 같은 모델이라도 설정에 따라 2~3x 속도 차이
  • 언제 사용하는가? llama.cpp 서버 신규 구축/리사이즈 시

Key Concepts

Concept Description Related
Quantization 모델 가중치 비트수 (Q4_0, Q5_K_M, F16) — VRAM vs 품질 트레이드오프 KV Cache
KV Cache attention 캐시 — context size가 클수록 VRAM 차지 Context Window
Speculative Decoding draft 모델로 후보 생성 + target 검증 → throughput 1.5~2x MTP
MoE (Mixture of Experts) 전체 파라미터 중 active만 forward → 30B 모델도 3B equivalent VRAM MoE
Batch / ubatch prefill 속도와 VRAM 트레이드오프 Throughput

인기 모델별 최적 설정

Qwen3.8-27B (dense)

  • 권장 퀀트: Q4_0 (27B x 4.5bit = 약 15GB) + KV q4_0
  • 1-GPU (24GB): --ctx-size 163840 --spec-type draft-mtp --spec-draft-n-max 2 → ~38 tok/s
  • 2-GPU (2x 24GB): --ctx-size 262144 --split-mode layer → ~45 tok/s
  • 3-GPU (3x 24GB): --ctx-size 294912 → ~48 tok/s (HPCMATE .203 baseline)
  • 주의: GPU 간 split은 layer-mode (PHB topology), row-split 불가

Qwen3-32B (dense)

  • 권장 퀀트: Q4_0 (32B x 4.5bit = 약 18GB)
  • 1-GPU (24GB): --ctx-size 131072 → ~30 tok/s
  • 2-GPU (2x 24GB): --ctx-size 262144 → ~40 tok/s
  • MoE 아님 → active = 전체 32B

Qwen3-Coder-30B-A3B (MoE)

  • 권장 퀀트: Q4_K_M
  • 1-GPU (24GB): active 3B만 forward → VRAM 효율적, ~80 tok/s
  • MoE 특성: 전체 30B 로딩하나 inference 시 3B만 사용
  • 코딩 특화 → IDE 통합, code completion에 적합

DeepSeek-V4-Flash

  • MoE 구조 (active 파라미터 확인 필요)
  • 1-GPU: MoE active 기준 VRAM 확인 후 ctx-size 결정
  • Flash = 빠른 추론 최적화

Ternary-Bonsai-2-27B

  • 3-bit ternary (-1, 0, +1) → VRAM 27B x 1.5bit = 약 5GB
  • 1-GPU (8GB 이상): --ctx-size 131072 → 저VRAM 환경 최적
  • 품질: Q4_0 대비 일부 degradation, 경량 배포용

하드웨어별 가이드

Hardware VRAM 1-GPU (24GB) 2-GPU 3-GPU 4-GPU
1x TITAN RTX 24GB 24 GB 27B Q4_0 + MTP (~38 tok/s) — — —
2x TITAN RTX 24GB 48 GB 32B Q4_0 (~30 tok/s) 27B Q4_0 ctx 262K (~45 tok/s) — —
3x TITAN RTX 24GB 72 GB 27B Q4_0 27B Q4_0 27B Q4_0 ctx 294K (~48 tok/s) —
1x A6000 48GB 48 GB 70B Q4_0 possible 32B Q4_0 ctx 262K — —
4x A6000 48GB 192 GB — 70B Q4_0 70B Q4_0 70B Q4_0 + large ctx

Best Practices

  • KV Cache q4_0: --cache-type-k q4_0 --cache-type-v q4_0 → VRAM 50% 절약, 품질 영향 최소
  • --jinja + chat template: 모델별 jinja 템플릿 필수 (Qwen3: qwen3.8-claude.jinja)
  • --spec-type draft-mtp: Qwen3.8 내장 MTP head 활용 → 1.3~1.5x throughput
  • --cache-prompt --cache-ram 24576: Hermes/agent 작업에서 prompt cache 재사용
  • MoE 모델은 1-GPU 우선: active 파라미터가 적어 multi-GPU split 불필요
  • --parallel 1: 단일 사용자 인터랙티브 (serverless inference)
  • --parallel N: multi-user concurrent (KV cache N배 필요)

Performance

Model Quant GPU Config Throughput Context Notes
Qwen3.8-27B Q4_0 1x TITAN RTX 24GB ~38 tok/s 163K MTP draft-2
Qwen3.8-27B Q4_0 3x TITAN RTX 24GB ~48 tok/s 294K MTP draft-2, HPCMATE .203
Qwen3-32B Q4_0 1x 24GB ~30 tok/s 131K
Qwen3-Coder-30B-A3B Q4_K_M 1x 24GB ~80 tok/s 131K MoE active 3B

Limitations

  • Throughput 수치는 TITAN RTX (GDDR6 672GB/s) 기준 — HBM(A6000/H100)에서는 prefill 우위
  • MoE 모델의 실제 active 파라미터는 모델별 상이 (A3B = 3B, V4-Flash = TBC)
  • Speculative Decoding 효과는 모델/작업 유형에 따라 1.2x~2x

References

Related Pages

Knowledge Graph

Related

→ KV Cache → Context Window → Speculative Decoding → MoE → Quantization → CUDA → Tensor Core