Enterprise Infrastructure Intelligence ● Certified engineers online · Fast response guaranteed
클라우드/DevOps

Kubernetes HPA로 오토스케일링 설정하기

트래픽이 몰릴 때 파드를 자동으로 늘리고 싶다면 HPA(Horizontal Pod Autoscaler)가 정답입니다. 기본 개념과 설정 예시를 정리했습니다.

2026.08.18  ·  35회  · 

HPA란

Horizontal Pod Autoscaler는 CPU·메모리 사용률 같은 지표를 기준으로 디플로이먼트의 파드 개수를 자동으로 늘리거나 줄여주는 쿠버네티스 리소스입니다.

사전 조건 — Metrics Server

HPA가 CPU/메모리 사용률을 읽으려면 클러스터에 Metrics Server가 설치되어 있어야 합니다.

kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml

HPA 리소스 예시

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: web-hpa
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web-deployment
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70

평균 CPU 사용률이 70%를 넘으면 파드를 늘리고, 낮아지면 최소 2개까지 줄입니다.

동작 확인

kubectl get hpa web-hpa --watch

주의할 점

파드가 늘어나는 데는 컨테이너 시작 시간만큼 지연이 있습니다. 트래픽 급증이 예상되는 이벤트가 있다면 minReplicas를 미리 올려두는 것이 안전합니다.

WIKIDATA WORKSTATION
AI·렌더링에 최적화된
전문가용 워크스테이션
NVIDIA RTX GPU · 최대 192GB 메모리 · ECC 지원