HPA란
Horizontal Pod Autoscaler는 CPU·메모리 사용률 같은 지표를 기준으로 디플로이먼트의 파드 개수를 자동으로 늘리거나 줄여주는 쿠버네티스 리소스입니다.
사전 조건 — Metrics Server
HPA가 CPU/메모리 사용률을 읽으려면 클러스터에 Metrics Server가 설치되어 있어야 합니다.
kubectl apply -f https://github.com/kubernetes-sigs/metrics-server/releases/latest/download/components.yaml
HPA 리소스 예시
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-deployment
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
평균 CPU 사용률이 70%를 넘으면 파드를 늘리고, 낮아지면 최소 2개까지 줄입니다.
동작 확인
kubectl get hpa web-hpa --watch
주의할 점
파드가 늘어나는 데는 컨테이너 시작 시간만큼 지연이 있습니다. 트래픽 급증이 예상되는 이벤트가 있다면 minReplicas를 미리 올려두는 것이 안전합니다.