Zone-Failure-Resilient OpenSearch® at Uber
Author: Smit Patel (Uber) | Source: Uber Engineering Blog | Published: 2026
한 줄 요약
Uber가 OpenSearch의 native shard allocation awareness와 사내 isolation group 인프라(Odin 기반)를 결합해, zone 전체 장애 + 추가 노드 장애(“zone + 1”)에도 쿼리·인제스천이 멈추지 않는 아키텍처를 구축.
핵심 주장/내용
- ZFR(Zone Failure Resilience)는 비협상 요건 — zone 전체 손실에도 쿼리·인제스천 유지
- Isolation Group: 노드를 failure domain(zone/rack)에 통제된·균등하게 분산하는 논리적 파티셔닝. failure domain 인식, role별 밸런싱(data/cluster manager가 한 domain에 집중 안 됨), stable membership(노드 교체돼도 같은 IG). Uber는 3 IG → 단일 zone 장애가 최대 ~33% 용량만 제거. 물리 zone은 보통 3개 초과이나 IG가 추상화 레이어
- Shard Allocation Awareness: shard copy를 IG에 균등 분산(5 copy면 2,2,1, IG 간 차이 최대 1). 단 IG 뒤 노드 풀이 균형이어야 함(아니면 unassigned)
- Forced Shard Allocation Awareness: 전체 예상 attribute(3 IG)를 명시 → IG 하나가 사라져도 남은 IG에 over-allocate 거부(shard는 unassigned, cluster는 yellow). zone 실패 시 surviving zone의 대규모 rebalancing(disk I/O·CPU·네트워크 폭발, cascading failure) 방지. recovery는 IG 복귀 또는 admin 명시 시
- zone + node 장애 대응: data node는 3 copy(zone 1개 + node 1개 잃어도 3번째 생존). cluster manager는 5개(3개 아님) —
cluster.auto_shrink_voting_configuration으로 zone 장애 후 2개 잃어도 3개 생존, 추가 1개 잃어도 2/3 quorum 유지
주요 수치 / 사실
- 최소 3 shard copy(2 replica) + 5 cluster manager → zone 장애 + 추가 node 장애 견딤(quorum 유지, 데이터 무손실)
- 물리 zone ID 사용 시 yellow state(zone별 노드 수 불균등) → IG가 100% shard assignment + green health 보장, disk skew·hot node 제거
관련 위키
Source: 원문 보기