Introducing Arrow UDFs in PySpark: A Faster, Leaner Replacement for Pandas UDFs
Author: Ruifeng Zheng, Yicong Huang (Databricks) | Source: Databricks Blog | Published: 2026-05-20
한 줄 요약
Pandas/Arrow 데이터 변환을 제거하고 Arrow 데이터 위에서 직접 동작하는 Native Arrow UDF(Databricks Runtime 18.0)로, Pandas UDF 대비 ~10% 빠르고 ~40% 메모리 절감 + 복잡 타입 지원 개선.
핵심 주장/내용
- 기존 Pandas UDF의 한계: Pandas/Arrow 변환 시 추가 데이터 복사(NULL 컬럼은 deep copy 유발), 중첩 StructType 등 복잡 타입 미지원
- Native Arrow UDF는 컬럼나 레이아웃을 end-to-end 유지, 불필요한 복사 회피, Arrow 네이티브 compute/memory 모델로 벡터화 처리.
@arrow_udf데코레이터 또는 타입 힌트 포함@udf로 정의 - 변형 지원: Scalar Functions(direct/iterator/iterator-of-tuples 입력 모드, iterator는 모델 로드 같은 일회성 초기화 분할상환), Aggregate Functions(groupBy/Window), Table Functions(UDTF, table-in table-out 컬럼나 변환)
- DataFrame API:
mapInArrow,applyInArrow(grouped),cogroup().applyInArrow()— Pandas 카운터파트와 동일하나 변환 오버헤드 회피 - scalar Python UDF와 인터페이스 정렬 → 익숙한 데코레이터 문법, 마이그레이션은 보통 몇 줄 변경
주요 수치 / 사실
- Arrow UDF가 Pandas UDF 대비 ~10% 빠름, 메모리 ~40% 절감(메모리 프로파일: 135.8 MiB → 71.7 MiB)
- 출력 행 수는 입력 행 수와 일치해야 함(scalar)
- Databricks Runtime 18.0부터 제공
관련 위키
Source: 원문 보기