검색 상세

RNS-CKKS의 리스케일 에러 최소화를 위한 최적 모듈러스 선택 기반 메모리 효율적 NTT 가속기 설계

A Memory-Efficient NTT Accelerator with Optimized Moduli Selection for Low-Rescaling-Error in RNS-CKKS

초록/요약

본 논문에서는 완전동형암호(Fully Homomorphic Encryption, FHE) 알고리즘을 대표하는 RNS-CKKS 에서 리스케일 에러 최소화에 기여할 수 있는 모듈러스 선택법에 기반한 메모리 효율적 Number Theoretic Transform(NTT) 가속기 구조를 제안한다. 제안하는 NTT 가속기는 Virtex UltraScale+ FPGA 디바이스에 구현하였으며, 세 가지 주요 특성을 갖는다. 첫째, 선택한 최적 모듈러스의 특성을 활용한 하드웨어 설계로 모듈러 곱셈기의 성능 및 면적 효율을 향상하였다. 둘째, 메모리 효율적인 실시간 회전 인자(Twiddle Factor, TF) 생성기 설계에 기반하여 성능 저하 없이 NTT 가속기 전체의 메모리 사용량을 효과적으로 절감하였다. 셋째, 전체 NTT 가속기 구조의 면적-성능 최적화 설계를 통해 단일 하드웨어 구조에서 NTT 와 역 NTT(Inverse NTT, INTT) 연산을 모두 지원하면서 동시에 263 MHz 의 높은 동작 주파수에서 구동할 수 있다. 결론적으로, 구현된 NTT 가속기는 선행 연구 대비 2.85 배 높은 Throughput-Per-Slice(TPS)와 90%의 BRAM(Block RAM) 사용량 절감을 달성하였다. 아울러, 구현된 NTT 가속기는 선택된 최적 모듈러스 지원에 특화되어 RNS- CKKS 알고리즘의 리스케일 에러를 줄이는 데 기여할 수 있으며, 따라서 제안하는 NTT 가속기를 활용할 경우 RNS-CKKS 알고리즘에 기반한 응용 시스템 전반의 정확도 향상을 기대할 수 있다. 주제어 — 완전동형암호, RNS-CKKS, 리스케일 에러 최소화, Number Theoretic Transform

more

초록/요약

This paper proposes a memory-efficient Number Theoretic Transform(NTT) accelerator based on optimized moduli selection that can contribute to minimizing cumulative rescaling errors in the RNS-CKKS algorithm. The proposed NTT accelerator is implemented on a Virtex UltraScale+ FPGA device, demonstrating three key features. First, the performance and area efficiency of the modular multiplier are improved through a hardware design that exploits the characteristics of the selected optimal moduli. Second, the overall memory overhead of the NTT accelerator is effectively reduced without performance degradation by employing a memoryefficient on-the-fly Twiddle Factor(TF) generator. Third, through designing an area- and performance-optimized architecture, the proposed NTT accelerator supports both NTT and inverse NTT(INTT) operations within a single hardware architecture while operating at a high clock frequency of 263 MHz. As a result, the implemented NTT accelerator achieves 2.85x higher Throughput- Per-Slice(TPS) and a 90% reduction in BRAM(Block RAM) usage compared with state-of-the-arts. Consequently, the proposed NTT accelerator can further improve the overall accuracy of systems based on the RNS-CKKS with its optimized hardware structure.

more

목차

Ⅰ 서론 1
Ⅱ 본론 5
1. 제안하는 RNS-CKKS 특화 NTT 가속기의 설계 5
1.1. RNS-CKKS의 정확도 향상을 위한 모듈러스 선택 5
1.2. NTT 가속기의 전체 구조 9
1.3. 모듈러 곱셈기와 재구성 가능한 butterfly 유닛 설계 10
1.4. 메모리 효율적인 실시간 twiddle factor 생성 기법 17
2. 결과 29
2.1. 최적 모듈러스 적용에 따른 RNS-CKKS 연산 에러 검증 29
2.2. NTT 가속기 하드웨어 구현 및 선행 연구와의 비교 31
Ⅲ 결론 38
참고문헌 39
Abstract 44

more