전자전 환경에서 DoA 추정을 활용한 계층적 심층 강화학습 기반 최적 재밍 알고리즘 연구
An Optimal Jamming Algorithm based on Hierarchical Deep Reinforcement Learning using DoA Estimation in an Electronic Warfare Environment
- 주제(키워드) 신호 대 재밍 잡음 비 , 무인 항공기 , 계층적 심층 강화학습 , 도래 방향
- 주제(DDC) 004.6
- 발행기관 아주대학교 일반대학원
- 지도교수 이호원
- 발행년도 2026
- 학위수여년월 2026. 8
- 학위명 석사
- 학과 및 전공 일반대학원 AI융합네트워크학과
- 실제URI http://www.dcollection.net/handler/ajou/000000036626
- 본문언어 한국어
- 저작권 아주대학교 논문은 저작권에 의해 보호받습니다.
초록/요약
최근 5G와 인공지능의 발전으로 무선 기기가 빠르게 확산되고 모바일 애플리케이션이 폭발적으로 증가하면서, 무선 서비스는 일상생활과 사회 통신 인프라의 핵심 요소로 자리 잡고 있다. 무선 서비스 의존도가 높아 짐에 따라 보안 위협이 심각한 문제로 대두되고 있으며, 재밍은 적의 통 신을 교란하여 정상적인 송수신을 방해하는 대표적 방식으로 주목받고 있 다. 재밍을 수행하는 방법 중 지상 재머를 활용하는 방법은 비용 효율적 일 수 있지만 지상의 고정된 위치에 배치되어 대상이 방해 범위 내에 있 을 때만 효과적이라는 구조적 한계를 갖는다. 무인항공기(unmanned aerial vehicle, UAV)는 다양한 수준의 자율성을 갖춘 항공기로 위치와 고도를 유연하게 조정할 수 있기 때문에 빠르고 비 용 효율적인 배치, 이동성, 높은 line-of-sight(LoS) 확률로 최근 무선 통 신 분야에서 상당한 주목을 받고 있다. 특히, 3차원 전장 환경에서 높은 기동성과 유연성을 기반으로 UAV를 통한 재밍은 기존의 고정 위치 기반 재밍보다 효율적이고 유연한 운용이 가능하다. 그러나 UAV에 탑재된 배 터리 용량은 제한적이므로 지속적으로 높은 전력을 사용하기에는 한계가 있어 재밍 성능과 전력 소모 간에 존재하는 절충 관계를 고려하여 재밍에 사용할 전력과 빔 폭을 결정해야 한다. 이에 본 논문에서는 재머 UAV의 위치, 빔 폭, 전송 전력 제어를 동적 으로 수행할 수 있도록 계층적 구조의 심층 강화학습(hierarchical Deep reinforcement learning, HDRL) 기반 재밍 기법을 제안한다. 각 UAV는 독립적인 에이전트로 존재하며, 상태를 기반으로 최적의 재밍 기법을 학 습한다. 또한, 재밍 성능과 에너지 소모를 동시에 고려하며, 동적인 환경 에 적응적으로 동작할 수 있도록 가중치 기반 다목적 보상 함수를 설계하 여 학습을 진행한다. 기존 재밍 연구는 대체로 적이 고정된 위치에 있다고 가정하였으나, 실 제 운용 환경에서는 적의 위치를 정확히 알기 어렵기 때문에 도래 방향 (direction-of-arrival, DoA) 추정을 활용할 수 있다. 본 논문은 이를 완화 하기 위해 DoA 추정을 도입하며, 고복잡도의 다중 신호 분류(multiple signal classification, MUSIC) 대신 저지연, 저복잡도의 상관 간섭계 (correlative interferometer, CI)를 활용하여 이동 표적의 방향을 추정하고, 이를 계층적 심층 강화학습과 결합해 근접 재밍 기법을 설계한다. 시뮬레이션 결과, 제안 방안은 완전 탐색 기반 기법(Exhaustive Search), 무작위 행동 기법(Random Action), 계층 구조가 없는 심층 강화 학습 기법(Non-Hierarchical DRL), 실제 위치 활용 기법(HDRL with Real Location), DoA 추정을 수행하지 않는 계층적 심층 강화학습 기법 (HDRL w/o DoA Estimation)의 여러 비교 알고리즘과 비교하였을 때, 적의 통신 품질을 효과적으로 저하함과 동시에 재머의 전력 소모를 낮추 는 데 있어 우수한 성능을 보였다. 특히, 계층적 구조와 DoA 추정 모듈을 결합함으로써 학습 수렴 과정의 안정성이 크게 향상됨을 확인하였다. 이 를 통해, 제안 방안이 적의 통신 품질을 효과적으로 저하하면서도 전력 소모를 최소화하는 방향으로 안정적인 학습을 수행함을 검증하였다.
more초록/요약
With the recent advances in 5G and artificial intelligence, wireless services have become a key component of everyday life and social communication infrastructure, driven by the rapid proliferation of wireless devices and the explosive growth of mobile applications. As dependence on wireless services increases, security threats are emerging as a serious concern, and jamming is attracting attention as a representative method for disrupting enemy communications and preventing normal transmission and reception. Among various jamming strategies, the use of ground-based jammers can be cost-effective, but it has a structural limitation in that it is effective only when the jammer is placed at a fixed ground location and the target remains within its jamming range. Unmanned aerial vehicles(UAVs) are aircraft with varying levels of autonomy that can flexibly adjust their positions and altitudes. They have recently attracted considerable attention in wireless communications due to their fast and cost-effective deployment, mobility, and high line-of-sight(LoS) probability. In particular, in a 3D battlefield environment, UAV-based jamming can be operated more efficiently and flexibly than conventional fixed-position jamming. However, since the onboard battery capacity of a UAV is limited, there are inherent constraints on continuous high-power operation. Therefore, the transmit power and beamwidth for jamming must be carefully determined by considering the trade-off between jamming performance and power consumption. To address this challenge, this paper proposes a hierarchical deep reinforcement learning(HDRL)-based jamming technique that dynamically controls the movement, beamwidth, and transmit power of jammer UAVs. Each UAV is modeled as an independent agent and learns an optimal jamming policy based on its observed state. In addition, both jamming performance and energy consumption are jointly considered by designing a weighted multi-objective reward function, enabling the agent to operate adaptively in a dynamic environment. The simulation results show that the proposed scheme achieves superior performance in degrading the malicious’s communication quality while reducing the jammer’s power consumption, compared to several benchmark algorithms, including the Exhaustive Search scheme, the Random Action scheme, the Non-Hierarchical DRL scheme, the HDRL with Real Location scheme that assumes ideal knowledge of the true positions, and the HDRL w/o DoA Estimation scheme that does not perform DoA estimation. In particular, combining the hierarchical structure with DoA estimation significantly improves the stability of the learning convergence process. These results confirm that the proposed scheme is capable of performing stable learning in a manner that effectively degrades the malicious’s communication quality while minimizing power consumption.
more목차
제1장. 서 론 1
1.1 배경 지식 1
1.2 논문의 개요 및 목적 3
1.3 논문 구성 4
제2장. 관련 연구 동향 6
2.1 계층적 구조를 갖는 강화학습 기법 연구 6
2.2 UAV 기반 재밍 기법 연구 8
2.3 빔 기반 재밍 기법 연구 10
제3장. 전자전 환경에서의 UAV 기반 재밍 시스템 및 채널 모델 12
3.1 네트워크 구성 및 시스템 개요 12
3.2 Air-to-Ground(A2G) 채널 모델 14
3.3 Air-to-Air(A2A) 채널 모델 15
3.4 RSSI 기반 거리 추정 모델 16
3.5 DoA 추정 모델 16
3.6 Signal-to-Jamming-plus-Noise Ratio 정의 21
제4장. 전자전 환경에서 DoA 추정을 활용한 계층적 심층 강화학습 기반 최적 재밍 알고리즘 24
4.1 재머 UAV의 MDP 설계 24
4.2 DQN 학습 과정 26
제5장. 시뮬레이션 결과 및 성능 분석 31
5.1 시뮬레이션 환경 31
5.2 비교 알고리즘 31
5.3 시뮬레이션 결과 및 분석 33
제6장. 결 론 68
참고문헌 70
논문요약 75

