검색 상세

군사환경에서 객체 탐지 성능 극대화를 위한 데이터 라벨링 비교 연구

초록/요약

본 연구는 무인 무기 체계 및 감시정찰(ISR) 분야의 핵심 전술 과제인 '정밀 타격(Precision Strike)'과 '정밀 조준(Precision Aiming)' 성능을 극대화하기 위해, 군사 위장 환경에서 데이터 라벨링 방식(Bounding Box vs. Polygon)이 객체 탐지 및 세그멘테이션 모델의 추론 정확도와 중앙점 탐지 정밀도에 미치는 영향을 실증적으로 분석하였다. 현대전의 군사 도메인 시각 데이 터는 적의 은엄폐를 위한 정교한 위장 패턴과 복잡한 지형 배경으로 인해 객체와 배경의 경계 가 모호하다는 특수성을 지닌다. 기존의 사각형 Bounding Box 라벨링 방식은 객체 외곽의 과 도한 전장 배경 노이즈(수목, 암석, 포구 연막 등)를 필연적으로 포함하게 되며, 이는 모델 학 습 및 추론 시 실제 타격 영역(차체나 신체 중심)으로부터 중심 좌표를 이탈시키는 '중앙점 편 향(Center Shift)' 문제를 야기하여 원거리 조준 실패의 근본 원인이 된다. 본 연구는 이러한 한계를 극복하기 위해, 픽셀 단위로 배경을 배제하고 그린 정리를 통해 순수 객체의 무게중심 (Centroid)을 보존하는 'Polygon 라벨링' 기반 세그멘테이션의 기술적 필연성을 가설로 수립하 였다. 이를 검증하기 위해 한국형 전장 노이즈(수목, 설원, 사막 도메인)를 반영한 군인 및 전 차 클래스의 군사 특화 데이터셋을 직접 구축하였다. 실험 대조군으로는 최신 앵커 프리 (Anchor-free) 모델(YOLOv12-S, FCOS, CenterNet)과 2단계 검출기 등 Bounding Box 계열 모 델 5종, 그리고 Mask R-CNN, SAM 등 세그멘테이션 계열 모델 3종을 포함한 총 10종의 최신 아키텍처를 설정하였다. 또한, 전통적인 mAP@50 지표와 더불어 조준 정밀도를 정량 평가하기 위해 실제 무게중심과 예측 중앙점 간의 유클리드 거리를 정규화한 '중앙점 오차율(Center Distance Error)' 지표를 새로 제안하여 교차 검증을 수행하였다. 실험 결과, 모든 비교 모델의 mAP@50 성능은 유사한 수준을 기록하였으나, 본 연구의 핵심 지표인 중앙점 오차율에서 극명 한 기술적 격차가 나타났다. Polygon 라벨링 기반의 Mask R-CNN은 군사 환경의 극심한 잡음 속에서도 5.88%라는 가장 낮은 중앙점 오차율을 달성하여 정밀 조준 환경에서의 압도적인 우 위를 입증하였다. 반면, 단일 포인트 히트맵에 의존하는 CenterNet은 군사 위장 특유의 저대비 텍스처로 인해 히트맵 붕괴 현상을 겪으며 mAP 0.1810 및 소형 객체 탐지 전면 실패(Small AP 0.0000)라는 심각한 구조적 취약성을 노출하였다. 픽셀 단위 Dense 거리를 회귀하는 FCOS 는 mAP 0.4600으로 CenterNet 대비 선방하였으나, 장단비가 극단적인 자주포 장비나 사격 시 발생하는 비정형 포구 연막 상황에서 사각형 박스 출력 구조의 한계로 인해 조준점이 허공으 로 치우치는 편향(중앙점 오차율 7.78%) 문제를 극복하지 못했다. 추가적으로 최신 YOLOv12 및 YOLO-seg 모델은 분포 회귀(DFL) 메커니즘을 통해 각각 7.10%와 7.00%의 비교적 강건한 조준점 고정 능력을 입증하였다. 본 연구는 전통적인 mAP 지표 중심의 평가 체계에서 벗어 나 화력 제어의 실제 타격 신뢰성을 보장할 수 있는 '중앙점 오차율' 지표의 국방 작전요구성 능(ROC) 표준화 정책을 제안한다. 나아가 전장 극한 조건(설한지 위장복에 의한 미탐지, 사격 화염에 의한 오탐지) 하에서의 단일 광학 모달리티의 한계를 논증하고, 이를 극복하기 위해 가 시광 및 열화상(EO/IR) 센서의 특징 맵을 유기적으로 결합하는 차세대 복합 센서 융합 하이브 리드 아키텍처를 기술적 이정표로 제시한다는 점에서 독보적인 학술적·전술적 의의를 가진다. 주제어(Keywords): 군사 객체 탐지, 데이터 라벨링, Bounding Box, Polygon, 인스턴스 세그멘 테이션, 중앙점 오차율(Center Distance Error), 정밀 타격, 센서 융합

more

초록/요약

To maximize the performance of "Precision Strike" and "Precision Aiming," which are core tactical pillars of unmanned weapon systems and Intelligence, Surveillance, and Reconnaissance (ISR), this study empirically analyzes the effects of data labeling methods (Bounding Box vs. Polygon) on the inference accuracy and center point detection precision of object detection and segmentation models within military camouflaged environments. Visual data in the military domain is characterized by high ambiguity between targets and backgrounds due to sophisticated camouflage patterns and complex natural terrains designed to evade detection. The conventional rectangular Bounding Box labeling method inevitably incorporates excessive battlefield background noise (such as vegetation, rocks, and muzzle smoke) inside the frame. During model training and inference, this noise creates a "Center Shift" (geometric bias), causing the predicted center coordinates to deviate away from the actual physical center of the hull or body—the primary target area—which serves as the root cause of long-range aiming failures. To overcome these structural limitations, this study hypothesizes the technical necessity of Polygon-based segmentation labeling, which filters out background interference at the pixel level to preserve the true physical centroid of the target object through Green's theorem integration. To validate this hypothesis, a military-specific dataset consisting of soldier and tank classes was constructed, capturing various Korean battlefield noises such as woodlands, snowfields, and deserts. We evaluated a comparative suite of ten state-of-the-art architectures, including five Bounding Box-based models—comprising state-of-the-art anchor-free models (YOLOv12-S, FCOS, CenterNet) and two-stage detectors—and three segmentation-based models, including Mask R-CNN and SAM. Along with the standard mean Average Precision (mAP@50), we proposed and cross-validated the "Center Distance Error" metric—which measures the Euclidean distance between the ground-truth centroid and the predicted center normalized by the diagonal length of the ground-truth box—to quantitatively evaluate precision aiming capabilities. The experimental results revealed that while the mAP@50 performance across all models remained at a comparable level, a stark technical contrast emerged in the core metric of Center Distance Error. The Polygon-labeled Mask R-CNN achieved the lowest Center Distance Error of 5.88% despite extreme background noise, demonstrating an overwhelming advantage in precision aiming environments. Conversely, CenterNet, an anchor-free model relying on a single-point heatmap, suffered severe "heatmap collapse" under low-contrast military camouflage textures, leading to catastrophic performance degradation with a military domain mAP of 0.1810 and a total failure in small object detection (Small AP: 0.0000). FCOS, which regresses pixel-level dense distances, outperformed CenterNet with an mAP of 0.4600. However, for objects with extreme aspect ratios like self-propelled howitzers or during dynamic muzzle smoke expansions, FCOS failed to overcome the center point bias (resulting in a Center Distance Error of 7.78%) due to the structural limitations of its rectangular output format, which caused the predicted center point to drift toward empty space. Meanwhile, the latest YOLOv12 and YOLO-seg models demonstrated relatively robust center-point targeting capabilities (7.10% and 7.00%, respectively) by utilizing Distribution Focal Loss (DFL) and Task Alignment. By shifting away from traditional mAP-centric evaluation frameworks, this study advocates for the policy standardization of the "Center Distance Error" metric in military Required Operational Capabilities (ROC) to guarantee the aiming reliability of fire control systems. Furthermore, by analyzing the physical limitations of single-modal optical sensors under extreme battlefield conditions—such as false negatives due to snow camouflage and false positives triggered by muzzle flashes—this research serves as an outstanding technical milestone by proposing a next-generation hybrid network architecture that organically fuses Visible (EO) and Thermal Infrared (IR) sensor feature maps. Keywords: Military Object Detection, Data Labeling, Bounding Box, Polygon, Instance Segmentation, Center Distance Error, Precision Strike, Sensor Fusion

more

목차

1. 서론 1
1.1 연구배경 1
1.2 연구목적 및 문제 제기 2
1.3 연구가설 3
2. 관련연구 4
2.1 앵커 프리 기반 객체 탐지 알고리즘의 발전 4
2.1.1 Anchor-based 알고리즘의 연산적 한계 4
2.1.2 히트맵 기반 핵심점 추정 아키텍처 (CenterNet) 5
2.1.3 Dense 픽셀 거리 회귀 및 센터니스 분기 아키텍처 (FCOS) 6
2.1.4 최신 분포 회귀 및 태스크 정렬 아키텍처 (YOLO 최신 모델) 6
2.2 인스턴스 세그멘테이션 및 라벨링 기하학 7
2.2.1 FCN 및 Mask R-CNN 기반 픽셀 마스킹 7
2.2.2 Bounding Box 회귀 손실 함수의 한계와 세그멘테이션 무게중심 7
2.3 군사 도메인 관련 연구 현황 8
2.3.1 수목 지역에서의 심층학습 기반 군인 탐지 연구 8
2.3.2 경량화 신경망 아키텍처 및 정규화된 모서리 거리 손실 함수 기반 연구 8
2.3.3 무인기 정찰 영상 기반의 다중 스케일 피처 헤드 및 흐림 변형 강건성 연구 9
2.3.4 계층적 특징 표현 및 강화학습 지능형 에이전트 기반 국소화 정밀화 연구 10
3. 실험 방법론 11
3.1 실험 환경 및 컴포넌트 구성 11
3.2 군사 특화 데이터셋 구성 및 전처리 프로토콜 11
3.2.1 데이터 수집 범위 및 환경적 도메인 분할 11
3.2.2 데이터 라벨링의 이원화 11
3.2.3 전장 노이즈 모사를 위한 정량적 데이터 증강 기술 12
3.3 실험 대상 모델 라인업 및 파라미터 설정 12
3.3.1 아키텍처적 특성에 따른 모델 그룹 분화 12
3.3.2 신경망 학습 파라미터 통제 기준 13
3.4 Polygon 기반 질량 중심 산출 메커니즘의 작동 원리 13
3.5 성능 평가 지표 및 분석 기준 13
3.5.1 전통적 시각 지능 평가 지표 (mAP) 13
3.5.2 제안하는 조준 정밀도 지표: 중앙점 오차율 (Center Distance Error) 15
4. 실험 결과 및 분석 16
4.1 정량적 결과 비교‧ 16
4.2 결과 분석 및 도메인간 변동 가설 검증 18
4.3 중앙점 기반 탐지 모델과의 메커니즘별 구조적 취약성 해부 19
4.3.1 단일 포인트 히트맵 의존에 따른 CenterNet의 성능저하 19
4.3.2 픽셀 단위 Dense 거리 회귀 알고리즘의 중앙점 표류 편향 기전 20
5. 종합 및 제언 22
5.1 전통적 평가지표의 한계성과 국방 화력제어 지표의 표준화 제언 22
5.1.1 mAP 지표의 구조적 맹점과 전술적 위험성 22
5.1.2 작전요구성능(ROC) 내 중앙점 오차율 표준화 정책 제언 22
5.2 다중 센서 융합 하이브리드 아키텍처 제안 23
5.2.1 가시광 센서 기반 단일 모달의 물리적 한계점 논증 23
5.2.2 가시광 및 열화상(EO/IR) 센서 융합 차세대 신경망 구조 제안 24
5.3 본 연구의 한계점 및 향후 과제 26
6. 결론 26
6.1 연구 요약 26
6.2 주요 연구 결론 및 군사적 시사점 27
7. 참고문헌 28
8. Abstract 30

more