검색 상세

LLM기반 임베딩 조합을 활용한 IP담보설정 특허 식별 성능 비교 연구

A Comparative Performance Study of LLM-Based Embedding Combinations for Identifying Patents with IP Collateral Registration

초록/요약

기존 특허가치 관련 연구는 주로 청구항 수, 피인용 수, 패밀리 규모, 권리 잔존기간, 기술분류 등 구조화된 정량지표를 중심으로 특허의 경제적·기술적 가치를 설명해 왔다. 그러나 특허 문서는 발명의 기술적 내용, 권리범위, 문제 해결 방식, 응용 가능성 등 정량지표만으로 포착하기 어려운 복합적인 정보를 포함하고 있다. 이에 본 연구는 특허 텍스트 정보가 구조화 특허정보와 결합되었을 때 담보설정이라는 실제 금융거래 관측 신호를 예측하는 데 추가적인 예측 정보를 제공하는지 검토 하였다. 이를 위해 본 연구는 입력특성 조합, 텍스트 변환 방식, 분류모형, 데이터의 양성·음성 클래스 비율 등 여러 축을 체계적으로 탐색하였다. 평가구조 측면에서는 무작위 분할 기반의 Non-OOT 평가 와 등록연도 기준의 OOT 평가를 병행함으로써, 표본 내 분류 성능과 시간 외 일반화 성능을 구분 하여 검토하였다. 또한 본 연구는 모형의 성능을 예측 성능지표에만 한정하지 않고, 우선 검토 후보 군 선별이라는 실무적 활용 가능성의 관점에서 검토하기 위해 Top-K 분석과 임계값 민감도 분석을 함께 수행하였다. 분석 결과, 구조화 특허정보와 특허 텍스트 정보를 결합한 특성 조합은 단일 정보 유형만을 활용 한 경우보다 전반적으로 높은 예측 성능을 보였다. 이는 특허 텍스트 정보가 기존 구조화 특허정보 를 보완하는 추가적인 예측 정보로 활용될 수 있음을 시사한다. 특히 예측점수 상위 후보군에서 실 제 담보설정특허가 무작위 선별보다 높은 비율로 포함되어, 본 연구의 모형은 개별 특허의 확정적 판별보다는 금융기관의 우선 검토 대상 순위화와 후보군 압축을 위한 보조 도구로 활용될 수 있음 을 확인하였다. 아울러 동일한 예측모형이라도 활용 목적에 따라 중시해야 할 성능지표와 임계값 설정이 달라질 수 있음을 확인하였다. 예를 들어 금융기관의 심사·검토 과정에서는 제한된 심사 자원을 고려하여 정밀도와 Top-K 성능을 중시할 수 있으며, 잠재적 IP금융 후보 특허를 폭넓게 발굴하고 IP 활용 저변을 확대하려는 정책적 관점에서는 재현율과 F1 점수를 함께 고려할 필요가 있다. 따라서 특허 담보설정 예측모형은 단일한 성능지표나 고정된 임계값으로 평가하기보다, 운영 목적에 맞는 성능지 표를 선택하고 그에 따라 임계값을 조정하는 방식으로 활용되어야 한다. Keywords : IP금융, 특허담보, LLM 임베딩, Top-K, 임계값 조정

more

초록/요약

Previous studies on patent value have mainly explained the economic and technological value of patents using structured quantitative indicators such as the number of claims, forward citations, patent family size, remaining patent term, and technology classification. However, patent documents contain complex information that is difficult to capture through quantitative indicators alone, including the technical content of inventions, scope of rights, problem-solving mechanisms, and potential applications. This study examines whether patent text information provides additional predictive value when combined with structured indicators for identifying patents associated with collateral registration, an observable signal in actual financial transactions. To this end, this study systematically explores multiple dimensions, including input feature combinations, text representation methods, classification models, and positive-to-negative class ratios. In terms of evaluation design, the study conducts both a Non-OOT evaluation based on random splits and an out-of-time (OOT) evaluation based on registration years, allowing a comparison between within-period classification performance and temporal generalization. In addition, model performance is evaluated not only in terms of conventional predictive performance metrics but also from the practical standpoint of identifying candidates for priority review, using Top-K analysis and threshold sensitivity analysis. The results show that feature combinations integrating structured information and patent text generally achieved higher predictive performance than those using only a single type of information across multiple experimental settings. This suggests that patent text can provide complementary predictive information beyond existing structured indicators. In particular, patents with collateral registration histories were represented at a higher rate among top-ranked candidates than under random selection, indicating that the proposed models can be used not as tools for definitive patent-level classification, but as screening tools for ranking and narrowing down candidates for priority review by financial institutions. Furthermore, the results suggest that the relative importance of precision and recall may vary depending on operational objectives and that threshold settings should be adjusted accordingly. Specifically, in financial contexts where minimizing false positives is important, thresholds may be set to emphasize precision. In contrast, when the objective is to expand the IP finance ecosystem and broadly identify potential candidates for IP finance, thresholds may be set to emphasize recall. Keywords: IP finance, patent collateral, LLM-based embedding, Top-K, threshold adjustment

more

목차

제1장 서론 1
제1절 연구 배경 및 필요성 1
제2장 관련 선행 연구 및 이론적 배경 3
제1절 특허 ․ IP금융 관련 선행연구 3
1. 특허 가치평가와 IP금융 연구 3
2. 특허 텍스트 분석과 딥러닝 기반 특허 연구 3
제2절 예측모형의 방법론적 배경 4
1. 불균형 데이터와 평가 지표 4
2. OOT 검증과 시간 일반화 5
제3장 실험방법 및 분류모형 설계 7
제1절 데이터 구성 및 실험설계 7
1. 분석자료와 표본구성 7
2. 예측문제와 기준 9
3. 입력특성 구성 11
4. 표본비율 및 평가구조 13
제2절 분류모형 및 성능 평가 15
1. 분류모형 15
2. 성능 평가 지표 16
3. 실험 수행 및 결과 산출 절차 18
제4장 실험결과 및 분석 20
제1절 실험결과 개요 20
제2절 성능 비교 21
1. 전체 성능 비교: Non-OOT와 OOT 21
가. Non-OOT와 OOT 전체 평균 성능 비교 21
나. 표본비율별 PR-AUC와 무작위 기준선 대비 향상도(lift) 23
2. 분류모형별 성능 비교 25
3. 특성 조합별 성능 비교 26
가. 입력특성 조합별 성능 비교 일반 26
나. 텍스트 입력 길이 조건별 성능 28
다. 분류모형과 특성 조합의 교차 성능 30
라. 단일 실험 조합 성능 비교 31
제3절 심화 분석 32
1. 임계값 조정에 따른 F1 점수 민감도 분석 32
2. Top-K 후보군 선별 성능 34
3. 계산 효율성 분석 36
가. 평가구조별 입력특성 조합별 평균 계산 시간 37
나. 분류모형별 평균 계산 시간 38
제4절 실험결과 종합 39
제5장 결론 41
제1절 연구 결과 41
제2절 연구의 의의 및 시사점 43
제3절 본 연구의 한계 및 향후 연구 44
참고문헌 45
ABSTRACT 48

more