Transformer-Based Semantic Similarity for Duplicate Detection in R&D Project Proposals: a Case Study on the TusaŞ Lift Up Program
8th International Congress on Human-Computer Interaction, Optimization, and Robotic Applications (ICHORA), Ankara, Türkiye, 21 - 23 Mayıs 2026, ss.1-5, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Doi Numarası: 10.1109/ichora69329.2026.11537006
- Basıldığı Şehir: Ankara
- Basıldığı Ülke: Türkiye
- Sayfa Sayıları: ss.1-5
- Kocaeli Üniversitesi Adresli: Evet
Özet
The increasing number of R&D projects in the defense and aerospace sectors makes it difficult to detect duplicate proposals during evaluation processes. Keywordbased systems are insufficient to capture semantic relationships between project texts. In this study, transformer-based models were applied to detect semantic similarity among R&D project proposals within the TUSAŞ LIFT UP program. Project titles, abstracts, and keywords were extracted from PDF proceedings and converted into a structured dataset. Pre-trained BERTbased models were used to generate embeddings, and cosine similarity was employed to measure overlap. Experimental results show that BERTurk and DistilBERT produced similarity scores above 85 % even for unrelated project pairs, indicating limited discriminative capability. In contrast, Sentence-BERT generated similarity scores below