Transformer-Based Semantic Similarity for Duplicate Detection in R&D Project Proposals: a Case Study on the TusaŞ Lift Up Program


Akbaş E., Yüce B., Firat H. A., Umay B., Solak S.

8th International Congress on Human-Computer Interaction, Optimization, and Robotic Applications (ICHORA), Ankara, Türkiye, 21 - 23 Mayıs 2026, ss.1-5, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/ichora69329.2026.11537006
  • Basıldığı Şehir: Ankara
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.1-5
  • Kocaeli Üniversitesi Adresli: Evet

Özet

The increasing number of R&D projects in the defense and aerospace sectors makes it difficult to detect duplicate proposals during evaluation processes. Keywordbased systems are insufficient to capture semantic relationships between project texts. In this study, transformer-based models were applied to detect semantic similarity among R&D project proposals within the TUSAŞ LIFT UP program. Project titles, abstracts, and keywords were extracted from PDF proceedings and converted into a structured dataset. Pre-trained BERTbased models were used to generate embeddings, and cosine similarity was employed to measure overlap. Experimental results show that BERTurk and DistilBERT produced similarity scores above 85 % even for unrelated project pairs, indicating limited discriminative capability. In contrast, Sentence-BERT generated similarity scores below 5% for unrelated proposals and provided a wider similarity distribution. The findings suggest that sentence-level transformer models are more appropriate for duplicate detection in R&D evaluation processes.