Resumen
The detection of semantically similar terrorist incidents remains a major challenge for security analysts, who often face time-consuming manual reviews and cognitive biases. Although vector database approaches supported by word embedding models offer a promising solution, systematic evaluations of these models for identifying similar attacks are limited. This study benchmarks seven text embedding models from the all-mpnet-base-v2, E5, and general text embedding (GTE) families to assess their effectiveness in identifying semantically related terrorist incidents using a vector database architecture. Using the Global Terrorism Database (GTD), narrative descriptions of attacks were embedded in a Qdrant vector database to uncover latent similarities. The evaluation used four query types with varying complexity and metadata usage. Results show that the GTE family, particularly gte-large, achieved the highest average precision. Moreover, metadata proved crucial, as complex queries with metadata filters yielded superior retrieval performance. This work provides empirical evidence on embedding-based semantic search for threat detection, supporting scalable AIassisted systems for intelligence agencies.
| Idioma original | Inglés |
|---|---|
| Publicación | Proceedings of the IEEE Central America and Panama Convention, CONCAPAN |
| N.º | 2025 |
| DOI | |
| Estado | Publicada - 2025 |
| Evento | 43rd IEEE Central America and Panama Convention, CONCAPAN 2025 - San Salvador, El Salvador Duración: 26 nov 2025 → 28 nov 2025 |
ODS de las Naciones Unidas
Este resultado contribuye a los siguientes Objetivos de Desarrollo Sostenible
-
ODS 7: Energía asequible y no contaminante
-
ODS 16: Paz, justicia e instituciones sólidas
Huella
Profundice en los temas de investigación de 'Benchmarking Text Embedding Models for Semantic Search in Terrorism Incident Detection'. En conjunto forman una huella única.Citar esto
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver