TY - GEN
T1 - Feature Densities for Uncertainty Quantification for Complex Text Detection in Spanish
AU - Abreu-Cardenas, Miguel
AU - Calderón-Ramírez, Saúl
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Text simplification is crucial for enhancing content accessibility, particularly for audiences with low literacy or sensory disabilities. Despite recent advances using large language models (LLMs), their computational expense and predominant control by private entities hinder practical deployment. Efficiently detecting complex text segments requiring simplification is thus vital for resource optimization. This work addresses uncertainty quantification for complex text detection in Spanish-an under-explored area -to improve model transparency and enable targeted retraining. We train a BETO-based classifier and compare three uncertainty-quantification techniques-MonteCarlo Dropout, Deep Ensembles, and a novel Feature Density Estimation that operates in latent space, leverages representations to model feature distributions, offering post-training computational efficiency. Experiments on a financial education dataset (5,314 text pairs) show that Feature Density Estimation matches Monte Carlo Dropout's performance (Jensen-Shannon distances: 0.42 vs. 0.43) at significantly lower computational cost, while Deep Ensembles underperformed (0.34). Statistical analysis confirms Feature Density Estimation as a lightweight, effective uncertainty quantification alternative for low-resource language applications.
AB - Text simplification is crucial for enhancing content accessibility, particularly for audiences with low literacy or sensory disabilities. Despite recent advances using large language models (LLMs), their computational expense and predominant control by private entities hinder practical deployment. Efficiently detecting complex text segments requiring simplification is thus vital for resource optimization. This work addresses uncertainty quantification for complex text detection in Spanish-an under-explored area -to improve model transparency and enable targeted retraining. We train a BETO-based classifier and compare three uncertainty-quantification techniques-MonteCarlo Dropout, Deep Ensembles, and a novel Feature Density Estimation that operates in latent space, leverages representations to model feature distributions, offering post-training computational efficiency. Experiments on a financial education dataset (5,314 text pairs) show that Feature Density Estimation matches Monte Carlo Dropout's performance (Jensen-Shannon distances: 0.42 vs. 0.43) at significantly lower computational cost, while Deep Ensembles underperformed (0.34). Statistical analysis confirms Feature Density Estimation as a lightweight, effective uncertainty quantification alternative for low-resource language applications.
KW - BERT
KW - Deep Learning
KW - Safe Artificial Intelligence
KW - Text complex prediction
KW - Transformers
KW - Uncertainty Quantification
UR - https://www.scopus.com/pages/publications/105038743672
U2 - 10.1109/BIP68491.2025.11489140
DO - 10.1109/BIP68491.2025.11489140
M3 - Contribución a la conferencia
AN - SCOPUS:105038743672
T3 - 2025 IEEE 7th International Conference on BioInspired Processing, BIP 2025
BT - 2025 IEEE 7th International Conference on BioInspired Processing, BIP 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 7th IEEE International Conference on BioInspired Processing, BIP 2025
Y2 - 3 December 2025 through 5 December 2025
ER -