Project Details
Description
In recent years, large language models (LLMs) have demonstrated remarkable capabilities in generating coherent text, answering complex questions, and performing symbolic reasoning. However, their reliability and accuracy in solving university-level Mathematics and Physics problems remain uncertain. These topics require deep conceptual understanding, rigorous application of theoretical principles, and structured procedural execution—skills that current LLMs only partially exhibit.
This research project aims to empirically assess the performance of LLMs in solving representative problems from university Mathematics and Physics. The study seeks to evaluate their precision, consistency, and explanatory quality, while identifying typical reasoning errors and cognitive limitations.
The study adopts a mixed-methods approach that combines quantitative and qualitative analyses. During the experimental phase, selected LLMs will be tested using authentic problems taken from ITCR course exams and assignments. Model outputs will be evaluated using a structured rubric that measures dimensions such as accuracy, procedural rigor, clarity, and reasoning. This rubric builds
upon established frameworks in the literature. The resulting data will be analyzed both statistically and qualitatively to identify strengths, weaknesses, and recurrent failure patterns.
At the national level, the project supports ITCR’s commitment to fostering educational innovation and the responsible use of artificial intelligence. Internationally, it aligns with current trends in generative model evaluation and AI-assisted education, contributing a Latin American perspective to the global discourse. The outcomes are expected to include peer-reviewed publications that disseminate empirical evidence, promote good practices, and help establish standards for the pedagogically sound use of LLMs in science and engineering education.
This research project aims to empirically assess the performance of LLMs in solving representative problems from university Mathematics and Physics. The study seeks to evaluate their precision, consistency, and explanatory quality, while identifying typical reasoning errors and cognitive limitations.
The study adopts a mixed-methods approach that combines quantitative and qualitative analyses. During the experimental phase, selected LLMs will be tested using authentic problems taken from ITCR course exams and assignments. Model outputs will be evaluated using a structured rubric that measures dimensions such as accuracy, procedural rigor, clarity, and reasoning. This rubric builds
upon established frameworks in the literature. The resulting data will be analyzed both statistically and qualitatively to identify strengths, weaknesses, and recurrent failure patterns.
At the national level, the project supports ITCR’s commitment to fostering educational innovation and the responsible use of artificial intelligence. Internationally, it aligns with current trends in generative model evaluation and AI-assisted education, contributing a Latin American perspective to the global discourse. The outcomes are expected to include peer-reviewed publications that disseminate empirical evidence, promote good practices, and help establish standards for the pedagogically sound use of LLMs in science and engineering education.
General Objective
Evaluar el desempeño de los modelos de lenguaje de gran escala (LLM) en la resolución de problemas de Física y Matemática de los cursos de servicio del ITCR, mediante la selección y análisis de ejercicios representativos, el uso de rúbricas fundamentadas en la literatura y la comparación de sus características y limitaciones; con el fin de valorar su pertinencia, alcances y limitaciones para su eventual integración en el contexto educativo universitario.
Research Lines
Inteligencia artificial, educación
| Status | Active |
|---|---|
| Effective start/end date | 1/01/26 → 31/12/26 |
Collaborative partners
Keywords
- large language models
- mathematical reasoning
- STEM education
- educational artificial intelligence
- empirical evaluation
Fingerprint
Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint.