Assessing PBL work in the LLM era
A Year-over-Year Comparison of Artifact Size and Quality in Individual Free-Topic PBL Assuming LLM Use
This study analyzes how students' work in a free-topic PBL course changed once generative AI was available, focusing on a third-year undergraduate software engineering course.
Show details Hide details
Problem
In LLM-assisted development exercises, the size and quality of what students build depend not only on their skills but also on how they use LLMs and how the course is designed. We need to clarify what can actually be assessed from the deliverables.
Approach
We compared the functional size of deliverables from 48 students in 2024 and 46 students in 2025, using LLM-assisted measurement based on the COSMIC method. For 2025, we also examined the relationship between basic proficiency (measured by CFRP) and test pass rates.
Key results
Average functional size in 2025 was about 3.2 times that of 2024. CFRP showed only a weak correlation with functional size, but a moderate positive correlation with test pass rates. The study also states the limitations of LLM-assisted assessment.