
জার্নাল পেপার2026
জার্নাল পেপার 102Five AI systems outscored the average OOP cohort but still failed compilation and advanced concepts
Marina Lepp, Joosep Kaimre
arXiv preprint
Lepp and Kaimre evaluated ChatGPT-5.2, DeepSeek-V3, Gemini 2.5 Flash, Claude Sonnet 4.5, and Microsoft 365 Copilot on authentic introductory OOP tests and examinations using student grading criteria. Systems exceeded the historical average and often solved long tasks, yet some code did not compile and interfaces, abstract classes, inheritance, and image-based questions remained difficult. The results challenge take-home assessment validity without proving student learning.
programming assessmentobject-oriented programminggenerative AI
৫০০-শব্দের সারাংশ পড়ুন →