This study examines 50 university assessments to compare evaluations and feed- back generated by a human instructor and two AI models (GPT and GEM), using shared rubrics and statistical analyses. Findings show strong alignment between the instructor and ChatGPT, while Gemini applies stricter grading. Qualitative data highlight differences in feedback clarity and usefulness. Results indicate that AI can complement, but not replace, human judgment, supporting more consistent and pedagogically meaningful assessment processes.
Lo studio analizza 50 prove universitarie per confrontare valutazioni e feedback prodotti da un docente e da due modelli di IA (GPT e GEM), utilizzando rubriche condivise e analisi statistiche. I risultati mostrano un’elevata concordanza tra docente e ChatGPT, mentre Gemini adotta criteri più severi. La valutazione qualitativa evidenzia differenze nella chiarezza e nell’utilità dei feedback. La ricerca suggerisce che l’IA possa integrare, ma non sostituire, il giudizio umano nei processi valutativi.
Dalla valutazione tradizionale alla valutazione simbiotica: analisi comparativa tra docenti e sistemi di IA nella valutazione accademica
Sarcina F.
;Cuzzi A.;Baldassarre M.
2026-01-01
Abstract
This study examines 50 university assessments to compare evaluations and feed- back generated by a human instructor and two AI models (GPT and GEM), using shared rubrics and statistical analyses. Findings show strong alignment between the instructor and ChatGPT, while Gemini applies stricter grading. Qualitative data highlight differences in feedback clarity and usefulness. Results indicate that AI can complement, but not replace, human judgment, supporting more consistent and pedagogically meaningful assessment processes.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


