Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

Abstract

Output correctness alone does not tell us whether a large language model actually reasons correctly about code. This work introduces CodeRQ-Bench, a benchmark for assessing LLM reasoning quality in coding tasks beyond output correctness, together with VERA, a two-stage evaluator for detecting flawed reasoning in model solutions. VERA achieves improvements of up to 0.26 in AUCROC and 0.21 in AUPRC over existing evaluation approaches, moving LLM evaluation for code toward transparency and reasoning-aware assessment.

Publication
arXiv preprint
Yuangang Li
Yuangang Li
PhD Student at UCI