Benchmarking

arXiv
Agents' Last Exam
arXiv
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
EMNLP 2025 Findings
NLP-ADBench: NLP Anomaly Detection Benchmark
ACL 2025 Findings
AD-LLM: Benchmarking Large Language Models for Anomaly Detection