EA-PRF: An Explainable AI and Human-in-the-Loop Framework for Scholarly Peer Review
ABSTRACT Scholarly peer review is under increasing strain: submission volumes are rising faster than the pool of qualified reviewers, turnaround times are lengthening, and inter-reviewer agreement on acceptance decisions remains low across many disciplines. This paper presents the Explainable AI-Assisted Peer Review Framework (EA-PRF), a system that pairs large language models (LLMs) with a dedicated explainability layer and a human-in-the-loop (HITL) validation stage so that automated review support augments, rather than replaces, editorial judgment. EA-PRF decomposes a submitted manuscript into five review dimensions — novelty, methodological soundness, clarity, significance, and coverage of related work — and generates a calibrated recommendation, a confidence score, and a natural-language rationale grounded in specific manuscript spans for each dimension. Human reviewers inspect, edit, or override every AI-generated judgment through a dedicated interface, and their actions are logged to an active-learning store that periodically recalibrates the underlying models. We evaluate EA-PRF on a corpus of 1,240 manuscripts spanning engineering, computer science, and applied sciences, comparing it against a human-only baseline and an LLM-only baseline without explainability or human oversight. EA-PRF improves F1-score for acceptance-decision prediction from 0.72 (LLM-only) to 0.84, raises explanation-agreement scores from 0.58 to 0.79, and reduces mean reviewer time per manuscript by 65% relative to unaided human review, while preserving inter-rater agreement (Cohen's κ = 0.68) close to that of experienced human reviewer pairs (κ = 0.71). These results suggest that explainable, human-supervised LLM assistance can meaningfully reduce reviewer burden without sacrificing decision quality or accountability.