Research paper / v1.0

MRAC Bench: A Software Engineering Evaluation Method and Infrastructure Based on Multi-Round Audit Convergence

刘, 明

2026-09-16

Abstract

Software engineering evaluation needs a continuing supply of tasks that reflect real development, distinguish system capabilities, and remain affordable to construct. This paper proposes MRAC Bench, an evaluation method and infrastructure based on Multi-Round Audit Convergence (MRAC). Its central idea is to use audit-and-repair trajectories as a conditional signal of whether an AI system can achieve stable, effective coverage of the relevant search space under given task and resource constraints: consistently identifying and coordinating the constraints, dependencies, and verification requirements that matter for delivery. The system works toward a fixed objective, undergoes independent audits, and repeatedly repairs its output. Evaluation considers convergence within budget, process trajectories, and resource consumption. Audits can examine requirement and design specifications or execution artifacts such as code changes. Candidate tasks are generated from real repositories and change requests, then screened through runs that calibrate difficulty and the applicable capability range. Construction does not require a complete reference implementation or a dedicated acceptance verifier to be prepared for every task, potentially lowering the barrier to task creation substantially. Varying change scope, cross-module effects, and combinations of constraints also offers a way to keep constructing tasks near system capability boundaries. As long as real engineering continues to provide more complex tasks that can be audited effectively, the evaluation mechanism may continue to scale with model progress. To control the cost of reproduction at scale, the pilot study uses specification audits. The available records show differences among three model tiers in persistent non-convergence, gradual stabilization, and rapid stabilization. These observations do not directly establish actual coverage or a general capability ranking. Missed defects and differences in audit preferences may affect the signal's meaning. Comparability across models, repeatability, cost advantages, and extensibility to greater difficulty still require systematic testing. The paper also discusses potential industry implications and develops three classes of counterexamples concerning the direction of the signal, limits to difficulty scaling, and unavoidable task-construction costs.

Loading the full paper. You can also open the PDF using the link above.

Original paper reproduced under CC BY 4.0, preserving author attribution and layout. CC BY 4.0