Research paper / v2.1.9

Multi-Round Audit Convergence: Coverage-State Assessment and Recursive Control for Complex AI Engineering Tasks

liu, ming

2026-08-16
Multi-Round Audit ConvergenceEngineering AGIautonomous software engineeringsearch-space governancetask decompositionSpeccode reviewagent harness

Abstract

Large language models and their agents can already handle some real-world software engineering tasks, yet a single successful execution is still insufficient to demonstrate that a specific result has reached a level at which it can be delivered and relied upon. Prior work proposed the SCR relationship (Search Space–Coverage Capability–Delivery Reliability), arguing that delivery reliability depends on how well the coverage capability of the complete AI development system matches the task search space. This matching relationship, however, cannot be observed directly. Engineering systems therefore need to infer it from evidence produced during execution and verification. This paper first analyzes the relationship among execution, independent audit, and repair. After execution produces a candidate result, an independent audit can search the concrete engineering object for omissions, conflicts, and verification gaps while operating within a problem scope that is more constrained than the original generation task. Once identified issues are repaired, the next audit operates on an updated engineering object. Execution, audit, and repair are therefore not simple repetitions of the same problem. Instead, they allow limited coverage capability to act on different problems in stages, making it possible for the complete system to handle some tasks beyond the range that a single execution can cover reliably. On this basis, this paper proposes Multi-Round Audit Convergence (MRAC). Each audit round produces findings on the current engineering object; successive audit, repair, and re-audit rounds further form a trajectory describing changes in the number, severity, and types of newly discovered issues. MRAC performs second-order observation on this multi-round issue trajectory: when newly discovered significant issues decrease overall and the trajectory gradually stabilizes, it exhibits convergence; when new significant issues continue to emerge despite repeated repairs, the trajectory exhibits non-convergence. The significance of multiple rounds is not simply to increase the number of audits, but to create the cross-round object required for this second-order observation. What MRAC observes directly is the issue trajectory. In conjunction with SCR, when task boundaries, audit protocols, and verification conditions remain comparable, and later audits retain an effective ability to discover remaining significant issues, convergence can serve as positive process evidence that the current system has sufficient coverage capability to handle the task consistently. It cannot prove that the result is absolutely correct, nor can it rule out shared blind spots. Persistent non-convergence more directly indicates that the current system has not demonstrated that it can handle the task consistently. The two forms of evidence are therefore asymmetric: non-convergence is relatively strong negative evidence, whereas convergence depends on the additional condition that auditing remains effective. As of the data cutoff for this paper, the current project's MRAC observation records cover 35 tasks or subtasks resulting from decomposition. Of these, 33 have reached the project's adopted convergence condition, completed verification, and been merged into the main branch; 2 remain in progress. In addition, 3 parent tasks triggered decomposition after persistent non-convergence within their original boundaries, producing 10 subtasks in total. All of these subtasks reached the convergence condition and completed verification, after which the 3 parent tasks were closed. No case has yet been observed in the current records in which, after decomposition, persistent non-convergence of a subtask prevented closure of the overall task. This small sample cannot prove that any engineering task can always be made to converge through recursive decomposition, nor can it by itself prove that true delivery reliability improves correspondingly. It does, however, show that the process of “non-convergence → decomposition → renewed convergence” has repeatedly occurred in multiple real engineering tasks. This paper therefore positions MRAC as a mechanism for search-space coverage assessment and task adjustment built on top of the SCR relationship. The evidentiary requirements and permissible scope of extrapolation for claims at different levels are left to subsequent MRAC L1–L5 research.

Loading the full paper. You can also open the PDF using the link above.

Original paper reproduced under CC BY 4.0, preserving author attribution and layout. CC BY 4.0