AGI Theory / v0.1
The Three-Axis Model of AGI Progress
A model of AGI progress using the scale axis, paradigm axis, and engineering axis, with the threshold of stable complex-task delivery as the key milestone.
The Three-Axis Model of AGI Progress
AGI should not be understood as a competition among several routes. It should be understood as evolution within a three-dimensional capability space. The three-axis model of AGI uses three approximately orthogonal dimensions, similar to the X, Y, and Z axes of three-dimensional space, to describe the main structure of AI capability progress: the scale axis, the paradigm axis, and the engineering axis.
The scale axis determines the upper bound of local capability.
The paradigm axis changes the capability structure and triggers capability leaps.
The engineering axis improves the capability conversion rate, meaning actual delivery capability in complex projects.
From the COG perspective, AGI is not the endpoint of any single route. It is the result of the three axes advancing together until system capability crosses the “stable complex-task delivery threshold.”
1. The Scale Axis: Expanding the Capability Ceiling and Its Boundaries
The scale axis asks whether more resources continue to produce capability improvements.
It includes the expansion of compute, data, parameter scale, training scale, and inference-time compute. During the GPT-3 period, scaled pretraining revealed a new capability structure. Since then, compute, data, and training scale have continued to expand, pushing large language model capabilities forward. By the reasoning-model stage, the scale axis extended further into inference-time compute: models consume large-scale resources not only during training, but also use more computation during inference to unlock stronger capabilities.
Therefore, the scale axis is not merely about “larger models.” It represents the process by which AI systems raise their capability ceiling through resource expansion.
But continued progress along the scale axis also faces clear boundaries.
First, compute expansion is extremely costly. Further scaling training and inference requires enormous capital investment, energy consumption, chip supply, and engineering infrastructure. The scale axis is not free natural growth; it is an increasingly expensive systemic investment.
Second, scaling laws have already shown signs of weakening or diminishing marginal returns. Larger scale may still bring gains, but we can no longer simply assume that those gains are sufficient to support a continuous leap toward AGI. The scale axis remains important, but it should not be understood as a deterministic monotonic path to AGI.
Third, today’s models are still several orders of magnitude away from independently solving complex engineering problems without relying on external engineering. Complex engineering tasks require long-term state management, task decomposition, verification loops, error rollback, and multi-stage delivery. It is difficult for gradual improvement in model capability alone to naturally cross this distance. Unless another new paradigm leap occurs, there is no sufficient reason to infer that the next emergence strong enough to solve complex engineering problems will automatically appear ahead on the scale axis.
Therefore, the scale axis can continue to raise the capability ceiling, but it cannot explain the arrival of AGI by itself. It needs the paradigm axis to open new capability structures, and it needs the engineering axis to convert existing capabilities into stable delivery capability.
2. The Paradigm Axis: Nonlinear Changes in Capability Structure
The paradigm axis asks whether the mechanism that generates capability has changed.
It includes changes in model architecture, training methods, reasoning mechanisms, and mechanisms of intelligence generation. But the key to the paradigm axis is not simply “switching to a different technical solution.” It is a structural change in how an AI system obtains capability.
The scale axis and engineering axis are relatively easier to approach continuously. The former can be advanced through more resource investment, while the latter can be continuously optimized through better tools, processes, tests, memory, task decomposition, and delivery loops. The paradigm axis is different. A paradigm leap is often not the direct result of linear engineering investment; it is closer to emergence. At some stage, a new training method, model mechanism, or capability-release method suddenly forms a higher-order capability structure.
This is also what makes the paradigm axis the most difficult: no one truly knows where the next paradigm leap is.
Today, many directions are regarded with hope, such as sparse attention, world models, reinforcement learning, self-improvement, recursive self-improvement (RSI), new memory mechanisms, multi-agent collaboration, and verifiable reasoning. But many of these directions still have not formed a fully self-consistent theory. They may become part of the next paradigm leap, or they may only be local optimizations, or they may even prove to have limited returns in large-scale systems.
The significance of GPT-3 was not merely that the model became larger. It was that scaled pretraining allowed language models to display a new capability structure. Models began to acquire broad language understanding, knowledge organization, and task generalization capabilities from large-scale text prediction. This was not a linear extension of traditional NLP tasks, but a new way of organizing capability.
The GPT o1 period can be seen as another kind of paradigm leap. Reasoning models and test-time reasoning changed how model capabilities are released. Models no longer simply output answers directly; they process problems through more complex reasoning procedures. Afterwards, inference-time compute expansion continued to produce returns along the scale axis.
This shows that the paradigm axis and the scale axis are highly related, but not identical. Scaling may nurture a paradigm leap, but it cannot guarantee one. Once a paradigm leap occurs, it opens a new space for scaling. What truly matters is not whether a specific technical direction looks advanced, but whether it changes the capability-generation mechanism and forms a new capability structure.
3. The Engineering Axis: Improving Capability Conversion Rate
The engineering axis asks whether model capability can be organized into stable delivery capability in real complex tasks.
Large language models already possess strong local intelligence, but local intelligence does not automatically equal stable delivery in complex tasks. Real complex tasks usually include goal understanding, task decomposition, long-term state management, tool use, error correction, test verification, and delivery loops. Even when a model performs strongly on local problems, it may still fail in long-chain tasks because of context loss, missing constraints, insufficient verification, or unstable execution.
Therefore, the core of the engineering axis is not making the model itself smarter, but improving the conversion rate of model capability. It asks how, on top of existing model capability, tools, processes, environments, and governance mechanisms can organize local intelligence into dependable system capability.
Most current progress in AI engineering revolves around this goal. For example, connecting models to terminals, IDEs, code repositories, and test environments lets them enter real workflows. Plan mode, spec-first work, task decomposition, and stage acceptance reduce the probability of losing control in complex tasks. Automated tests, logs, CI, benchmarks, and human acceptance turn model output from “seems reasonable” into “can be verified.” Context management, project documentation, external memory, and multi-session collaboration extend the effective working radius of models in complex tasks.
The emergence of Claude Code can be seen as an important advance on the engineering axis. It does not merely make the model smarter. It places the model in a real engineering environment, iterates through a CLI single-session mode, and develops effective engineering methods such as plan mode, subagents, tool use, and test loops. It demonstrates an important fact: for model capability to convert into productivity, it must enter real task environments and be constrained by tools, processes, and verification mechanisms.
But the engineering axis should not be understood as the product progress of a single coding agent. The larger direction is to build a complete collaborative governance layer around stable delivery of complex tasks. This governance layer needs to manage task decomposition, context allocation, executor selection, test verification, evidence recording, human intervention, result reproduction, and multi-session collaboration.
This is precisely the direction COG focuses on. It does not try to replace foundation models, and it does not understand AGI as the sudden completion of a single model capability. Instead, it seriously advances the engineering axis: organizing models’ local intelligence into stable delivery capability in complex tasks. Expressed through the three-axis model, COG focuses on how to maximize capability conversion rate under existing capability ceilings and capability structures, and how to move AI from “locally smart” toward “systemically deliverable.”
4. The Approximate Orthogonality of the Three Axes
The “orthogonality” in the three-axis model of AGI does not mean the three axes are completely independent in a mathematical sense. It means they represent three approximately independent directions of effort.
The scale axis addresses whether more resources continue to bring capability improvements. The paradigm axis addresses whether the capability-generation mechanism has changed. The engineering axis addresses whether model capability can be organized into stable delivery capability in real complex tasks. These three questions are related, but they cannot replace one another.
Strictly speaking, engineering structures are affected by model scale: stronger models can often handle longer chains, more complex tool use, and higher-level task decomposition. Paradigm changes can also qualitatively alter the return on investment of the scale axis: once a new capability-generation mechanism appears, resource investment that previously faced diminishing marginal returns may open a new space for scaling.
Conversely, engineering can amplify the gains from scale and paradigm, but it cannot change the capability ceiling of the base model. Scale expansion can nurture paradigm leaps, but it cannot guarantee that they happen. Paradigm leaps can change capability structure, but they still need engineering to enter real complex tasks.
Therefore, the three axes are not isolated from one another. They are approximately orthogonal as an analytical abstraction. The value of this abstraction is not to describe every causal relationship with strict precision, but to guide practical directions of effort. When an AI system lacks capability, we need to distinguish whether the problem mainly comes from the capability ceiling, the capability structure, or the capability conversion rate. Only if this distinction holds can AGI progress avoid being reduced to a single variable, and avoid being misunderstood as the lone victory of one route.
5. The Stable Complex-Task Delivery Threshold
From the COG perspective, AGI is not the endpoint of any single route. It is the result of the three axes advancing together until system capability crosses the “stable complex-task delivery threshold.”
This threshold means that an AI system is not only smart in local tasks, but can also stably complete decomposition, execution, verification, and delivery in long-term, complex, dynamic, constrained real tasks.
In this framework, the scale axis raises the capability ceiling, the paradigm axis changes capability structure, and the engineering axis improves capability conversion rate. Together, the three determine how far an AI system is from AGI.
Different stages may be dominated by different axes. During the GPT-3 period, a paradigm leap and scale expansion jointly drove the rise of large language models. During the GPT o1 period, the reasoning paradigm and inference-time compute expansion further released capability. Engineering progress represented by Claude Code has allowed model capability to enter real engineering scenarios more deeply.
Therefore, AGI progress should not be understood as a competition among “three paths.” It should be understood as continuous evolution within a three-axis coordinate system. The truly important question is not which route wins alone, but how the three axes jointly push AI systems across the stable complex-task delivery threshold.