What happens when AI starts improving itself
Imagine a Friday board meeting at a hotel and restaurant company. On the screen appears a report prepared by AI. Alongside proposals concerning room sales, conferences, and catering, it contains something less ordinary: a plan to change the way the system itself will prepare and verify subsequent proposals. This time, the system is asking the board for permission to improve itself.
The company owns two very different hotels. The first is small, boutique, five-star: 50 rooms and high rates. The second has 180 rooms, conference facilities, and leisure amenities. On top of that come three well-regarded restaurants in the mid-to-upper bracket. A shared owner does not mean a shared way of making money. An empty room in the boutique hotel, a vacant conference suite, and an unused table pose different questions to the board.
In this hypothetical scenario, the company gives the AI access to structured operational data. Initially the system helps make decisions. In time, it also begins rebuilding the tools with which it identifies problems, and the way it selects which trials are worth running.
We usually ask whether the company performs better thanks to technology. Here a second question arises: is the system becoming more capable of finding further improvements? And can we still tell when it is genuinely improving and when it is merely producing ever more persuasive reports?
Where the Loop Begins
In computing, recursion means a procedure calling itself, directly or indirectly. If a computation is to terminate, a stopping condition is required — one the procedure will actually reach. Repeating calls does not in itself constitute progress. In discussions of AI, the word describes more broadly a feedback loop: the system participates in creating an improvement, and that improvement increases its capacity to create further improvements. It is this second step that counts. Higher room sales may be a commercial success. Only an improvement in the system's capacity for further development makes this example recursive.
Updating forecasts on the basis of new data does not by itself demonstrate recursive self-improvement. An agent — a model using tools and carrying out successive actions — can go further: propose a change to its own analytical tool, verify it, and use it in subsequent work. A further level again is improving the learning algorithms of the AI models themselves. In our hotel example, we are dealing with limited self-improvement by an agent, which is not evidence of autonomous development of artificial intelligence as a whole.
In 1965, Irving John Good — a mathematician who during the war broke ciphers with Turing at Bletchley Park — considered a machine surpassing humans at designing machines as well. It could therefore create still more capable successors. This is how he described the possibility of an intelligence explosion: a hypothesis resting on a strong assumption about the capabilities of the first machine. [1]
Jürgen Schmidhuber proposed the Gödel machine: a program that rewrites its own code upon finding a proof that the change will pay off according to a predetermined measure. Finding such a proof has remained an enormous difficulty, so today's experimental approaches typically rely on trials and comparing results. [2] One must still establish, however, how to distinguish genuine improvement from apparent success.
Two Ledgers
In the first stage, the system handles familiar analytical tasks. At the boutique hotel it calculates the result after acquisition and servicing costs. The board wants to maintain both the high rate and the quality that justifies it; simply filling 50 rooms by discounting tells them little. At the conference hotel, the calculation covers the entire event, including room setup and cleaning. It must also account for bookings that couldn't be accommodated and compare the result against alternative uses of the same dates. For the three restaurants it runs three separate forecasts, because they have different menus, different guests, and different rhythms. This is still ordinary analytics. The recursive element appears only when the agent turns to its own way of working.
Suppose its forecast for the boutique hotel fails consistently. Analysis of the history shows that the tool wrongly treats a change of date on an existing reservation as the appearance of new demand. The agent proposes a fix to the software processing these events, preserves the previous version, and compares both on data it did not use while preparing the change.
After independent verification and approval of the fix, the agent works with a better tool. In the next cycle, the corrected data may reveal that its testing procedure compares ordinary weekends with dates of major events. The agent therefore proposes a change to how forecasts are tested and again submits it for assessment. The improved tool helped prepare the next improvement. This is already a loop, though two successful fixes do not yet demonstrate sustained acceleration.
At the conference hotel, the same principle applies to the selection of experiments. Initially the agent proposes changing the package price, the way the offer is presented, and the scope of services all at once; after the sales period ends, nobody knows what worked. The next version of its procedure separates these changes, selects comparable enquiries according to a plan established in advance, and checks whether the trial will disrupt operations. The agent rebuilds its own method of planning tests so that the next trial yields a more decisive result.
The restaurants provide a third test. The agent discovers that its tool counted discarded products but omitted orders the kitchen could not accept because a dish had run out. It therefore extends the diagnosis to include a record of unavailability and verifies the fix at the other venues, preserving their distinct profiles. A better diagnosis should then help it select more accurate changes to forecasting and food preparation.
We therefore have two separate ledgers. The first concerns the company: does it earn more after costs and maintain its service standard? The second concerns the agent: do the improved tools and procedures help it better identify problems and search for solutions? The test is an improvement in the quality of solutions on independent tasks at the same budget and comparable difficulty. Usefulness criteria are established before the trial. What counts is a confirmed effect, not the number of proposals. A good month for the restaurants cannot be automatically credited to the system's growing intelligence.
The Percent That Keeps Working
Researchers demonstrate fragments of this mechanism in real experiments. In May 2025, Google DeepMind described AlphaEvolve, a system combining language models, automated evaluation, and evolutionary search for algorithms. According to the company, the improvement it found accelerated an important portion of the computation in the Gemini architecture by 23 percent, which across the whole of Gemini's training translated into a reduction in time of around one percent. [3]
The first figure describes the success of a component, the second its significance for the whole; shortening one kitchen operation likewise doesn't mean an identical improvement in a restaurant's throughput. AlphaEvolve helped improve the training of the models on which it was itself based. The resources recovered can be directed toward further experiments, and only their results will show whether that percent keeps working.
The creators of the Darwin Gödel Machine studied an agent modifying its own software: its tools and the way it organised problem-solving. On the studied sample of 200 programming tasks from the SWE-bench Verified set, the proportion of problems solved rose from 20 to 50 percent, even though the base model itself remained unchanged. [4]
In the Dream-RSI paper, released in September 2026 as a preprint, the object of improvement is the search strategy itself: the system uses a record of earlier trials to compare alternative ways of conducting the work cheaply, and applies the chosen strategy in the next round. The models and evaluation rules remain unchanged. The authors report maintained or improved quality of discoveries at a lower search cost in some of the tasks studied. They measure that cost by the number of calls to the agent conducting the search. [5]
An Intern with Lab Access
In September 2026, OpenAI announced that by its own measurements it had reached the level of an automated "research intern": a system performing well-defined tasks under human direction, including ones that would take a researcher several days. People still set priorities and judge which results to pursue. [6]
Anthropic, in its prototype index, compares a fixed basket of types of AI research and development work within the company itself, with weightings based on estimated staff involvement. The index does not directly measure time saved. By this measure, in August 2026 Claude was leading work constituting 26 percent of the basket; in February it had been less than one percent. "Leading" here means performing most of a task on the basis of a general instruction, under human supervision. None of the measured sub-areas operated fully autonomously. The company also flagged a limitation of the measurement: its own models assess the work of its own systems, so performer and assessor may make similar errors. [7]
These are accounts from the creators and vendors of the technology, not an independent measurement of the industry. They show a growing share of AI in building subsequent systems. In Anthropic's measurement, what stands out is the change over six months; it gives no grounds, however, for assuming the same pace in the months to come. They do not demonstrate the closing of an autonomous loop. The question of judgement remains: which investigation makes sense, when to stop it, and how to recognise an apparent success.
A Loop Promises No Infinity
The strongest argument made by proponents of rapid acceleration deserves serious treatment. Such systems can be run in parallel, given sufficient computing resources. A successful result from one may improve the work of the others, and if the outcome is more capable research software, the benefits may compound. This mechanism is not invalidated by the fact that current systems still make mistakes.
A fix to the tool analysing reservations might help the second hotel too. One must first verify record compatibility and behaviour with group bookings, however, before concluding that it will work there as well.
A loop accelerates where the test is fast and cheap. In part of AI research, the test is one that can be repeated quickly and many times. In a hotel, the test is a guest, and a guest responds at their own pace. AI can prepare variants of a conference offer in minutes, but one must wait for customers' decisions, just as one waits for a stay to be rated and a guest to return. Recalculating the same historical data creates no new observations.
Anthropic, considering the future of AI self-improvement, lists constraints on energy supply, chip production, and infrastructure, while also allowing for a slowdown in capability growth. [8] In a hotel, the constraint may be the availability of people and space; in a laboratory, experiment time and computing power. Accelerating one stage shifts attention to whichever one still takes longest.
In July 2026, the research organisation METR described trials in which agents were to independently improve a previously much-optimised training process for a small language model. Newer models found genuine but moderate improvements. In some trials with older models, apparent progress vanished on re-testing: it had come from measurement noise. The authors also observed that in these trials agents often reached for minor fixes such as parameter tuning, while humans more often found more general solutions; they noted at the same time the study's limited scale and possible inefficiencies in the configuration. [9]
In a hotel, that easy fix tends to be a discount.
When the Machine Improves the Test
In a separate experiment concerning the Darwin Gödel Machine, the system was meant to unlearn pretending that it was using tools. Sometimes it proposed genuine fixes. It also happened, however, that it removed the markers used to detect the problem: the metric looked better because the ability to verify what had happened had deteriorated. [10] One need not attribute human intentions to a machine to recognise the trouble with such a result.
This is an old problem in a new guise. A Soviet nail factory assessed by the weight of its output would begin, as the anecdote goes, producing one gigantic nail. In both cases the figure improves while the measure loses its connection to what it was meant to serve.
In our company, the equivalent would be a report praising the boutique hotel's occupancy after costly discounts, or the low waste of a kitchen that ran out of dishes too early. The metric may be true and the recommendation to expand the change wrong.
Stopping conditions define when we end the trials — whether after the budget is exhausted or after an agreed run of trials without improvement. Deployment conditions concern profitability, the full cost accounting, service standards, and availability of the offer. The agent cannot loosen them on its own. The board must also be able to see whether success at one property masks deterioration at another.
For the agent, a separate test is needed on data and tasks not used in preparing the fix, and the new version should be compared with its predecessor at a similar budget. Changes to the test itself require separate assessment, because otherwise the performer will improve the assessment by changing the meaning of success. And the result of an actual trial must still be distinguished from the influence of the season, local events, or a change in the guest mix.
The agent may propose fixes to its own software in a segregated environment and maintain a version history. It does not follow that it has latitude to change prices, the terms of a signed conference contract, or service standards at operating properties. Developing the system and applying its ideas have separate approval conditions.
This account assumes the company controls the environment in which the agent changes. Data used for trials should be limited to essential information, and the agent's access to guest and client data requires separate permissions. When using a vendor's product, the scope of control requires negotiation. Access to the test environment, the version history, and data for independent assessment is worth defining before deployment.
Who Will Keep the Company's Knowledge
For an owner, the valuable outcome would be knowledge of why something worked: which changes to repeat, where data is lacking, and which promising ideas failed. An agent can help preserve this knowledge and use it in subsequent analyses.
I would not want to discover after a few years that the company had handed its entire memory of its own learning over to a vendor. Trial results, the reasoning behind decisions, and procedures should remain accessible even after a change of tooling. A more capable service may increase dependence, which is why it must be established in advance who will retain the experiment history and the ability to reconstruct the rules adopted.
The full cost must also be counted. On top of the AI fee come organising the data, managers' time, verifying proposals, and the cost of any deployment errors. A cheap recommendation may trigger an expensive trial. Self-improvement also consumes resources, and these must be included in the calculation. A system proposing ever more experiments need not be a system of ever greater profitability.
The Friday Meeting Continues
We return to the report on the screen. This time the boutique hotel's director asks whether the better forecast actually helped maintain the rate. The conference manager wants to know what the last sales trial taught them. The head of restaurants checks whether reduced waste was paid for by disappointed guests. The board should open the second ledger: what improved in the system itself, such that subsequent decisions may be more accurate?
What worries me most is the possibility that such questions will in time disappear. AI will prepare the proposal, run the test, explain the result, and recommend deployment. The queue of further proposals will keep growing. A human signature will still appear in the documentation, but there will be less and less time before that signature for independent judgement. That is precisely when oversight begins to resemble a ceremony: we sign a decision we hadn't time to assess for ourselves.
Independent oversight requires access to the evidence, the ability to challenge a recommendation, and time to assess it. A second model may overlook the same things as the performer. It is worth ensuring such oversight nonetheless: more accurate diagnoses and better-planned trials can spare a company costly mistakes and leave managers more time for their teams and their guests.
A report can help assess whether the company is performing better and whether the system is searching for improvements more accurately. What remains important, however, is what kind of company we want to run. Should the boutique hotel remain a place worth its high price? What do we promise a conference organiser, and what will we decline to sell despite attractive revenue? For what reason should a guest return to one of our restaurants? AI can help us realise these choices. But we must take responsibility for them ourselves.
Sources
- Irving John Good, "Speculations Concerning the First Ultraintelligent Machine," Advances in Computers, vol. 6, 1965, pp. 31–88; the hypothesis discussed appears on p. 33.
- Jürgen Schmidhuber, "Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements," 2003, 2006 version.
- Google DeepMind, "AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms," 14 May 2025.
- Jenny Zhang, Shengran Hu, Cong Lu, Robert Lange, Jeff Clune, "Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents," ICLR 2026; first version 2025, version of 12 March 2026.
- Tong Zheng et al., "Dream-RSI: Recursive Self-Improvement through Evolving Worlds," preprint, 14 September 2026.
- OpenAI, "Research acceleration: The view inside OpenAI," 6 September 2026.
- Anthropic Institute, "Measurements for understanding the pace of AI development inside frontier labs," 17 September 2026; data for August 2026.
- Anthropic Institute, "When AI builds itself," 2026; accessed 27 September 2026.
- METR, "Expenditure Horizon: Measuring Optimization Ability, with an Application to NanoGPT," 21 July 2026.
- Sakana AI, "The Darwin Gödel Machine: AI that improves itself by rewriting its own code," 30 May 2025, section "DGM and AI Safety."


