AI disclosure: This text was drafted with AI assistance from my RSI explainer. It presents a proposed way to think about oversight, not new empirical results.
Suppose a research organization runs a thousand AI agents. One discovers a better tool. The organization tests it and makes it available to the others.
That is an attractive use of horizontal scaling: a useful discovery can benefit work beyond the experiment that produced it. If better tools help agents develop further improvements, the process also has a recursive component.
The acceptance decision now has consequences across the organization. A mistake in that decision can spread along with the software.
This seems relevant to oversight even before we have a system capable of autonomously training a successor foundation model. Agents can modify tools or working methods around an unchanged model. The question of which changes are safe to distribute already exists at that level.
A useful distinction is between authority to propose a change and authority to accept it. An agent could have broad freedom to search in a sandbox while lacking permission to modify the held-out evaluation or distribute a new version to other agents.
That separation would not solve alignment. Evaluators can miss failures, and several evaluators built around the same model can share blind spots. It would, however, make an otherwise easy-to-hide decision explicit: who or what approved this change, on what evidence, and with what scope?
I'd want retained versions, recorded evaluation conditions and a way to stop further distribution when a problem appears. Rollback has limits: reverting software cannot undo information already disclosed or an external action already taken.
The last-invention hypothesis asks whether humans could become unnecessary to the discovery of further technology. Even in that possible future, deciding which discoveries to adopt would remain consequential. Automating more of the search does not settle what we should permit the resulting systems to do.
For people evaluating automated AI research, which part of this acceptance process seems hardest to preserve as the number of parallel experiments grows?
Background and sources: RSI AI explained