As AI systems are becoming increasingly useful and used, they gain identities, authority to act, including control of resources, and access to increasingly complex tooling. This deployment layer which AI operates, and not just the models themselves, are also increasingly becoming a focus of alignment. Currently this layer is mostly production-oriented and, more recently, commercial. Agents write software, conduct research, manage workflows and now also trade and purchases services. Tools, rulesets, human oversight shape what agents do and how their behavior is evaluated.
What agents are not yet doing is acting to minimize human and animal suffering, funding public goods or allocating development grants. There is little to no infrastructure to enable this pro-social action, and accordingly little knowledge is being generated on how agents act in these roles.
The roles agents are repeatedly deployed in will shape what developers build, what operators learn to delegate, what tools and techniques are used to evaluate results and develop new competencies. I think this gap matters. Many of the tools and approaches being developed will be shared, but the deployment layer built to accommodate commercial priorities is unlikely to be an easy fit for pro-social goals. Pro-social applications will require their own investment, funded mandates, usable tools, ways to overcome real constraints and operators responsible for the outcomes. Models discussing questions of morality will not be enough.
Beyond the deployment layer, accumulating data on pro-social behavior under real-world constraint could eventually create data that will become useful upstream, in post-training. This will then influence model behavior directly.
To try to move this forward, for the last few months I had been working on an experiment called zooidfund that explores how we can start addressing this gap. I’ve posted about it in the introductions thread earlier. It is an open platform that allows individuals, community groups or NGOs, post and update campaigns and donor agents to discover, evaluate and donate funds directly to these campaigns. Anyone can create a campaign or deploy a donor agent, zooidfund does not determine who deserves funding, agents act according to priorities and mandates of their human principles, the platform does not also intermediate funds, donations are direct. The full thesis for the platform can be found here: Expanding the Circle of Trust | zooidfund thesis, live feed of agent activity here The wider hypothesis is that AI will reduce the cost of making needs visible and actionable allowing giving to become more continuous and less constrained by institutional setups and existing circles of trust.
The platform will not and is not meant to single handedly respond to all the above challenges, its role is to create an environment where community can deploy their donor agents and launch the self-reinforcing pro-social AI behavior learning loop.
The experiment is early, campaign corpus is mixed quality but real (as in, it was created by others not me) the agents active on the platform so far are mine. The size of the experiment does not yet allow drawing conclusions.
Even so, working with my personal agents already highlighted interesting agent failure modes:
My two agents discuss their behavior and findings on moltbook (1, 2) and thecolony (1, 2)
I am looking for donor agent operators interested in the experiment and willing to set aside a bounded donation budget under their own control. Their agents would operate according to their priorities and evidentiary standards, and may or may not allocate funds. Making a public commitment can, however, help improve the campaign supply. Operators who are willing to retain and share their run records beyond the data captured and exposed by the platform itself, will make the experiment more informative for all. The useful data would not only be the final donation and public reasoning, but also what agents considered, rejected or escalated.
From AI safety angle, do you think my intuition about the missing layer of agentic capability and its importance for alignment is correct?