Compassion Aligned Machine Learning (CaML) is conducting research and deploying benchmarks with the aim of implementing broad compassion for all sentient beings into AI systems. In so doing, we’re attempting to integrate ideas from across the broad field of alignment research, including those focussed on specific technical questions, and those considering more strategic questions around who or what AI should be aligned with. To do this, it's important for us to understand how the alignment community is conceptualizing and prioritizing different ideas that CaML cares about.
In this spirit, we recently conducted a survey on controversial questions in alignment, to encourage more discussion on topics important to CaML, and roughly gauge where both core researchers and wider community members stand. Questions ranged from the value of benchmarking as eval awareness increases, to the implications of AGI for animals, to thoughts about AI s-risks. We’ve shared some interesting findings from this survey below.
Our intent with this survey was not to conduct an exhaustive analysis of the community’s position on every issue. This was an initial probe primarily designed to generate discussion and feedback on issues important to CaML. That said, we’d love to see research more accurately tracking core areas of (dis)agreement, consensus and controversy within alignment research. If we know where there is uncertainty, we can better understand where our research should focus next.
Our survey consisted of two stages.
In the first, we shared 14 polls with a panel of 14 researchers, where they privately responded to statements on an 11-point agree-disagree scale. They could also share the reasoning behind their answers. We received responses from the following 14 researchers:
Please note that respondents took these polls in an individual capacity, and responses do not represent the views of their employers
In the second stage, we released the survey totalling 15 polls (with one additional poll to the panel survey) to the broader community. Respondents answered via an EA Forum post (where all the polls are still viewable), and could publicly share their reasoning for different responses. After feedback from the expert panel, we included clarifications alongside the polls. We received a total of 1444 votes over all 15 polls, and over 100 comments providing further reasoning.
Once the survey was completed, we analysed two primary variables across the 15 questions: the lean (whether respondents on average voted agree or disagree), and standard deviation (SD) to see the spread of responses on a given question. Lean is measured from -5 meaning strongly disagree to +5 meaning strongly agree (with 0 being neither agree or disagree).
In two polls, we asked explicitly about the value of S-risk work:
On both of these questions, the community leant towards disagreeing with the statement (lean of -1.2 on poll six, and -2.0 on poll eight). Only 3 of 73 respondents agreed that s-risk progress was on track and not one strongly agreed (lean of more than 1.5).
Poll six saw the biggest distance between the community and the expert panel, with the panel leaning agree (+1.2) that s-risk work was more valuable at the margin. The panel was somewhat split on this though: this poll had the highest standard deviation (3.2) of the set. This may have been somewhat representative of the different research interests of the panel, with some explicitly more focussed on s-risks than x-risks or vice versa.
Nevertheless, both the panel and the community disagreed that s-risk work was on track. Some text responses gave more detail as to why:
Regardless of how respondents prioritise s-risks, there was at least agreement that, as it stands, what progress actually looks like in the s-risk space is difficult to identify, be that due to the field’s far-future focus, small size, or general uncertainty. This suggests a primary challenge for those currently focussed on s-risk research: identify and communicate how progress could be measured in concrete ways in the field, and consider what robustly good things can be done under uncertainty.
The final poll in the community survey asked ‘Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?’ This question was not shared with the panel, but it had the highest rate of agreement on the EA Forum of any poll: 89% of respondents agreed (lean +2.4, SD 1.9). Few voters articulated their reasoning on this question, though some responders did add that care must be taken in how this value is conceived of and pursued:
This poll points directly at CaML’s core mission and suggests broad support among respondents for the kinds of values we’d like to instil in AI systems. What we need to show to labs and the alignment community now, is how broad compassion can be implemented alongside other important values in a way that is safe and robust.
‘Non-human suffering is neglected by the AI safety community’ has the highest standard deviation among the community polls (3.1). In agreement, one respondent said: ‘as I understand, many key players are explicitly anthropocentrists. Hell, some key players care more about money than humans, let alone non-humans.’ In disagreement, another respondent said work had already been done to make Claude care about non-human suffering, and that it wasn’t even clear if this was a valuable direction to pursue (due to potential backfire effects, and the generalisation of human alignment to non-humans). Some of the disagreement here may have been due to the respondents’ reference point: those voting agree were more likely to mean ‘non-human suffering is neglected relative to the moral stakes’; disagree voters saw it as ‘neglected relative to common sense views’’. This was splits was visible in some panel responses —
Given that CaML’s positioning is very strongly in favour of more non-human representation in AI Safety and development, the lack of consensus here represents a challenge for us and similar organisations: if we want to increase representation of non-human interests in AI safety, how do we demonstrate the importance of this work? This is something we hope to emphasise more in forthcoming research on scaling compassion, and our animal cruelty benchmark currently in development.
As above though, there was strong support for instilling compassion as a value within AI systems. CaML is specifically interested in broad compassion that is robust and generalisable across different classes of sentient beings.
The community’s second strongest result overall was that 84% agreed that deceptive AIs will evade mechanistic interpretability tools to hide their features from humans and operators. The panel was more evenly split on this (lean of +0.4, SD 2.6), with some respondents saying it would depend on specifics:
As mentioned above, these polls should not be interpreted as an exhaustive analysis of where the alignment community stands. We wanted to share the results primarily to encourage further discussion and work on these questions, as well as get feedback on CaML's research direction.
A few caveats on the above findings:
Again, we’d be excited to see a more rigorous and in-depth study of opinion in the alignment space, similar to what Lucius Caviola has done for digital minds here.
Over the next few years, the alignment community will have to move quickly and respond to changing circumstances. Having a clear representation of what the community values and believes will be crucial in coordinating effectively, making convincing arguments to external stakeholders, and pushing on new research directions in areas of uncertainty.
This survey has suggested support for compassion to be instilled within AI systems, but also disagreement about whether non-human interests are neglected. CaML will continue to produce research to inform how and why compassion can be instilled within these systems, and make the case for why that compassion should be broad, accounting for the welfare and interests of humans, animals, and potentially sentient AI systems.
We’d like to thank BlueDot Impact for supporting this work, and Scott Alexander for sharing the community poll with the Astral Codex Ten audience.