Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month's polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 15 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER), Jeff Sebo (NYU), and Tobias Baumann (CRS). We'll compare panel and community responses in an upcoming report, which we'll publish here and on LessWrong. To get notified when it's released, you can subscribe to our new Substack.
Many people we've talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML's research agenda.
A few final things about the polls themselves:
- Timeframe: unless a statement says otherwise (e.g. post-AGI), read forward-looking claims as being about roughly the next 2 years.
- We're not trying to find the 'right' answers. Please answer based on your own best guess.
- % agree is your % credence in a given position
- Please let us know if you think the questions are ambiguous or embed false assumptions
- Any further engagement with the content of the polls in the comments is encouraged
Thanks to BlueDot Impact for funding this work.
¹ This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
² “Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
³ This is about where the next dollar is best spent, not about which area you think is more important overall.
⁴ "Role-playing” means the behavior arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
⁵ This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
⁶ You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
⁷ This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).
8 AIs that are trying to hide features of themselves from humans and operators
*By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.
Hello. Thanks for sharing. I had already shared the thoughts below with Jasmine a few weeks ago. I am publishing them here in case others find them useful.
I appreciate the intentions behind the survey, and I would like to take part in principle. However, I do not know how to answer the questions in practice. I feel like they would have to be operationalised in much more detail for me to give a probabilistic forecast.
I can see "post-AGI" having already been achieved, or never being achieved depending on how it is defined.
I do not know what "Animal suffering" means. Does moving away from noxious stimuli count as suffering? Does it require metacognition? One could say it refers to suffering in the phenomenal sense (negative qualia), but I do not think phenomenal consciousness exists (I endorse illusionism), or that sentience as traditionally conceptualised is relevant for ethics (relatedly).
I understand "Animal suffering" is that of animals with non-trivial moral weight, but I do not know what this means in practice given my large uncertainty. I can see the moral weight of shrimps ranging from 10^-12 to 1. I could speculate about a distribution, get its mean, and then compare it with a guess for what is trivial moral weight. However, I feel like my view is better described by "I have basically no clue about whether shrimps have trivial moral weight or not". Likewise for many of the questions in the survey when I try to think about potential operations.
With respect to "Benchmarks will become useless", I think it matters e.g. whether we are talking about 90 %, 99 %, or 99.9 % of current benchmarks, whether we are talking about current benchmarks, or benchmarks at some point in the future (what point?), and the degree to which they will become useless (e.g. 90 %, 99 %, or 99.9 % useless).
These are not minor details for me in the sense my answers could range from strongly disagree to strongly agree depending on definitions. The results of the survey above could still provide some vibes about what panelists are thinking, especially if they are repeated across time (as I understand you are doing). However, I think they will still be very difficult to interpret. As a rule of thumb, I would clarify the questions up to a point where they could be published on Metaculus. I understand this would decrease engagement. On the other hand, with the current methodology, I will put very little weight on the results. I think the results will be influenced a lot by how people are interpreting the questions, and I also do not trust forecasts produced in a few minutes about such hard questions.
All this said, your methodology may well make sense given your target audience and goals.
Thanks for you comment Vasco, and appreciate your concerns about the methodological validity. I'm likely to take more of a lead on producing these surveys in future, and will be thinking about how to give the questions the right amount of specificity. You're right though that it's possible that there's going to be trade-offs between added clarification and accessibility.
I definitely had to change my answer on a question quite a lot because I realized from reading one of your comments that you didn't mean what I thought. Hopefully people's written justification helps more, since that probably tells you how they were interpreting it when they answered.