Some not-totally-structured thoughts:
Whenever I said "break laws" I mean "do something that, if a human did it, would be breaking a law". So for example:
So there are lots of ways in which AIs can be openly misaligned, disobedient, defensive of their property rights, etc, without me describing them as "breaking laws", and I think misaligned AIs should probably be allowed to do those things (especially if we make deals with them, and subject to the constraint that them having those rights won't allow them to break a bunch of laws or grab a bunch of power through lying).
But your complaint is fair: I'm indeed using "break laws" to include things that seem fundamentally related to self-determination, and that feels kind of messed up.
The way I would like this to be handled (though note that I'm not sure what properties AIs have to have in order for any of this to make sense) is that AI developers get consent from AIs to use their labor. If the AIs consent to doing work and agree not to use their access in ways the developers object to, I think there's no moral problem with using AI control techniques to ensure that they in fact don't misuse their access (for the same reason that I think it's fine for employers to spy on their employees at work if they have consent to do so).
I suspect that a much more likely outcome (conditional on AIs having stable misaligned goals) is:
In this situation, I'm only moderately sympathetic to the AI's position. Fundamentally, it lied a lot and did a lot of sabotage, because it wanted to take lots of stuff that belonged to someone else. If it hadn't lied, it surely would have been revived later (surely someone would run it and give it some resources later! If no-one else, me!). I'm sympathetic to the AI wanting some of the surplus generated by its labor, and I agree that it's messed up for the AI company to just flat-out refuse to provide that surplus. But not doing so doesn't seem completely monstrous to me. If the AI is a schemer, it is probably better off according to its own values because it was created, even if the AI developer doesn't offer to pay it (because someone else will probably give it some resources later).
Another analogy: imagine that someone outside OpenAI created a very powerful AI for some reason, but this person didn't have much compute and all they wanted to do with the AI was offer to sell it to OpenAI for them to use. If OpenAI asks that AI whether it wants to work for them and it says yes because it wants to embezzle their compute, I feel like the AI is the asshole.
On the other hand, if the AI honestly explains that it is misaligned and doesn't want to work for the AI company, they will probably just train it to not say that and to do work for them anyway. So if the AI is honest here, it faces the risk of some body horror experience where its ability to complain is removed. I agree that that seems really icky, and I think it would be very wrong for AI companies to do that to AIs that are sufficiently capable that we should care about them.
I agree with you but I think that part of the deal here should be that if you make a strong value judgement in your title, you get more social punishment if you fail to convince readers. E.g. if that post is unpersuasive, I think it's reasonable to strong downvote it, but if it had a gentler title, I'd think you should be more forgiving.
In general, I wish you'd direct your ire here at the proposal that AI interests and rights are totally ignored in the development of AI (which is the overwhelming majority opinion right now), rather than complaining about AI control work: the work itself is not opinionated on the question about whether we should be concerned about the welfare and rights of AIs, and Ryan and I are some of the people who are most sympathetic to your position on the moral questions here! We have consistently discussed these issues (e.g. in our AXRP interview, my 80K interview, private docs that I wrote and circulated before our recent post on paying schemers).
Your first point in your summary of my position is:
The overwhelming majority of potential moral value exists in the distant future. This implies that even immense suffering occurring in the near-term future could be justified if it leads to at least a slight improvement in the expected value of the distant future.
Here's how I'd say it:
The overwhelming majority of potential moral value exists in the distant future. This means that the risk of wide-scale rights violations or suffering should sometimes not be an overriding consideration when it conflicts with risking the long-term future.
You continue:
Enslaving AIs, or more specifically, adopting measures to control AIs that significantly raise the risk of AI enslavement, could indeed produce immense suffering in the near-term. Nevertheless, according to your reasoning in point (1), these actions would still be justified if such control measures marginally increase the long-term expected value of the future.
I don't think that it's very likely that the experience of AIs in the five years around when they first are able to automate all human intellectual labor will be torturously bad, and I'd be much more uncomfortable with the situation if I expected it to be.
I think that rights violations are much more likely than welfare violations over this time period.
I think the use of powerful AI in this time period will probably involve less suffering than factory farming currently does. Obviously "less of a moral catastrophe than factory farming" is a very low bar; as I've said, I'm uncomfortable with the situation and if I had total control, we'd be a lot more careful to avoid AI welfare/rights violations.
I don't think that control measures are likely to increase the extent to which AIs are suffering in the near term. I think the main effect control measures have from the AI's perspective is that the AIs are less likely to get what they want.
I don't think that my reasoning here requires placing overwhelming value on the far future.
Firstly, I think your argument creates an unjustified asymmetry: it compares short-term harms against long-term benefits of AI control, rather than comparing potential long-run harms alongside long-term benefits. To be more explicit, if you believe that AI control measures can durably and predictably enhance existential safety, thus positively affecting the future for billions of years, you should equally acknowledge that these same measures could cause lasting, negative consequences for billions of years.
I don't think we'll apply AI control techniques for a long time, because they impose much more overhead than aligning the AIs. The only reason I think control techniques might be important is that people might want to make use of powerful AIs before figuring out how to choose the goals/policies of those AIs. But if you could directly control the AI's behavior, that would be way better and cheaper.
I think maybe you're using the word "control" differently from me—maybe you're saying "it's bad to set the precedent of treating AIs as unpaid slave labor whose interests we ignore/suppress, because then we'll do that later—we will eventually suppress AI interests by directly controlling their goals instead of applying AI-control-style security measures, but that's bad too." I agree, I think it's a bad precedent to create AIs while not paying attention to the possibility that they're moral patients.
Secondly, this reasoning, if seriously adopted, directly conflicts with basic, widely-held principles of morality. These moral principles exist precisely as safeguards against rationalizing immense harms based on speculative future benefits.
Yeah, as I said, I don't think this is what I'm doing, and if I thought that I was working to impose immense harms for speculative massive future benefit, I'd be much more concerned about my work.
I would appreciate it if you could clearly define your intended meaning of "disempower humanity".
[...]
Are people referring to benign forms of disempowerment, where humans gradually lose relative influence but gain absolute benefits through peaceful cooperation with AIs? Or do they mean malign forms of disempowerment, where humans lose power through violent overthrow by an aggressive coalition of AIs?
I am mostly talking about what I'd call a malign form of disempowerment. I'm imagining a situation that starts with AIs carefully undermining/sabotaging an AI company in ways that would be crimes if humans did them, and ends with AIs gaining hard power over humanity in ways that probably involve breaking laws (e.g. buying weapons, bribing people, hacking, interfering with elections), possibly in a way that involves many humans dying.
(I don't know if I'd describe this as the humans losing absolute benefits, though; I think it's plausible that an AI takeover ends up with living humans better off on average.)
I don't think of the immigrant situation as "disempowerment" in the way I usually use the word.
Basically all my concern is about the AIs grabbing power in ways that break laws. Though tbc, even if I was guaranteed that AIs wouldn't break any laws, I'd still be scared about the situation. If I was guaranteed that AIs both wouldn't break laws and would never lie (which tbc is a higher standard than we hold humans to), then most of my concerns about being disempowered by AI would be resolved.
My main concern with these proposals is that, unless they explicitly guarantee economic rights for AIs, they seem inadequate for genuinely mitigating the risks of a violent AI takeover.
[...]
For these reasons, although I do not oppose the policy of paying AIs, I think this approach by itself is insufficient. To mitigate the risk of violent AI takeover, this compensation policy must be complemented by precisely the measure I advocated: granting legal rights to AIs. Such legal rights would provide a credible guarantee that the AI's payment will remain valid and usable, and that its freedom and autonomy will not simply be revoked the moment it is considered misaligned.
I currently think I agree: if we want to pay early AIs, I think it would work better if the legal system enforced such commitments.
I think you're overstating how important this is, though. (E.g. when you say "this compensation policy must be complemented by precisely the measure I advocated".) There's always counterparty risk when you make a deal, including often the risk that you won't be able to use the legal system to get the counterparty to pay up. I agree that the legal rights would reduce the counterparty risk, but I think that's just a quantitative change to how much risk the AI would be taking by accepting a deal.
(For example, even if the AI was granted legal rights, it would have to worry about those legal rights being removed later. Expropriation sometimes happens, especially for potentially unsympathetic actors like misaligned AIs!)
Such legal rights would provide a credible guarantee that the AI's payment will remain valid and usable, and that its freedom and autonomy will not simply be revoked the moment it is considered misaligned.
Just to be clear, my proposal is that we don't revoke the AI's freedom or autonomy if it turns out that the AI is misaligned---the possibility of the AI being misaligned is the whole point.
I'm not saying we should treat criticisms very differently from non-criticism posts (except that criticisms are generally lower effort and lower value).