Effective Altruism Forum
EA Forum

Hide table of contents

Comment Permalink

Answer by Steven ByrnesOct 03, 20216

[note that I have a COI here]

Hmm, I guess I've been thinking that the choice is between (A) "the AI is trying to do what a human wants it to try to do" vs (B) "the AI is trying to do something kinda weirdly and vaguely related to what a human wants it to try to do". I don't think (C) "the AI is trying to do something totally random" is really on the table as a likely option, even if the AGI safety/alignment community didn't exist at all.

That's because everybody wants the AI to do the thing they want it to do, not just long-term AGI risk people. And I think there are really obvious things that anyone would immediately think to try, and these really obvious techniques would be good enough to get us from (C) to (B) but not good enough to get us to (A).

[Warning: This claim is somewhat specific to a particular type of AGI architecture that I work on and consider most likely—see e.g. here. Other people have different types of AGIs in mind and would disagree. In particular, in the "deceptive mesa-optimizer" failure mode (which relates to a different AGI architecture than mine) we would plausibly expect failures to have random goals like "I want my field-of-view to be all white", even after reasonable effort to avoid that. So maybe people working in other areas would have different answers, I dunno.]

I agree that it's at least superficially plausible that (C) might be better than (B) from an s-risk perspective. But if (C) is off the table and the choice is between (A) and (B), I think (A) is preferable for both s-risks and x-risks.

Showing 3 of 13 replies (Click to show all)

Vasco Grilo🔸Dec 9 20222

Hi Steven,

I really appreciate the dartboard analogy! It helped me understand your view.

MichaelStJules

Oct 7 2021

Ya, I think this is the crux. Also, considerations like the cosmic ray flips a bit tend to force a lot of things into the second category when they otherwise wouldn't have been, although I'm not specifically worried about cosmic ray bit flips, since they seems sufficiently unlikely and easy to avoid. (Fair.) This is actually what I'm thinking is happening, though (not like the firefighter example), but we aren't really talking much about the specifics. There might indeed be specific cases where I agree that we shouldn't be clueless if we worked through them, but I think there are important potential tradeoffs between incidental and agential s-risks, between s-risks and other existential risks, even between the same kinds of s-risks, etc., and there is a ton of uncertainty in the expected harm from these risks, so much that it's inappropriate to use a single distribution (without sensitivity analysis to "reasonable" distributions, and with this sensitivity analysis, things look ambiguous), similar to this example, and we're talking about "sweetening" one side or the other i, but that's totally swamped by our uncertainty. What I have in mind is more symmetric in upsides and downsides (or at least, I'm interested in hearing why people think it isn't in practice), and I don't really distinguish between effects by order*. My post points out potential reasons that I actually think could dominate. The standard I'm aiming for is "Could a reasonable person disagree?", and I default to believing a reasonable person could disagree when I point out such tradeoffs until we actually carefully work through them in detail and it turns out it's pretty unreasonable to disagree. *Although thinking more about it now, I suppose longer chains are more fragile and likely to have unaccounted for effects going in the opposite direction, so maybe we ought to give them less weight, and maybe this solves the issue if we did this formally? I think ignoring higher-order effects is formally i

Steven Byrnes

Oct 7 2021

I agree that direct and indirect effects of an action are fundamentally equally important (in this kind of outcome-focused context) and I hadn't intended to imply otherwise.

See in context

[ Question ]

Why does (any particular) AI safety work reduce s-risks more than it increases them?

by MichaelStJules

Oct 3 20211 min read2 answers 4