An eccentric dreamer in search of truth and happiness for all. I formerly posted on Felicifia back in the day under the name Darklight and still use that name on Less Wrong. I've been loosely involved in Effective Altruism to varying degrees since roughly 2013.
That is sorta the idea yes. Agents would choose this decision criteria mostly because it vastly increases their odds of survival, which allows them to further whatever goals they have. I would hope that this result is obvious enough that many agents will be able to converge on it, increasing the proportion using it, and thus increasing the overall survival rate.
The other takeaway is that, given that humans will be weaker than AGI/ASI, any game theoretic reason for such entities to still cooperate with us can potentially help reduce the existential risk.
I agree that the model requires further scrutiny to determine if it is realistic enough to matter.
I have this game theory thing I've been working on that involves modifying the Iterated Prisoner's Dilemma to include death, asymmetric power, and aggressor reputation. Agents' points are their "power" that dynamically impacts their payoff matrix.
The basic takeaway is that this simple simulation seems to make a case for cooperating with weaker agents, by showing how the cooperative strategies outcompete the aggressive ones in the long run. I think, before I can make a proper post about it, I'll need to run some analysis to graph out how, for instance, having a higher percentage of cooperative agents increases the odds of survival, which implies a kind of Veil of Ignorance logic towards being cooperative.
Note that I mean cooperative in the sense that you don't defect first except against aggressors that have defected first against non-aggressors.
With default settings, the most common result of any given run is that a significant number of the cooperative strategies survive and almost all of the aggressive ones die out. Very occasionally, particularly if you adjust the settings are certain way, a single "Opportunist" strategy, that Tit-For-Tats against stronger agents and Defects against weaker ones, will be the only survivor. This seems to imply, at least, to me, that being a cooperative strategy significantly increases your odds of survival, as the alternative is to hope to win a "Highlander" scenario.
I think this is relevant to AI alignment as a variation on Anthropic Capture, the "Hail Mary" approach that Bostrom mentions in Superintelligence. It could work as part of a defence-in-depth, a kind of "infoblessing" that could persuade some AGI to spare us as a kind of Superrational Signalling. While you might assume this only works if aliens are probable, it also functions in a multi-agent scenario where there are several AGI at near peer levels of power. It also potentially could be a way to align a previously unaligned AGI even after it is deployed. If enough AGIs are aligned in this way, their alliance could defeat the unaligned AGIs.
You can run the simulation yourself here: https://paxscientia.com/power/
I have the code and initial analysis here: https://github.com/josephius/power
I realize that a very obvious critique of this work is that the simulation is probably too simple. I intentionally tried to keep it an MVP in its first iteration. I also should, as mentioned earlier, complete a more thorough and rigorous analysis of the apparent results. I'm also keenly aware that it seems like this is a "neglected" path towards alignment, and I'm uncertain whether this is because the idea is a bad one that's already been discarded by others who are more competent. I know that there are related ideas around Decision Theory, Acausal Trade, and Superrationality, but I've never seen this particular kind of effort, which confuses me, because it seems obvious and trivial to try.
My main question to ask is simply, does this seem like something worth pursuing and expanding further, or am I wasting my time on a foolish endeavour?
Yeah, getting something to be both meaningful and fun at the same time is hard. I took a quick look at your prototype. It could have some potential, but at the same time, it's not the only game out there with a similar idea. I recently came across The Choice Before Us, which is in a similar vein.
Given how fast things are moving in terms of AI developments, I'm not sure it's realistic to try to make a game that's polished enough to go viral before things change enough that the game is essentially obsolete.
Also, games are hard. Creative work seems very lottery-like in terms of success. Maybe you can argue from a risk neutral EV perspective that it's worth it, but that doesn't pay the bills.
Interesting.
What if there was a way to conceal your RKV or invasion fleet, like some kind of cloaking device or camouflage technique?
And what about spamming an overwhelming number of RKVs or invasion fleets? In theory, you might be able to saturate defences and win a war of attrition if you have a significantly stronger industrial base? If one civilization is already a billion years ahead of another, wouldn't that one be likely to have such an insurmountable lead in control of energy resources that they could simply use sheer numbers?
Also, it's a big assumption to make that there's no wormholes/FTL/time travel shenanigans possible.
With wormholes, it would be possible to, for instance, send an invasion probe with one end of a wormhole, and when it arrives, send forces through from the other end of the wormhole. However, defenders could also have a vast network of wormholes that would allow near instant transfer of forces. If the defenders have good sensor networks, they could also destroy the incoming wormhole probes. Still, if even one wormhole probe gets into the galaxy, the supply lines suddenly don't look so one-sided, and you could, again, spam lots of these probes in the hopes that one makes it there, and then you flood into the galaxy through the wormhole bridgehead.
With FTL, in theory, you could send an RKV with it that their sensors might not be able to detect in time, unless there are FTL sensors, which would negate this. But, if FTL is possible, it's probably equivalent to time travel...
A civilization with time travel is likely to at least know in advance when an attack is coming. This would benefit defence by making the element of surprise even more impossible. And taken to its logical extreme, time travel would likely be used to ensure that no other threatening civilizations come into existence in our lightcone in the first place. This would, in practice, create a universe where every civilization exists in its own otherwise empty bubble, not unlike your predicted universe, but for different reasons than defence dominance.
At least, those are my initial thoughts.
Edit: Also, being able to see the acceleration and deceleration plumes assumes we never develop something more efficient, like say, a light sail, or some kind of beam thruster with almost no beam divergence. These, I imagine, would be harder to detect?
I will take a much stronger stand and also say that I'm very sympathetic to the argument by creatives that using generative AI is inherently unethical for a wide variety of reasons, ranging from the nature of how the data is collected without permission, to the environmental and societal impacts, to the way it degrades our culture through slop.
I'll also throw in from a more doomer perspective that by using these models, we are supporting the AI industry, giving them data and (if you are subscribing) money to rush toward building ever more dangerous systems (even if LLMs don't pan out, the sheer amount of research effort now going into AI makes things likely to reach critical mass) that can already cause things like AI psychosis, are nearing the point of mass replacement of labour with capital, and could one day kill us all.
I do not subscribe to any chatbot services, and have only experimented with free versions to a extent, and have never used them for my coding or writing.
I personally, am seriously considering joining PauseAI and possibly, additionally, boycotting AI products. This from someone who used to be an ML researcher before it was cool.
I'm very much aggrieved to see my life's work used for such tremendous evil as it is now.
I think EA should take a much stronger stand against AI. The public backlash is already starting, and for once, we should pick a side. If the most recent METR trendlines are right, we don't have much time left.
I have a particular writing style that I consider my "voice", and I fundamentally take pride in my writing skill and see writing as a craft and art form, so I refuse to use AI to write a single word of what I would publish to the world.
To me, using AI for writing is equivalent to having someone else write it for you.
Not quite a draft amnesty thing, but I have been playing with the idea of writing short stories or perhaps even novels that use time travellers as a vehicle for Longtermism. The idea is that time travellers from the far distant future are our descendents, the very people that Longtermism cares about, so their perspective could be something worth exploring in fiction.
Given, I'm more of a soft Longtermist, and creative writing is notoriously hard to make any kind of living out of, so I'm not sure to what extent this is worth doing/trying/exploring, even as just a side project.
I've been on this forum since 2014 and I -still- feel this way sometimes. Although, I will say Less Wrong is notably worse for this.
It does get better after you make a few comments/posts and notice people aren't jumping all over you. I used to be much more terrified, but now, I'm only kinda apprehensive whenever I post.
I've explored very similar ideas before in things like this simulation based on the Iterated Prisoner's Dilemma but with Death, Asymmetric Power, and Aggressor Reputation. Long story short, the cooperative strategies do generally outlast the aggressive ones in the long run. It's also an idea I've tried to discuss (albeit less rigorously) before as The Alpha Omega Theorem and Superrational Signalling. The first of those was from 2017 and got downvoted to oblivion, while the second was probably too long-winded and got mostly ignored.
There are a bunch of random people like James Miller and A.V. Turchin and Ryo who have had similar ideas that can broadly be categorized under Bostrom's concept of Anthropic Capture, or Game Theoretic Alignment, or possibly a subset of Agent Foundations. The ideas are mostly not taken very seriously by the greater LW and EA communities, so I'd be prepared for a similar reception.
We tried earlier. Carrick Flynn received substantial support from EA and the result was mediocre, with criticisms of EA actually having a negative effect on his campaign, as people pointed out the connection to the "billionaires and techbros" who apparently fund EA and such.
Also, the head of RAND, Jason Matheny, is an EA, and there's some connections between EA and the American NatSec establishment. CSET for instance was funded partly by OpenPhil. There is a tendency among a lot of EAs is to try not to be partisan and mostly support effective governance and policy kind of things.
That being said, Dustin Moskovitz, the billionaire who is the main donor behind what was previously called Open Philanthropy and is now Coefficient Giving, has donated significantly and repeatedly to Democrats. OpenPhil has historically been by far the largest funder of EA stuff, particularly since SBF fell from grace, so Dustin's contributions can be seen tacitly as EA support for the Dems.
So, I don't think it's accurate to say EAs have made absolutely no effort on this front. We have, and it has stupidly backfired before and we're in this very awkward position politically where the whole TESCREAL controversy makes the EA brand tarnished to the Left, even though past surveys have shown that most rank and file EAs are centre-left to left. It's a frustrating situation.