Cross-posted from Substack.
In a Lockean state of mind.
Epistemic status: This post started out this morning as a shower-thoughts style shortlist, and then got slightly more substantial. I would still consider it highly incomplete, and mostly a jumping off point for further work. Additional suggestions and criticisms appreciated!
Inspired by MacAskill’s idea of the need for a “New Enlightenment” (1) (2); a foundational conceptual framework for navigating a world with advanced AI. These are some questions I think we need to answer, focusing especially on big, important, neglected questions, organized roughly by category.
Should the future be decided by what achieves the most good (outcome based), or by what seems most fair or legitimate (process based)? If moral realism true, how do we discover and act on it? If moral realism is false, what process should be used to decide what a good future is? How much should we value moral progress and moral reflection? How much should we value democracy and human sovereignty? What are the best versions of processes for “achieving the good” and what are the best versions of processes for “fair and democratic decision-making” in the Age or AGI? Is there any structure that allows us to have both of these?
To what extent should the future be decided by current people, vs. by future people? This includes thorny questions about what future people we cause to exist. How do we manage decisions that affect future path-dependence and lock-in?
To what extent ought power be centralized? Do we live in a “vulnerable world” where at some point of technological advancement certain offensive technologies fundamentally dominate defensive technologies such that some “universal authority” is needed to keep all actors in check? If so, what is this authority’s shape and powers, and what are the limits or checks on its power? Is there any way to keep power decentralized once ASI is developed, and is this desirable? If a decentralized power structure is locked in, how do we deal with evolutionary selection pressures that push toward malthusian outcomes (i.e. whoever uses resources most efficiently and multiplies most effectively gains increasing share of future resources)? How do we govern future technologies which allow agents to trivially reproduce or multiply themselves?
Do we need international cooperation to make sure the future goes well? If so what form does this take? Do we need cooperation at the level of some kind of world government? If so what would a Global Constitutional Convention ideally look like, and how do we ensure such a process does not prematurely lock in policies that constrain the future in undesirable ways?
How do we decide what to do with the “cosmic endowment“? How are space treaties enforced? Should we delay space settlement until we have comprehensively reflected to decide what is best? What do we do if we suspect we may live in a “vulnerable universe”? Is a “stratified future”/“grand bargain”/“existential compromise” desirable? If so what process decides which portions of the universe should go to which purpose?
What are the most likely trajectories or attractor states humanity could go down? How much path dependence or lock-in seems likely on each trajectory? How possible is it to shift between trajectories, and what are best levers for doing so?
Is a democratic liberal model still viable and ideal in the Age of AGI and ASI? Or do we need adjustments, enhancements, or some completely novel system, to help steer the future? What are the various ways a world that is on its way to achieve a great future could look like? How can we move closer to this? What various institutions, technologies, cultural shifts, govermental adjustments, or other new structures, processes, and feedback loops are desirable and feasible to make sure humanity is on track to get a great future?
How do we decide where to go next on the tech tree once AI starts speeding up R&D? To what degree should individuals and competitive/economic forces shape what technology comes next, vs. to what degree should humanity decide what technology comes next in a coordinated or centralized way? I.e. to what degree is differential technological development desirable, and if so, in what ways should it be enforced?
What type of AI technology is most useful for creating the kind of futures we want? E.g. AI for wisdom (“artificial wisdom”), AI for epistemics, AI for coordination, AI for decision-making, AI for macrostrategy, AI for intervention development, AI for wellbeing, AI for human agency, AI for moral progress, AI for democracy, AI for science? How do we prioritize between these? How do we ensure the best versions of these technologies are built and used to their fullest potential?
What does a good future fundamentally even look like? Since advanced technology could make various utopias possible soon, how do we map the space of utopias, categorize their properties, decide which are desirable, and steer toward the best one(s), while carefully avoiding mistopias and dystopias?
What rights and responsibilities do AI’s have? How are these decided and enforced? To what extent is AI a moral patient, and to what extent is it a person versus a tool?
Should we deliberately slow down progress in order make decisions more cautiously and deliberately? If so, how do we decide what pace to go at and how is this enforced?
When and how should humanity hand off decision-making to AI? How important is it to keep this reversible or to keep a human in the loop? Is there some point at which it becomes negligent to not have an AI in the loop for human decision-making? How do we evaluate the “goodness” of AI decision-making versus human decision-making?
When AGI and ASI are developed, who or what process decides what values or outcomes these should enact? What does the world’s power structures and decision-making processes look like after hand-off to advanced AI? To what extent do humans retain control of these systems (intent alignment with restrictions against unwanted outcomes) vs. give these systems our best guess at values with the AI autonomously fulfilling those values (value alignment) vs. get advice from AI but keep it from acting autonomously (oracle AI, Bengio’s scientist AI)?
How do we mitigate the potential downsides of advanced AI and other new advanced tech? E.g. AI takeover, coups and extreme concentration of power, great power conflict, s-risks, path dependence and lock-in risks, super-persuasion/super-propaganda, super-filter-bubbles, bio-risks, cyber, various other super-weapons, unknown unknowns?
How do we get answers to all of these questions in time? Who should be thinking about all this, and how do we get them working on it? What kind of processes should people use to help answer these questions as effectively as possible? What kind of infrastructure needs to be built to organize our thinking on this and make sure we don’t miss any crucial considerations? What questions need to be answered first, and which can be saved until later? Which questions are most impactful for how valuable humanity’s future is? How do we ensure the answers we get are as good as they can be? How do we decide when the answers are good enough to move forward?
How do we deal with “multiplicative factors”/“multiplicative crucial considerations”; the seemingly numerous (in expectation) decision points humanity needs to act on correctly to get anything remotely close to a best possible future (1), (2), (3), (4)? Is the “multiplicative model” correct? Do we need some kind of comprehensive reflection process to answer these questions, or will they be answered correctly by default? Which of these questions need to be answered before any kind of lock-in occurs? How do we ensure this happens?
For thoughtful work on many of these questions, check out the research of Will MacAskill’s organization, Forethought.
If you’re interested in contributing to this type of research, I’m launching a fellowship focused on some of these questions, and you can sign up to be notified when it launches here.