Do you have anything you recommend reading on that?
I guess I see a lot of the value of people at labs happening around the time of AGI and in the period leading up to ASI (if we get there). At that point I expect things to be very locked down such that external researchers don't really know what's happening and have a tough time interacting with lab insiders. I thought this recent post from you kind of supported the claim that working inside the labs would be good? - i.e. surely 11 people on the inside is better than 10? (and 30 far far better)
I do agree OS models help with all this and I guess it's true that we kinda know the architecture and maybe internal models won't diverge in any fundamental way from what's available OS. To the extent OS keeps going warning shots do seem more likely - I guess it'll be pretty decisive if the PRC lets Deepseek keep OSing their stuff (I kinda suspect not? But no idea really).
I guess rather than concrete implications I should indicate these are more 'updates given more internal deployment' some of which are pushed back against by surprisingly capable OS models (maybe I'll add some caveats)
I think it's very reasonable to say that 2008 and 2012 were unusual. Obama is widely recognized as a generational political talent among those in Dem politics. People seem to look back on, especially 2008, as a game-changing election year with really impressive work by the Obama team. This could be rationalization of what were effectively normal margins of victory (assuming this model is correct) but I think it matches the comparative vibes pretty well at the time vs now.
As for changes over the past 20+ years, I think it's reasonable to say that there's been fundamental shifts since the 90s:
Polarization has increased a lot
The analytical and moneyball nature of campaigns has increased by a ton. Campaigns now know far more about what's happening on the ground, how much adversaries spend, and what works.
Trump is a highly unusual figure which seems likely to lead to some divergence
The internet & good targeting have become major things
Agree that 5-10% probability isn't cause for rejection of the hypothesis but given we're working with 6 data points, I think it should be cause for suspicion. I wouldn't put a ton of weight on this but 5% is at the level of statistical significance so it seems reasonable to tentatively reject that formulation of the model.
Trump vs Biden favorability was +3 for Trump in 2020, Obama was +7 on McCain around election day (average likely >7 points in Sept/Oct 2008). Kamala is +3 vs Trump today. So that's some indication of when things are close. Couldn't quickly find this for the 2000 election.
I think this is all very reasonable and I have been working under the assumption of one votes in PA leading to a 1 in 2 million chance of flipping the election. That said, I think this might be too conservative, potentially by a lot (and maybe I need to update my estimate).
Of the past 6 elections 3 were exceedingly close. Probably in the 95th percentile (for 2016 & 2020) and 99.99th percentile (for 2000) for models based off polling alone. For 2020 this was even the case when the popular vote for Biden was +8-10 points all year (so maybe that one would also have been a 99th percentile result?). Seems like if the model performs this badly it may be missing something crucial (or it's just a coincidental series of outliers).
I don't really understand the underlying dynamics and don't have a good guess as to what mechanisms might explain them. However, it seems to suggest that maybe extrapolating purely from polling data is insufficient and there's some background processes that lead to much tighter elections than one might expect.
Some incredibly rough guesses for mechanisms that could be at play here (I suspect these are mostly wrong but maybe have something to them):
Something something polarization, steady voting blocs for Rep & Dem aren't shifting much year to year. This means we should expect similar margins this year as 2016 & 2020.
Some balancing out process where politicians are adjusting their platform, messaging, etc to react to their adversary and this ends up increasing how close elections get.
Maybe something where voters have local information on whether the person they don't like is more likely to win and they then feel more motivated to vote? Turns out, in aggregate, this local information is pretty accurate and leads to tighter-than-expected elections.
Maybe political parties/donors observe how much their adversary spends in a given state and are consistently able to spend to counteract their efforts. This maybe provides a balancing effect that tightens the race. This would have the unfortunate consequence that visible spending is much less effective - but maybe implies that smaller, more under-the-radar, projects are better.
Hey Mathilde! Thanks for your thoughtful comment. Curious to hear the mechanism behind eating too many healthy things leading to your issues.
Also, interesting about the supplements, hadn't heard that before. I am a bit ignorant on these things but try to offset that by buying the more expensive versions of supplements when trying them for the first time.
Heartening to hear that you figured it out after a few years!
I see a lot of talk with digital people about making copies but wouldn't a dominant strategy (presuming more compute = more intelligence/ability to multitask) be to just add compute to any given actor? In general, why copy people when you can just make one actor, who you know to be relatively aligned, much more powerful? Seems likely, though not totally clear, that having one mind with 1000 compute units would be strictly better for seeking power than 100 minds with 10 compute units each.
For example, companies might compete with one another to have the smartest/most able CEO by giving them more compute. The marginal benefit of more intelligence might be really high such that Tim Cook being 1% more intelligent than Mark Zuckerberg could mean Apple becomes dominant. This would trigger an intense race for compute. The same should go for governments. At some point we should have a multipolar superintelligence scenario but with human minds.
That seems true for many cases (including some I described) but you could also have a contingent of forward-looking digital people who are optimizing hard for future bliss (a potentially much more appealing prospect than expansion or procreation). Seems unclear that they would necessarily be interested in this being widespread.
Could also be that digital people find that more compute = more bliss without any bounds. Then there is plenty of interest in the rat race with the end goal of monopolizing compute. I guess this could matter more if there were just one or a few relatively powerful digital people. Then you could have similar problems as you would with AGI alignment. E.g. infrastructure profusion in order to better reward hack. (very low confidence in these arguments)
One thing that seems interesting to consider for digital people is the possibility of reward hacking. While humans certainly have quite a complex reward function, once we have full understanding of the human mind (having very good understanding could be a prerequisite to digital people anyway) then we should be able to figure out how to game it.
A key idea here is that humans have built-in limiters to their pleasure. I.e. if we eat good food that feeling of pleasure must subside quickly or else we'll just sit around satisfied until we die of hunger. Digital people need not have limits to pleasure. They would have none of the obstacles that we have to experiencing constant pleasure (e.g. we only have so much serotonin, money, and stamina so we can't just be on heroin all the time). Drugs and video games are our rudimentary and imperfect attempts to reward hack ourselves. Clearly we already have this desire. By becoming digital we could actually do it all the time and there would be no downsides.
This would bring up some interesting dynamics. Would the first ems have the problem of quickly becoming useless to humans as they turn to wireheading instead of interacting with humans? Would pleasure-seekers just be kind of socially evolved away from and some reality fundamentalists would get to drive the future while many ems sit in bliss? Would reward-hacked digital humans care about one another? Would they want to expand? If digital people optimize for personal 'good feelings' probably that won't need to coincide with interacting with the real world except so as to maintain the compute substrate, right?
That's my bad, I did say 'automated' and should have been 'automatable'. Have now corrected to clarify
Do you have anything you recommend reading on that?
I guess I see a lot of the value of people at labs happening around the time of AGI and in the period leading up to ASI (if we get there). At that point I expect things to be very locked down such that external researchers don't really know what's happening and have a tough time interacting with lab insiders. I thought this recent post from you kind of supported the claim that working inside the labs would be good? - i.e. surely 11 people on the inside is better than 10? (and 30 far far better)
I do agree OS models help with all this and I guess it's true that we kinda know the architecture and maybe internal models won't diverge in any fundamental way from what's available OS. To the extent OS keeps going warning shots do seem more likely - I guess it'll be pretty decisive if the PRC lets Deepseek keep OSing their stuff (I kinda suspect not? But no idea really).
I guess rather than concrete implications I should indicate these are more 'updates given more internal deployment' some of which are pushed back against by surprisingly capable OS models (maybe I'll add some caveats)
I think it's very reasonable to say that 2008 and 2012 were unusual. Obama is widely recognized as a generational political talent among those in Dem politics. People seem to look back on, especially 2008, as a game-changing election year with really impressive work by the Obama team. This could be rationalization of what were effectively normal margins of victory (assuming this model is correct) but I think it matches the comparative vibes pretty well at the time vs now.
As for changes over the past 20+ years, I think it's reasonable to say that there's been fundamental shifts since the 90s:
Agree that 5-10% probability isn't cause for rejection of the hypothesis but given we're working with 6 data points, I think it should be cause for suspicion. I wouldn't put a ton of weight on this but 5% is at the level of statistical significance so it seems reasonable to tentatively reject that formulation of the model.
Trump vs Biden favorability was +3 for Trump in 2020, Obama was +7 on McCain around election day (average likely >7 points in Sept/Oct 2008). Kamala is +3 vs Trump today. So that's some indication of when things are close. Couldn't quickly find this for the 2000 election.
I think this is all very reasonable and I have been working under the assumption of one votes in PA leading to a 1 in 2 million chance of flipping the election. That said, I think this might be too conservative, potentially by a lot (and maybe I need to update my estimate).
Of the past 6 elections 3 were exceedingly close. Probably in the 95th percentile (for 2016 & 2020) and 99.99th percentile (for 2000) for models based off polling alone. For 2020 this was even the case when the popular vote for Biden was +8-10 points all year (so maybe that one would also have been a 99th percentile result?). Seems like if the model performs this badly it may be missing something crucial (or it's just a coincidental series of outliers).
I don't really understand the underlying dynamics and don't have a good guess as to what mechanisms might explain them. However, it seems to suggest that maybe extrapolating purely from polling data is insufficient and there's some background processes that lead to much tighter elections than one might expect.
Some incredibly rough guesses for mechanisms that could be at play here (I suspect these are mostly wrong but maybe have something to them):
Thanks for this!
My thinking has moved in this direction as well somewhat since writing this. I'm working on a post which tells a story more or less following what you lay out above - in doc form here: https://docs.google.com/document/d/1msp5JXVHP9rge9C30TL87sau63c7rXqeKMI5OAkzpIA/edit#
I agree this danger level for capabilities could be an interesting addition to the model.
I do feel like the model remains useful in my thinking, so I might try a re-write + some extensions at some point (but probably not very soon)
Hey Mathilde! Thanks for your thoughtful comment. Curious to hear the mechanism behind eating too many healthy things leading to your issues.
Also, interesting about the supplements, hadn't heard that before. I am a bit ignorant on these things but try to offset that by buying the more expensive versions of supplements when trying them for the first time.
Heartening to hear that you figured it out after a few years!
In response to an earlier version of this question (since taken down) weeatquince responded with the following helpful comment:
Regulatory type interventions (pre-deployment):
Defence in depth type interventions (post-deployment):
I see a lot of talk with digital people about making copies but wouldn't a dominant strategy (presuming more compute = more intelligence/ability to multitask) be to just add compute to any given actor? In general, why copy people when you can just make one actor, who you know to be relatively aligned, much more powerful? Seems likely, though not totally clear, that having one mind with 1000 compute units would be strictly better for seeking power than 100 minds with 10 compute units each.
For example, companies might compete with one another to have the smartest/most able CEO by giving them more compute. The marginal benefit of more intelligence might be really high such that Tim Cook being 1% more intelligent than Mark Zuckerberg could mean Apple becomes dominant. This would trigger an intense race for compute. The same should go for governments. At some point we should have a multipolar superintelligence scenario but with human minds.
That seems true for many cases (including some I described) but you could also have a contingent of forward-looking digital people who are optimizing hard for future bliss (a potentially much more appealing prospect than expansion or procreation). Seems unclear that they would necessarily be interested in this being widespread.
Could also be that digital people find that more compute = more bliss without any bounds. Then there is plenty of interest in the rat race with the end goal of monopolizing compute. I guess this could matter more if there were just one or a few relatively powerful digital people. Then you could have similar problems as you would with AGI alignment. E.g. infrastructure profusion in order to better reward hack. (very low confidence in these arguments)
One thing that seems interesting to consider for digital people is the possibility of reward hacking. While humans certainly have quite a complex reward function, once we have full understanding of the human mind (having very good understanding could be a prerequisite to digital people anyway) then we should be able to figure out how to game it.
A key idea here is that humans have built-in limiters to their pleasure. I.e. if we eat good food that feeling of pleasure must subside quickly or else we'll just sit around satisfied until we die of hunger. Digital people need not have limits to pleasure. They would have none of the obstacles that we have to experiencing constant pleasure (e.g. we only have so much serotonin, money, and stamina so we can't just be on heroin all the time). Drugs and video games are our rudimentary and imperfect attempts to reward hack ourselves. Clearly we already have this desire. By becoming digital we could actually do it all the time and there would be no downsides.
This would bring up some interesting dynamics. Would the first ems have the problem of quickly becoming useless to humans as they turn to wireheading instead of interacting with humans? Would pleasure-seekers just be kind of socially evolved away from and some reality fundamentalists would get to drive the future while many ems sit in bliss? Would reward-hacked digital humans care about one another? Would they want to expand? If digital people optimize for personal 'good feelings' probably that won't need to coincide with interacting with the real world except so as to maintain the compute substrate, right?