I'm a computational physicist, I generally donate to global health. I am skeptical of AI x-risk and of big R Rationalism, and I intend on explaining why in great detail.
I'm okay with balancing trade-offs between rigor and reach when the rigor is reasonably high. But it's important to remember that rigor and effectiveness are correlated with each other.
When it comes to longtermism, as in trying to affect the future a thousand the rigor is essentially non-existent, and precisely because of that, I doubt there is any reach either.
First, of all, I want to say that this is an interesting experiment and a nice write-up.
However, I would caution you against taking this as a literal indicator of LLM morality, or a good indicator of how they would act in a real world context.
No actual animals are being harmed in the game you set up here, and at this point I believe LLM's are smart enough to be aware of this. This is more akin to a silly videogame than an actual real world scenario.
If I successfully train an LLM agent to play grand theft auto, it does not mean that the LLM agent loves murdering cops and running over innocent bystanders.
All too often, I see people reject societal norms as "socially dominant superstitions" only to later find out that those norms were there for a damn good reason. I might point here to SBF's maverick way of running a business, that later turned out to be a symptom of widescale fraud. A similarly cavalier attitude toward sexual norms can cause real harm. I have no issue with polyamory or orgies or whatever is going on in principle, but combine it with conflicts of interests and large power imbalances and things can go pretty badly.
One could distinguish this in a pop-bayesian way: an extraordinary claim is one that we should put an extremely low prior probability on being true, based on the nature claim alone. So, for example, quantum physics is extremely weird and out of line with what came before, before the quantum revolution we would place an insanely low probability of it being true.
But then you assess the evidence, and calculate the posterior probability based on that evidence. In order to be confident that an extraordinary claim is true, you need an extraordinary amount of evidence, like what we found for quantum physics. So I would say quantum physics is an extraordinary claim that is nonetheless extremely likely to be true.
This is not exactly rigorous (but none of this stuff is anyway). I do think existential risk is an extraordinary claim in this sense, and the prior probability for near-term extinction really should be extremely low. I personally don't think the evidence offered by the x-risk community comes anywhere close to "extraordinary", and so I believe that the overall risk is low as well.
I am a little concerned that you have immediately jumped to respond to the most controversial part of the post, rather than respond to any of the 9 uncontroversial reasons put forward that your decision to hire a convicted fraudster was a terrible decision.
I personally still don't trust Hanania, and I certainly don't think whether one supports shrimp welfare should have any bearing at all about how we feel about racism accusations, but is it really more important to discuss this rather than the fact that half of the two employees at Manifund have committed billion dollars worth of fraud?
One thing that worries me about discussions around reputation is that it seems like there is an assumption that ones reputation is independent of whether one is actually doing good or not.
Very often, the reason organisations get a bad reputation is because they are doing bad things, and the wider world is correctly associating them with those bad things. Obviously the general public can be irrational or wrong, but there are a significant amount of moral and intelligent people outside the ingroup, and their opinions deserve to be considered.
Often, the reason that decisions are perceived as bad is because they really are bad!
To add further context: The Manifund team currently consists of just two employees: CEO Austin Chen and "Carol N", the pseudonym for Caroline Ellison. Literally 50% of the employees in the organization are convicted financial fraudsters.
Now when people have their funds approved, they will know that decision was made by a convicted fraudster. The article recruiting new hires was written by a convicted fraudster (who was hiding it at the time). Irrespective of her actual role, the fact that she makes up a large percentage of the employee base gives her a significant power in the organization: it's crazy that she is apparently the first hire here.
This is a terrible decision on every dimension you can look at it. It's certainly a brave decision: that does not make it a sane or moral one.
For context, Caroline Ellison has been recently convicted for financial fraud amounting to billions of dollars. Manifund has just hired her for a finance role. They couldn't find a single qualified person that hadn't committed billions of dollars worth of financial fraud?
Nobody has to be diplomatic or mince words here, or couch this in terms of "reputation": this decision is rank incompetence, plainly obvious to anyone with a whit of common sense. Nobody should trust Manifund after this.
I think critiquing EA for a lack of moral imagination is very off the mark. Effective altruism has moral imagination in abundance, and it's one of the movements most admirable qualities.
I would say my problems with the movement is in the opposite direction: is that it lacks sufficient skepticism to temper it's vast moral imagination. Most novel ideas, including moral ideas, are simply wrong for one reason or another. The EA valorisation of novel moral ideas lets ideas be entrenched that have not passed the standards of scrutiny that is present in the actual scientific process that EA emulates.
If you look at the examples given in the article linked, they discuss a new great moral development occuring every couple of centuries. Do you truly believe that EA has come up with 5 or so such revolutionary ideas in a mere two decades?
Three anecdotes from a blog post is not exactly strong evidence. I think these anecdotes do not make the case very well. The economy returned to trend in world war II, but would it have done the same if the Axis had won? It says the american revolution had no effect on the development of freedom and democracy, but is there proof of that? It certainly played a role in the french revolution, for example.
Overall, despite the weak case, there probably are some historical trends that exist, but this doesn't mean that the world can't be radically altered by small changes. The latter is absolutely true. There is no contradiction between the two positions.
I'm okay with balancing trade-offs between rigor and reach when the rigor is reasonably high. But it's important to remember that rigor and effectiveness are correlated with each other.
When it comes to longtermism, as in trying to affect the future a thousand the rigor is essentially non-existent, and precisely because of that, I doubt there is any reach either.
First, of all, I want to say that this is an interesting experiment and a nice write-up.
However, I would caution you against taking this as a literal indicator of LLM morality, or a good indicator of how they would act in a real world context.
No actual animals are being harmed in the game you set up here, and at this point I believe LLM's are smart enough to be aware of this. This is more akin to a silly videogame than an actual real world scenario.
If I successfully train an LLM agent to play grand theft auto, it does not mean that the LLM agent loves murdering cops and running over innocent bystanders.
All too often, I see people reject societal norms as "socially dominant superstitions" only to later find out that those norms were there for a damn good reason. I might point here to SBF's maverick way of running a business, that later turned out to be a symptom of widescale fraud. A similarly cavalier attitude toward sexual norms can cause real harm. I have no issue with polyamory or orgies or whatever is going on in principle, but combine it with conflicts of interests and large power imbalances and things can go pretty badly.
One could distinguish this in a pop-bayesian way: an extraordinary claim is one that we should put an extremely low prior probability on being true, based on the nature claim alone. So, for example, quantum physics is extremely weird and out of line with what came before, before the quantum revolution we would place an insanely low probability of it being true.
But then you assess the evidence, and calculate the posterior probability based on that evidence. In order to be confident that an extraordinary claim is true, you need an extraordinary amount of evidence, like what we found for quantum physics. So I would say quantum physics is an extraordinary claim that is nonetheless extremely likely to be true.
This is not exactly rigorous (but none of this stuff is anyway). I do think existential risk is an extraordinary claim in this sense, and the prior probability for near-term extinction really should be extremely low. I personally don't think the evidence offered by the x-risk community comes anywhere close to "extraordinary", and so I believe that the overall risk is low as well.
I am a little concerned that you have immediately jumped to respond to the most controversial part of the post, rather than respond to any of the 9 uncontroversial reasons put forward that your decision to hire a convicted fraudster was a terrible decision.
I personally still don't trust Hanania, and I certainly don't think whether one supports shrimp welfare should have any bearing at all about how we feel about racism accusations, but is it really more important to discuss this rather than the fact that half of the two employees at Manifund have committed billion dollars worth of fraud?
One thing that worries me about discussions around reputation is that it seems like there is an assumption that ones reputation is independent of whether one is actually doing good or not.
Very often, the reason organisations get a bad reputation is because they are doing bad things, and the wider world is correctly associating them with those bad things. Obviously the general public can be irrational or wrong, but there are a significant amount of moral and intelligent people outside the ingroup, and their opinions deserve to be considered.
Often, the reason that decisions are perceived as bad is because they really are bad!
To add further context: The Manifund team currently consists of just two employees: CEO Austin Chen and "Carol N", the pseudonym for Caroline Ellison. Literally 50% of the employees in the organization are convicted financial fraudsters.
Now when people have their funds approved, they will know that decision was made by a convicted fraudster. The article recruiting new hires was written by a convicted fraudster (who was hiding it at the time). Irrespective of her actual role, the fact that she makes up a large percentage of the employee base gives her a significant power in the organization: it's crazy that she is apparently the first hire here.
This is a terrible decision on every dimension you can look at it. It's certainly a brave decision: that does not make it a sane or moral one.
For context, Caroline Ellison has been recently convicted for financial fraud amounting to billions of dollars. Manifund has just hired her for a finance role. They couldn't find a single qualified person that hadn't committed billions of dollars worth of financial fraud?
Nobody has to be diplomatic or mince words here, or couch this in terms of "reputation": this decision is rank incompetence, plainly obvious to anyone with a whit of common sense. Nobody should trust Manifund after this.
I think critiquing EA for a lack of moral imagination is very off the mark. Effective altruism has moral imagination in abundance, and it's one of the movements most admirable qualities.
I would say my problems with the movement is in the opposite direction: is that it lacks sufficient skepticism to temper it's vast moral imagination. Most novel ideas, including moral ideas, are simply wrong for one reason or another. The EA valorisation of novel moral ideas lets ideas be entrenched that have not passed the standards of scrutiny that is present in the actual scientific process that EA emulates.
If you look at the examples given in the article linked, they discuss a new great moral development occuring every couple of centuries. Do you truly believe that EA has come up with 5 or so such revolutionary ideas in a mere two decades?
Three anecdotes from a blog post is not exactly strong evidence. I think these anecdotes do not make the case very well. The economy returned to trend in world war II, but would it have done the same if the Axis had won? It says the american revolution had no effect on the development of freedom and democracy, but is there proof of that? It certainly played a role in the french revolution, for example.
Overall, despite the weak case, there probably are some historical trends that exist, but this doesn't mean that the world can't be radically altered by small changes. The latter is absolutely true. There is no contradiction between the two positions.