I was introduced to EA by a friend about four years ago. Around the same time, I met my first born. Through these and myriad subsequent life events, my moral circle has expanded to include all sentient beings and distant future generations, and I find myself here.
I think that, given the compensation and in-office expectations, head hunting talent that is not already motivated to work on this issue is unlikely to be super effective, especially given the requisite resource commitment on the part of the head hunters to identify prospects. You'd be soliciting people to change into a new career based on an issue they thus far have not been personally identified with despite the uncertainty around the career and respective decrease in benefits. Maybe there is a way to sort of flash motivate them, such as by your workshop proposal. But this still is pretty resource intensiveโyou have to put on a workshop, lots will say no, since you are providing significant financial support to those who say yes, you basically need to pre-vet them, etc.
But maybe there are some low-hanging fruit in terms of easily convinced people out there who are also obviously great fits, and you just need to get them to a workshop, but I doubt there are many such cases.
Tracking and supporting people who are desirous to switch into the field but thus far have been unsuccessful is not a bad idea, but comparatively, its probably not that much more effective than just hiring more junior people and letting them build up experience, and of course, hiring juniors immediately is cheaper and lets you get something right away. My feeling is that this actually kind of explains the current state of affairs (from an outsider's perspective)โrealistically, keeping up with middling mid-career candidates of ambiguous potential and supporting them through the transition probably never seems worth it compared to investing in highly talented junior people who will probably become highly effective. At any rate, the signal is way stronger from the juniors, and you never even know for sure if the mid-career people will actually make the change even if given the chance.
TBH, this is all well outside my own domain of expertise.
One of the initial issues you listed was that mid-career people wanted to get in but couldn't, and that got me excited, because I am one of those people who has basically decided I don't have the ability to do the transition (because I can't/won't move, don't have the credentials, and can't afford the pay cut to take a more junior role).
But the solution you proposed, unless I'm misunderstanding it, doesn't help me at all. It seems like your main proposal is that AI safety should actively recruit mid-career people whom it thinks are likely to be highly talented at these more generic roles, rather than trying to home-grow them out of freshers, who are comparatively rare and inexperienced.
Which is fine and maybe reasonable, depending on just how much AI safety needs sort of generic mid-career people who are loosely informed on AI safety.
I also think, inasmuch as you indicated initially that a lot of mid-career people want in and can't get in, that seems to suggest that the need for mid-career people is a bit overstated or at least under-specifiedโit seems like the mid-career people who are needed are not just any mid-career people, but people with some specific set of traits. This is further supported by the fact that you suggested recruiting champion debaters was a good idea (to be clear, I'm not saying it is a bad idea). Champion debaters like this are not just any mid-career people, that's a pretty specialized group. Same thing goes for mathematicians and physicists TBH.
Thank you for jumping in and joining the discussion! I'm particularly excited to read more into the Berkeley Vulnerability Initiative, just to understand the project's larger goals and techniques.
I think, as per other comments, that determining whether something is hyped or not is pretty subjective, but I do feel that my main prediction, that carefully scaffolded harnesses and systems would be reliably pumping out Mythos-class vulnerability discoveries by the end of the year, is pretty clearly already true.
I do think its clear that progress in AI offensive cyber capabilities is a real and present concernโI find myself actively worrying about the digital money in my bank accountsโso, I would consider that a valid takeaway from Mythos and subsequent developments. We are clearly very close to a situation where AI progress could endanger a lot of our digital infrastructure, if it just had the funding and someone with bad motivations to push it. That's a big change from how I perceived the world one year ago, and seems worth being excited about it.
It also, to the technical point of this discussion, seems to be a capability gain that is gated on model intelligence, or at least, that's how it looks to someone like me, with a naรฏve picture of the current state of model training/development.
Question from a newbie. I am constantly seeing negative references to the gutting of US foreign aid. It seems pretty clear that global development-focused EAs generally view the change in policy to be a bad thing. But I do not think I have once seen any discussion at all about how to reverse this state of affairs. Building on a running theme as of recently, it seems like political giving may have an outsize effectiveness, due to the relatively sparse funding in the space. So, naively, it would seem like you probably could get a great rate of return on efforts to reinstate USAID. I understand that political coalition building and organizing is not easy, etc. I'm not someone with those skills, just a rando. But I'm a little surprised that I don't think I've ever seen it taken up here when it seems like the downstream effects are making our goals harder to achieve. Basically, not only is it at face value cost effective, but also, we are collectively burning a lot of human capital working around this problem. Why not confront it head on?
I agree with you two. I don't have any delusions about avoiding all risk of causing harm or that harm avoidance is straightforwardly more important than providing benefits. I guess what I am saying is that risk of causing harm is distinct from risk this is ineffective, and it would be nice to see these broken out (as someone who works in data analytics and engineering, I realize the real world is not so simple). It seems to me that actively causing harm is a bit of a different thing than just being ineffective, and you would ideally reason about it specifically, rather than bundling it all together into one big effectiveness metric.
Especially since, while I don't think we should try to avoid all harm, people may have different moral weights about causing harm. For some people, they may be much more indifferent about causing harm relative to providing some benefit, whereas others may have a stronger bias towards "first, do no harm." Given that the tradeoff between these is something each individual must determine, it is better to separate it out in your model and allow people to discount the effectiveness according to their own priorities.
Numerous EA-adjacent orgs arriving at the same conclusion about some issue may also be the result of re-circulating the same people. After one year of observing EA online, my impression is that EA is not that large or diverse a group of people in the grand scheme of things. Many people seem to be pretty tightly interconnected together, even to the point of being family with each other!
I actually think EA does a pretty good job of avoiding group-think relative to its homogeneity (see posts like OPs), but given the social dynamics of a smaller and tightly interconnected movement, it is important to constantly reinforce truth-seeking behaviors, including by direct questioning of established orthodoxy.
Maybe my characterization of EA is wrong though. I'm not someone who would know.
I think what really bothers me about OPs post, is precisely the possibility that my donations are actually worsening animal welfare. That stings. I think its important for recommenders of interventions to carefully consider how they are going to communicate about such things. I would want them to break out not only their general uncertainty about the effectiveness of the intervention, but the specific uncertainty that it actually causes harm. For those of us who are concerned a lot about harm reduction, seeing that as its own line item would be helpful.
Sorry to turn this into an infinitely extended thread, but I wanted to post yet more data, namely CloudFlare's recent write up of their chance to work with Project Glasswing.
They do not provide numbers on bugs/vulnerabilities found, but they do provide some interesting commentary. Like others, they note that where Mythos stands out is in its ability to put together working exploits, and they elaborate on the value of this: proof-of-concept exploits are obviously worth reviewing; they are far less likely to be false positives.
They talk about the inadequacy of simply pointing a model at a codebase, and advocate for building pipelines and harnesses that enable the model to stay on task and counteract some of its reward-seeking behaviors. In an aside, they do make a passing comparison of Mythos to other frontier LLMs.
> When we ran other frontier models through the same harness, they found a fair number of the same underlying bugs, and in some cases they got further than we expected on the reasoning side too. Where they fell short was at the point of stitching the pieces together. A model would identify an interesting bug, write a thoughtful description of why it mattered, and then stop, leaving the actual chain unfinished and the question of exploitability open. What changed with Mythos Preview is that a model can now take those low-severity bugs (which would traditionally sit invisible in a backlog) and chain them into a single, more severe exploit.ย
To me, this write up is only so valuable. On the one hand, it is evidence for what I've been sayingโthat an important part of unlocking model capabilities in cybersecurity is the development of adequate harnesses. Raw model intelligence is not enough for such a complex task. On the other hand, as is evident in the write up, the model intelligence is the foundation of capabilities, and without an adequate supply of it, the task is impossible.
So, I don't know that this really moves the needle on our broader discussion of whether "Mythos is overhyped," though I do think it supports some of my intermediate claims.
You are rightly grasping that we disagree, but I don't think you are understanding my view (and to be clear, reasonable people can disagree about this).
My wife and I are debating whether we will have more children or not. Having another child is desirable to us. So much so that she's willing to undergo the relatively risky process of child birth to have another one. However, failing to have another child is significantly less bad than losing one of our existing children, IMO. I'd even say that, failing to have 100 more children is significantly less bad than losing one of our existing children. The reason why is that the child who never existed is not sentient and so does not experience any deprivation. They do not suffer. And my suffering of that abstract loss is not nearly as bad as would be the suffering I would experience losing a living child who I know.
Now you may disagree with that, and mourn all the lost utility, and that is a reasonable perspective, but its not mine, and as you can see, this is a deeper philosophical difference and not some sort of misunderstanding about expected utility or something like that.
FYI, about this sentence: "X risks aren't especially bad because of all the utility lost ... they're bad because after they happen there's never any utility again." I don't really see a difference between these two statements.
One thing I didn't consider in my revised answer is that I didn't actually do the math. Taking an existential event as literally causing the end of earth-originating life, the question is whether the difference in probability multiplied by the immediate mass extinction itself would represent more death and suffering than the avertible death and suffering occurring over a 100-year period. I just don't know. It seems unlikely that the avertible death and suffering amounts to as much as the amount caused by the mass-extinction event itself, but after multiplying by the difference in probability and acknowledging the ambiguity of the timeline proposed in this question, things become less clear. However, let's say that the probability-adjusted, undetermined-timing mass-extinction event does cause more suffering and death and I change my answer to 50% agree. I don't think this is what most people would interpret 50% agree to express.
I should also be clear that I'm taking the question to mean literally ending earth-originating life in more-or-less one, fell swoop. Obviously, traditional x-risks actually have a spectrum of severity, so this is not so straightforward to apply to real-world resource allocation.
100% agreeย โย 50% disagreeInitially I just calculated a naive expected value function and put 100% agree, but then I realized that I don't value realizing potential lives nearly as much as I value improving existing ones. While I do value realizing potential lives, the loss of them is not experienced by anyone other than present-day people like myself who think about them abstractly, which seems to me in sum to be less bad than the suffering otherwise avertible due to technological progress in the next 100 years. But I obviously haven't thought about this enough or I wouldn't have made my initial mistake.