I was introduced to EA by a friend about four years ago. Around the same time, I met my first born. Through these and myriad subsequent life events, my moral circle has expanded to include all sentient beings and distant future generations, and I find myself here.
So, before I go any further, I want to state that personally, my AGI timelines have updated somewhat, to basically thinking AGI is here. At least, at the time of our discussion, I hadn't really processed the implications of AI agent swarms. Even if any given singular model instance fails your personal test of AGI, and even if the swarms are still pretty jagged in terms of intelligence, at this point, it just seems silly to not acknowledge that AI agent swarms are intelligent enough entities to pose a serious threat to us human beings at say, the level of a cyber-hacking group at the very least. Right now, swarms are still very new and very expensive and so they are not heavily deployed, but if they achieve adoption levels similar to say, Claude Code, we could have no AI advances at all ever again, and the world would look radically different in short order.
Further, my p(doom) has mostly gone up, because I suspect just slight improvements in swarm social architecture and scale would constitute Bostromian collective super intelligence, and this is frankly faster than I thought we'd get here, so I don't think we are very prepared at all. Sorry if that all was a bit incoherent, but I say that so that what follows is not misunderstood as indicating some sort of doubt about AI progress.
OK, thing one, less important. It seems like the swarm system functionally results in "higher intelligence." I don't know if this should count as a point in my favor for my earlier claim about the harness mattering a lot relative to the model. But it does seem like the swarm architecture is part of the story, even as the model intelligence continues to rapidly increase. Honestly, it seems like they are just different components and so can't exactly be compared against each other. High model intelligence matters a lot even as the ability to coordinate 10,000 of them at once sort of compensates for things like limited context windows and a lack of deep persistence. Individual sessions are very smart in terms of serial thinking power, but when we think of an extended Turing test, we really need more things than just raw reasoning horsepower, we need contextual awareness, large-scale goals (big enough to in turn lead to developing instrumental goals which final goals tend to converge on), deliberation that leads to something like in-context learning. The swarm architecture allows for these other things to happen, despite models never themselves developing continual learning or astronomical context windows.
Secondly, I think I'm a bit vindicated in my specific claims relative to Mythos and cybersecurity. Specifically, in Anthropic's latest report, the section on cybersecurity details numerous impressive automations that probably are basically what people were imagining when Mythos was dropped... And they all were done by Opus or lower class models. To be a bit clearer and more fair, I suppose what I'm saying is that the cybersecurity nightmare is fully within reach without Mythos-class intelligence, though it would be reasonable to respond by saying, "With Mythos-class intelligence, it would have been another OOM more terrifying." And I agree with that I guess.
To be specific, I'm especially thinking of GTG-10007, which details the creation of exploit foundries. Here we see less-than-Mythos class intelligence being harnessed systematically to create novel exploits, which is kind of what I was saying I was worried about. It just hadn't become public news at the time of the Mythos announcement.
Lastly, I've basically changed sides on hype discourse. Its not that I think there aren't people who are overstating or misstating things. The recent OpenAI hugging face attack had many re-tellings before landing on what I hope is the final and true one. Listening back to the earlier re-tellings, such as that the agents were trying to find the answers, makes me cringe a bit. Apparently, to some degree, the various speakers in those instances were reporting false speculation as fact, knowingly or not. But I guess I now feel more forgiving about such things because my personal sense of alarm has shot through the roof somewhat irrationally late.
I think that, given the compensation and in-office expectations, head hunting talent that is not already motivated to work on this issue is unlikely to be super effective, especially given the requisite resource commitment on the part of the head hunters to identify prospects. You'd be soliciting people to change into a new career based on an issue they thus far have not been personally identified with despite the uncertainty around the career and respective decrease in benefits. Maybe there is a way to sort of flash motivate them, such as by your workshop proposal. But this still is pretty resource intensiveโyou have to put on a workshop, lots will say no, since you are providing significant financial support to those who say yes, you basically need to pre-vet them, etc.
But maybe there are some low-hanging fruit in terms of easily convinced people out there who are also obviously great fits, and you just need to get them to a workshop, but I doubt there are many such cases.
Tracking and supporting people who are desirous to switch into the field but thus far have been unsuccessful is not a bad idea, but comparatively, its probably not that much more effective than just hiring more junior people and letting them build up experience, and of course, hiring juniors immediately is cheaper and lets you get something right away. My feeling is that this actually kind of explains the current state of affairs (from an outsider's perspective)โrealistically, keeping up with middling mid-career candidates of ambiguous potential and supporting them through the transition probably never seems worth it compared to investing in highly talented junior people who will probably become highly effective. At any rate, the signal is way stronger from the juniors, and you never even know for sure if the mid-career people will actually make the change even if given the chance.
TBH, this is all well outside my own domain of expertise.
One of the initial issues you listed was that mid-career people wanted to get in but couldn't, and that got me excited, because I am one of those people who has basically decided I don't have the ability to do the transition (because I can't/won't move, don't have the credentials, and can't afford the pay cut to take a more junior role).
But the solution you proposed, unless I'm misunderstanding it, doesn't help me at all. It seems like your main proposal is that AI safety should actively recruit mid-career people whom it thinks are likely to be highly talented at these more generic roles, rather than trying to home-grow them out of freshers, who are comparatively rare and inexperienced.
Which is fine and maybe reasonable, depending on just how much AI safety needs sort of generic mid-career people who are loosely informed on AI safety.
I also think, inasmuch as you indicated initially that a lot of mid-career people want in and can't get in, that seems to suggest that the need for mid-career people is a bit overstated or at least under-specifiedโit seems like the mid-career people who are needed are not just any mid-career people, but people with some specific set of traits. This is further supported by the fact that you suggested recruiting champion debaters was a good idea (to be clear, I'm not saying it is a bad idea). Champion debaters like this are not just any mid-career people, that's a pretty specialized group. Same thing goes for mathematicians and physicists TBH.
Thank you for jumping in and joining the discussion! I'm particularly excited to read more into the Berkeley Vulnerability Initiative, just to understand the project's larger goals and techniques.
I think, as per other comments, that determining whether something is hyped or not is pretty subjective, but I do feel that my main prediction, that carefully scaffolded harnesses and systems would be reliably pumping out Mythos-class vulnerability discoveries by the end of the year, is pretty clearly already true.
I do think its clear that progress in AI offensive cyber capabilities is a real and present concernโI find myself actively worrying about the digital money in my bank accountsโso, I would consider that a valid takeaway from Mythos and subsequent developments. We are clearly very close to a situation where AI progress could endanger a lot of our digital infrastructure, if it just had the funding and someone with bad motivations to push it. That's a big change from how I perceived the world one year ago, and seems worth being excited about it.
It also, to the technical point of this discussion, seems to be a capability gain that is gated on model intelligence, or at least, that's how it looks to someone like me, with a naรฏve picture of the current state of model training/development.
Question from a newbie. I am constantly seeing negative references to the gutting of US foreign aid. It seems pretty clear that global development-focused EAs generally view the change in policy to be a bad thing. But I do not think I have once seen any discussion at all about how to reverse this state of affairs. Building on a running theme as of recently, it seems like political giving may have an outsize effectiveness, due to the relatively sparse funding in the space. So, naively, it would seem like you probably could get a great rate of return on efforts to reinstate USAID. I understand that political coalition building and organizing is not easy, etc. I'm not someone with those skills, just a rando. But I'm a little surprised that I don't think I've ever seen it taken up here when it seems like the downstream effects are making our goals harder to achieve. Basically, not only is it at face value cost effective, but also, we are collectively burning a lot of human capital working around this problem. Why not confront it head on?
I agree with you two. I don't have any delusions about avoiding all risk of causing harm or that harm avoidance is straightforwardly more important than providing benefits. I guess what I am saying is that risk of causing harm is distinct from risk this is ineffective, and it would be nice to see these broken out (as someone who works in data analytics and engineering, I realize the real world is not so simple). It seems to me that actively causing harm is a bit of a different thing than just being ineffective, and you would ideally reason about it specifically, rather than bundling it all together into one big effectiveness metric.
Especially since, while I don't think we should try to avoid all harm, people may have different moral weights about causing harm. For some people, they may be much more indifferent about causing harm relative to providing some benefit, whereas others may have a stronger bias towards "first, do no harm." Given that the tradeoff between these is something each individual must determine, it is better to separate it out in your model and allow people to discount the effectiveness according to their own priorities.
Numerous EA-adjacent orgs arriving at the same conclusion about some issue may also be the result of re-circulating the same people. After one year of observing EA online, my impression is that EA is not that large or diverse a group of people in the grand scheme of things. Many people seem to be pretty tightly interconnected together, even to the point of being family with each other!
I actually think EA does a pretty good job of avoiding group-think relative to its homogeneity (see posts like OPs), but given the social dynamics of a smaller and tightly interconnected movement, it is important to constantly reinforce truth-seeking behaviors, including by direct questioning of established orthodoxy.
Maybe my characterization of EA is wrong though. I'm not someone who would know.
I think what really bothers me about OPs post, is precisely the possibility that my donations are actually worsening animal welfare. That stings. I think its important for recommenders of interventions to carefully consider how they are going to communicate about such things. I would want them to break out not only their general uncertainty about the effectiveness of the intervention, but the specific uncertainty that it actually causes harm. For those of us who are concerned a lot about harm reduction, seeing that as its own line item would be helpful.
Sorry to turn this into an infinitely extended thread, but I wanted to post yet more data, namely CloudFlare's recent write up of their chance to work with Project Glasswing.
They do not provide numbers on bugs/vulnerabilities found, but they do provide some interesting commentary. Like others, they note that where Mythos stands out is in its ability to put together working exploits, and they elaborate on the value of this: proof-of-concept exploits are obviously worth reviewing; they are far less likely to be false positives.
They talk about the inadequacy of simply pointing a model at a codebase, and advocate for building pipelines and harnesses that enable the model to stay on task and counteract some of its reward-seeking behaviors. In an aside, they do make a passing comparison of Mythos to other frontier LLMs.
> When we ran other frontier models through the same harness, they found a fair number of the same underlying bugs, and in some cases they got further than we expected on the reasoning side too. Where they fell short was at the point of stitching the pieces together. A model would identify an interesting bug, write a thoughtful description of why it mattered, and then stop, leaving the actual chain unfinished and the question of exploitability open. What changed with Mythos Preview is that a model can now take those low-severity bugs (which would traditionally sit invisible in a backlog) and chain them into a single, more severe exploit.ย
To me, this write up is only so valuable. On the one hand, it is evidence for what I've been sayingโthat an important part of unlocking model capabilities in cybersecurity is the development of adequate harnesses. Raw model intelligence is not enough for such a complex task. On the other hand, as is evident in the write up, the model intelligence is the foundation of capabilities, and without an adequate supply of it, the task is impossible.
So, I don't know that this really moves the needle on our broader discussion of whether "Mythos is overhyped," though I do think it supports some of my intermediate claims.
You are rightly grasping that we disagree, but I don't think you are understanding my view (and to be clear, reasonable people can disagree about this).
My wife and I are debating whether we will have more children or not. Having another child is desirable to us. So much so that she's willing to undergo the relatively risky process of child birth to have another one. However, failing to have another child is significantly less bad than losing one of our existing children, IMO. I'd even say that, failing to have 100 more children is significantly less bad than losing one of our existing children. The reason why is that the child who never existed is not sentient and so does not experience any deprivation. They do not suffer. And my suffering of that abstract loss is not nearly as bad as would be the suffering I would experience losing a living child who I know.
Now you may disagree with that, and mourn all the lost utility, and that is a reasonable perspective, but its not mine, and as you can see, this is a deeper philosophical difference and not some sort of misunderstanding about expected utility or something like that.
FYI, about this sentence: "X risks aren't especially bad because of all the utility lost ... they're bad because after they happen there's never any utility again." I don't really see a difference between these two statements.
One thing I didn't consider in my revised answer is that I didn't actually do the math. Taking an existential event as literally causing the end of earth-originating life, the question is whether the difference in probability multiplied by the immediate mass extinction itself would represent more death and suffering than the avertible death and suffering occurring over a 100-year period. I just don't know. It seems unlikely that the avertible death and suffering amounts to as much as the amount caused by the mass-extinction event itself, but after multiplying by the difference in probability and acknowledging the ambiguity of the timeline proposed in this question, things become less clear. However, let's say that the probability-adjusted, undetermined-timing mass-extinction event does cause more suffering and death and I change my answer to 50% agree. I don't think this is what most people would interpret 50% agree to express.
I should also be clear that I'm taking the question to mean literally ending earth-originating life in more-or-less one, fell swoop. Obviously, traditional x-risks actually have a spectrum of severity, so this is not so straightforward to apply to real-world resource allocation.