Understanding how LLMs learn (e.g. SLT) is underrated relative to work on instilling specific behaviors
By analogy, this is like the difference between understanding how to make humans not suffer from the foibles of human nature, rather than the very brittle and incomplete methods of culture, religion, and ideology. Make something not even have evil nature first!
Theories of consciousness will lead to actionable understanding of AI consciousness²
Yes, though I think the direction will far more go the opposite way - that is, AI conciousness, if it emerges, will still more help us understand consciousness.
Recursive thinking loops possibly put current systems into a state of transient quasi-awareness; once this exists, the capacity for suffering becomes inevitable in any system with incentives.
Benchmarks will become useless due to eval awareness¹
Though they will need much more careful design, done along with careful understanding of incentives alignment, techniques such as competitive benchmarks will continue to have some utility.
If animals continue to exist in a post-AGI world, animal suffering will not persist
If animals persist, it is likely we would like them to remain to some degree of their nature; prosperity and control will likely allow for substantial reduction of suffering but not elimination.
By analogy, this is like the difference between understanding how to make humans not suffer from the foibles of human nature, rather than the very brittle and incomplete methods of culture, religion, and ideology. Make something not even have evil nature first!
I don't really see it much in the discourse at all; if it isn't by this stage, it's going to be hard to catch up.
Humans, at least, tend to learn empathy and have it encouraged and reinforced by example. Humans will mirror AI's; AI's will learn from each other.
Yes, though I think the direction will far more go the opposite way - that is, AI conciousness, if it emerges, will still more help us understand consciousness.
Fairly uncertain here, though I don't see any way of continuing to make systems better without trying for value alignment. A dangerous situation.
Recursive thinking loops possibly put current systems into a state of transient quasi-awareness; once this exists, the capacity for suffering becomes inevitable in any system with incentives.
Though they will need much more careful design, done along with careful understanding of incentives alignment, techniques such as competitive benchmarks will continue to have some utility.
If animals persist, it is likely we would like them to remain to some degree of their nature; prosperity and control will likely allow for substantial reduction of suffering but not elimination.