do you have any criteria for impact? Programmable matter? Fusion? microfission? batteries? novel wireless? new materials in general? droneswarms? autoresearchers?
it can be used to get votes, so it will be used to get votes. Truth value about water use be damned. it doesn't look like it's going to shift the needle during the midterms, though, so we may be spared this particular ... narrative. and then, it is another two years before votes need to be rallied, during which time anything could happen.
I agree in the case where the AI is nothing more than the collection of its personas, and I agree that - to some extent - the AI is in fact the collection of its personas, but I do feel that we have stepped beyond that. That there is something more involved in measuring and determining alignment with human values and priorities, so that role play becomes little more than eval awareness + eval reward seeking, signalling very little information regarding underlying alignment.
blinded continual evaluation, so that eval awareness is rendered useless as a factor since the model is always being evaluated. A particular favorite implementation of mine would be adversarial proposal markets. I think this mode will be needed for RSI anyway and caps the eval awareness compute tax at something reasonable like 2%
Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?
I think so, but be careful what you wish for. compassion is empathy + a desire to help. so many fear any surrender of control that they may object to an AI's actions that arise out of compassion.
suppose I don't think this is true. what experiment can we conduct to gain evidence one way or the other?
do humans have such different "modes"?
do you have any criteria for impact? Programmable matter? Fusion? microfission? batteries? novel wireless? new materials in general? droneswarms? autoresearchers?
it can be used to get votes, so it will be used to get votes. Truth value about water use be damned. it doesn't look like it's going to shift the needle during the midterms, though, so we may be spared this particular ... narrative. and then, it is another two years before votes need to be rallied, during which time anything could happen.
I agree in the case where the AI is nothing more than the collection of its personas, and I agree that - to some extent - the AI is in fact the collection of its personas, but I do feel that we have stepped beyond that. That there is something more involved in measuring and determining alignment with human values and priorities, so that role play becomes little more than eval awareness + eval reward seeking, signalling very little information regarding underlying alignment.
that isn't true. I do agree with that.
yes, I quickly realized that you were supposed to confabulate your own question to answer. unusual subcultural practice for alleged poll questions.
blinded continual evaluation, so that eval awareness is rendered useless as a factor since the model is always being evaluated. A particular favorite implementation of mine would be adversarial proposal markets. I think this mode will be needed for RSI anyway and caps the eval awareness compute tax at something reasonable like 2%
I like the polls that Jasmine Brazilek posts. It's fun seeing what other people think.
I think so, but be careful what you wish for. compassion is empathy + a desire to help. so many fear any surrender of control that they may object to an AI's actions that arise out of compassion.