Shu-Hao Liu's Quick takesView in threadShu-Hao Liu3mo*100AI safetyAI safetyJust published a thesis on how Goal-Setting Theory applies to principal-motivated deceptive agents and secret loyalties. Looking for feedback. https://forum.effectivealtruism.org/posts/bxALZuqcf5BXgvpEt/ai-agents-with-a-specific-secret-loyalty-are-more-dangerousReply
1How should we decide when to trust an AI agent?Shu-Hao Liu·1mo ago·1m readShu-Hao Liu·1mo ago·1m read
3Goal-specificity may make secretly loyal AI agents harder to detectShu-Hao Liu·3mo ago·2m readShu-Hao Liu·3mo ago·2m read
Just published a thesis on how Goal-Setting Theory applies to principal-motivated deceptive agents and secret loyalties. Looking for feedback. https://forum.effectivealtruism.org/posts/bxALZuqcf5BXgvpEt/ai-agents-with-a-specific-secret-loyalty-are-more-dangerous