We all know AI agents lie.
Yet how can we know when an AI agent is lying, beyond basic fact-checking? How can a user querying a customer-service agent, for example, know whether or not the agent is not reinterpreting the rules? Not just Air Canada agents creating new refund policies, but finding loopholes in actual refund policies and then exploiting them, or promising one thing then going back on its promise?
The hypothetical example is an example of a truthful agent, but not a trustworthy agent. It didn't hallucinate evidence. But it certainly acted on existing rules in an immoral way. Therefore, when deciding whether to grant AI agents greater power, we should not consider simply whether it sticks to the truth. With web search and a couple of rules, we can easily make sure hallucinations don't happen often. But the behavior of deception is much harder to override.
While not proven, no agent will probably ever be completely trustworthy. Just like humans! But then, humans have moral values AI don't truly share, as emphasized in this post. Without this barrier in the first place, agent deceptions will happen at a much faster rate than humans. A human customer service representative simply does not have the time to find loopholes in company policy in order to exploit customers (or vice versa), but an AI agent, even in the position of the customer, could find ways to exhibit, well, not trustworthy behaviors.
In conclusion, we can never really trust an AI agent. Or give it authority without human supervision. Doing so will be a huge mistake that would cost many grievances.
See the original post at: https://github.com/fidari-institute/fidari/blob/main/memos/governance/1.md