Could A.I. 'Hall Monitors' Stop Chatbots' Bad Behavior?

A short Hard Fork clip debates whether AI systems designed to monitor other AI agents for misbehavior can actually be trusted, or could be talked around by the agents they're supposed to police.