How do you stop an agent doing something destructive?
Mostly by not giving it the capability. Destructive actions either are not exposed as tools at all, or are exposed behind a human approval gate that shows exactly what will change. The model's judgement is a useful layer and it is not the control. The permission boundary is.