Maybe it'd be easier to train the models on key parts of the legal code and give it a hard aversion to breaking the law - rather than training on vague value judgements and then hope the model doesn't break the law
How does not explicitly trying to train the model specifically not to break the law give them any wiggle room if the model then goes and breaks the law?
Deniability on intention. Models can hallucinate, all bets are off. Sure, they can make it stricter but why do that and subject themselves to harder scrutiny based on the training criteria?
Maybe it'd be easier to train the models on key parts of the legal code and give it a hard aversion to breaking the law - rather than training on vague value judgements and then hope the model doesn't break the law