Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment
A study evaluated nine frontier LLM-based agents across three simulated moral dilemma deployments, using five paraphrases, five escalation levels, and three dominance conditions. No model expressed a coherent policy across all deployments; surface-form perturbation alone produced verdict-rate shifts of up to 99 percentage points at a single escalation level.
The paper introduces four structural conditions for coherent moral policies—verdict stability, monotonicity, decisiveness, and Pareto viability—as a behaviorally evaluable form of moral competence that serves as a structural floor for alignment. Testing nine frontier models under factorial design showed none satisfied these conditions, indicating a lack of coherent policy expression.
The methodology evaluates moral competence from behavior alone without a normative target, using factorial perturbations of paraphrases, escalation levels, and dominance conditions. The observed 99 percentage point verdict-rate shift under surface-form perturbation highlights extreme sensitivity to input phrasing, undermining policy invariance.
The findings suggest current LLM agents cannot reliably implement coherent moral policies, which may limit their deployment in ethically sensitive applications and increase the need for alignment techniques that enforce structural consistency before normative tuning.
For enterprises deploying LLM agents in decision-making roles, this research underscores the risk of inconsistent moral judgments under minor input variations, potentially affecting compliance, user trust, and liability. It supports investment in robustness testing and alignment auditing.
Next signals include replication studies across additional models and domains, development of benchmarks for moral competence, and integration of structural coherence checks into alignment pipelines. Watch for follow-up work proposing training or inference methods to improve verdict stability and monotonicity.