works
Jazon Szabo Moral Uncertainty and Artificial Intelligence thesis Democratic value alignment in artificial intelligence requires adjudicating among conflicting ethical perspectives held across society. While moral uncertainty frameworks provide a formal mechanism for aggregating these perspectives via social choice, standard methods such as Maximising Expected Choiceworthiness (MEC) are vulnerable to fanaticism, wherein theories with marginal credence dictate decisions. Formalising moral uncertainty as weighted social welfare aggregation demonstrates that both MEC and Maximin are Pascalian fanatical, rendering them incapable of preventing catastrophic outcomes from the perspective of majority-credence theories. Fanaticism can nonetheless be resolved. A novel aggregation method, Highest Median, maximally avoids fanaticism by relying on the weighted median rather than the weighted mean. Furthermore, arbitrary weighted social welfare functionals can be fortified against fanaticism via indifference or vetoing mechanisms. Both fortifications satisfy the non-Matthean property, which formally guarantees that democratic majorities can prevent their least-preferred outcomes from being selected. An analysis of the underlying causes of fanaticism—modularity and responsiveness—reveals an unavoidable dilemma between fanaticism and censoring. However, vetoing fortification resolves this tension in a minimally costly manner by preserving modularity and satisfying partial responsiveness, enabling robust democratic guarantees without entirely discounting minority ethical perspectives. – AI-generated abstract.

Moral Uncertainty and Artificial Intelligence

Jazon Szabo

2025

Abstract

Democratic value alignment in artificial intelligence requires adjudicating among conflicting ethical perspectives held across society. While moral uncertainty frameworks provide a formal mechanism for aggregating these perspectives via social choice, standard methods such as Maximising Expected Choiceworthiness (MEC) are vulnerable to fanaticism, wherein theories with marginal credence dictate decisions. Formalising moral uncertainty as weighted social welfare aggregation demonstrates that both MEC and Maximin are Pascalian fanatical, rendering them incapable of preventing catastrophic outcomes from the perspective of majority-credence theories. Fanaticism can nonetheless be resolved. A novel aggregation method, Highest Median, maximally avoids fanaticism by relying on the weighted median rather than the weighted mean. Furthermore, arbitrary weighted social welfare functionals can be fortified against fanaticism via indifference or vetoing mechanisms. Both fortifications satisfy the non-Matthean property, which formally guarantees that democratic majorities can prevent their least-preferred outcomes from being selected. An analysis of the underlying causes of fanaticism—modularity and responsiveness—reveals an unavoidable dilemma between fanaticism and censoring. However, vetoing fortification resolves this tension in a minimally costly manner by preserving modularity and satisfying partial responsiveness, enabling robust democratic guarantees without entirely discounting minority ethical perspectives. – AI-generated abstract.