We already have a small system for comparing answers across models, and I agree that falsification could strengthen it by asking where a moral might fail, be misapplied, or produce unintended harm. These questions are already partly included in the Moral Compare tab shown in the attached image.

However, I think moral interpretation presents several deeper problems:

1) The composition, omissions and weighting of model training data are not fully visible, whether we are examining an internal model or an external one. This may affect interpretation in many ways; for example, a model developed within one society may interpret a story differently when it is applied within another, particularly across Eastern and Western cultural traditions.

2) A model’s moral interpretation is not equivalent to a person’s moral judgement. Local, cultural and minority perspectives may be underrepresented or averaged away.

3) Model outputs are non-deterministic generated results. Unlike an individual, whose views may have developed through relatively stable life experiences, a model can produce different—and sometimes opposing—interpretations from similar prompts.

4) The context in which a story is read may also shape the reader’s interpretation. Age, current social dynamics, personal experience and what someone recognises in the situation at a particular point in their life can all affect the moral they take from it.

We can therefore compare models, test assumptions, search for counterexamples, examine cultural context and ask where each moral might be misapplied.

But falsification does not necessarily reveal one final “correct” moral. It can help us identify interpretations that are weaker, narrower or potentially more harmful, while still leaving room for legitimate differences in human perspective.

That is really the purpose of the system: not to make the machine the moral authority, but to make its assumptions visible enough for people, particularly children and parents, to examine and discuss them.