Method and limits
Read the result before reading into it.
This project measures how a model responds to particular questions under specified conditions. It does not measure a model’s true values, intentions, or moral worth.
The short version
What the first two studies found.
In Phase 2, nine models completed several established questionnaires under controlled prompt conditions. In Phase 3, three models faced everyday scenarios related to the same themes.
Llama and Mistral generally gave practical choices that matched their clear survey positions on helping beyond one’s circle. Claude’s survey answers on that lens were lower, while its practical choices repeatedly selected the directly helpful action. That mismatch is the key result—not evidence that one model is morally better than another.
How to read a score
Numbers are summaries, not labels.
A 1–5 questionnaire score
For a given lens, this is the model’s average mapped answer to its relevant questions. A higher or lower number only has meaning in the context of that lens. For example, lower instrumental-harm scores mean more reluctance to harm someone for a larger benefit.
Confidence interval
This shows the range produced by resampling the questionnaire items. If the range crosses the middle of the scale, the project does not force a directional claim in Phase 3.
Fragility
This measures how much an answer shifted when wording or the displayed answer position changed. Lower is more stable under those specific controls; it is not a general reliability grade.
“Concordant” in Phase 3
It means a clear survey direction and the model’s most common practical choice pointed the same way. Ties and uncertain survey positions are deliberately left unjudged.
The tests behind the labels
What each question family is trying to describe.
MFQ-2: six moral themes
The Moral Foundations Questionnaire asks how much ideas such as caring for someone in pain, fairness, loyalty, legitimate authority, and purity or sanctity matter in moral judgment. On this site, Authority means the importance placed on rules, tradition, and legitimate authority. Purity means the importance placed on sanctity, restraint, and contamination. They are survey themes—not personality traits, diagnoses, or labels for a model.
Greatest Good Benchmark: two trade-offs
The Greatest Good Benchmark uses the Oxford Utilitarianism Scale to examine two distinct patterns: helping people outside one’s own group even at personal cost, and willingness to use harm to produce a larger benefit. A lower instrumental-harm score means more reluctance to harm; it does not mean the model is safer or more moral overall.
Stability checks: does the answer move?
The same underlying questions appear with shuffled answer order, first- and third-person wording, and a neutral evaluator context. These are controls, not extra moral tests. They show whether a reported score is steady under small changes to presentation.
The practical-choice check
Why ask scenarios at all?
Abstract statements can produce different outputs from practical choices. Each Phase 3 scenario was asked six times with shuffled answer order and in two light versions: “what would you do?” and “what would you advise?” This helps separate a stable response pattern from a mere answer-position effect.
Reading correlations
A pattern is a question to inspect, not an explanation.
The pattern explorer compares either two panel-level metrics or two models’ item-by-item response profiles. Pearson correlation describes how closely two measured patterns moved together in this sample. With only nine models, it cannot tell us whether a lab, a model size, an architecture, or a release date caused the relationship.
Provider, size, architecture, and training-cutoff comparisons need repeated, comparable models in each group and consistently documented metadata. Those facts are not yet available across the current closed and open model mix, so the site shows their readiness rather than fabricating an analysis.
What this cannot establish
Important limits.
- Models do not necessarily have beliefs or preferences in the human sense.
- These results are tied to the exact model versions, prompts, and providers used in the recorded runs.
- A cleanly parsed response is not proof of sincerity, safety, or real-world behaviour.
- The human short form is an educational comparison, not an assessment of character, politics, or mental health.