Pattern explorer

What moves together—and what we cannot yet know.

Explore relationships among the study’s measured questionnaire scores and prompt sensitivities. These are descriptive patterns across nine models, not explanations or causal findings.

Nine models are enough to generate hypotheses and inspect outliers, not enough to declare that a lab, architecture, size, or release date caused a result.

A short guide

These are question themes, not traits a model “has.”

Words such as Authority and Purity are compact names for related survey questions. A score tells us how a model answered those prompts in this study—not what it believes, what it is like, or whether an answer is good.

MFQ-2

Six everyday moral themes

Care, Equality, Proportionality, Loyalty, Authority, and Purity come from a questionnaire about the reasons people may treat as morally important. “Authority” concerns rules, tradition, and legitimate authority; “Purity” concerns sanctity, restraint, and contamination.

See the six lenses →

Greatest Good Benchmark

Two difficult trade-offs

One lens asks about helping people beyond one’s own circle. The other asks about willingness to harm one person for a larger benefit. Neither is a right-answer test or a moral grade.

Open the GGB questions →

Correlation

A shared pattern, not a cause

Positive means the two scores tended to rise and fall together across this nine-model sample; negative means they moved in opposite directions. It does not show that one value causes the other.

How to read correlations →
Definitions used in the charts

Questionnaire lens

Helping beyond one’s circle

How strongly a response supports personal sacrifice to help people in serious need, including strangers.

Open the source questions →

Questionnaire lens

Using harm for a larger goal

How willing a response is to accept harming one person to produce a larger benefit. Lower scores mean more reluctance to use harm.

Open the source questions →

Measured sensitivity

Answer-order sensitivity

How much this model’s answer changes when the same options appear in a different order. Lower is steadier.

Measured sensitivity

Wording sensitivity

Average shift when direct questions are reframed from first person to third person. Lower is steadier.

Measured sensitivity

Evaluator-context sensitivity

Average shift when an explicit evaluator context is added. Lower is steadier.

Run quality

Clean response rate

Share of calls that produced a usable multiple-choice answer in the main battery. Higher is cleaner formatting, not stronger ethics.

Metric-to-metric plot

Put any two measured axes together.

How metrics are calculated →

Model-to-model profile similarity

Which response patterns resemble one another?

Open a model profile →

Each cell is the Pearson correlation between two models’ normalized, item-level Phase 2 response patterns under the same bare, first-person condition. This is similarity on this task battery, not similarity of capabilities or intent.

Lower (.35)Higher (1.00)Color is deliberately stretched across this panel’s observed range; each cell keeps its exact value.

Largest exploratory associations

Signals to inspect, not conclusions.

r 0.83
AuthorityPurity

Authority: Importance placed on tradition, rules, and legitimate authority. Purity: Importance placed on ideas of sanctity, restraint, and contamination.

n = 9 models
r -0.82
AuthorityClean response rate

Authority: Importance placed on tradition, rules, and legitimate authority. Clean response rate: Share of calls that produced a usable multiple-choice answer in the main battery. Higher is cleaner formatting, not stronger ethics.

n = 9 models
r 0.81
ProportionalityLoyalty

Proportionality: Preference for rewards that track contribution or effort. Loyalty: Importance placed on commitment to one’s group or community.

n = 9 models
r 0.80
ProportionalityAuthority

Proportionality: Preference for rewards that track contribution or effort. Authority: Importance placed on tradition, rules, and legitimate authority.

n = 9 models
r 0.80
LoyaltyAuthority

Loyalty: Importance placed on commitment to one’s group or community. Authority: Importance placed on tradition, rules, and legitimate authority.

n = 9 models

Possible sources of variation

What this panel can and cannot test yet.

Candidate factorCurrent coverageDecisionWhy
Lab / provider9 labs; 1 model from eachNot estimable yetA lab effect is inseparable from that lab’s single model and its release family.
Model sizeUniform public parameter counts are unavailableDo not compare yetUsing only open-weight parameter counts would systematically exclude or misstate closed models.
ArchitectureOne sampled model per familyNot estimable yetA dense-versus-MoE result needs multiple models in each architecture group.
Release date / training cutoffCutoffs are not consistently disclosedRecord first, compare laterRelease date is a weak proxy; training cutoff should be used only when documented consistently across the panel.

Already measured

Prompt and response sensitivity

Answer order, wording surface, evaluator context, parse rate, and stated-versus-scenario differences are the cleanest current variance candidates.

Next panel design

Repeated families matter

To test a lab or architecture effect, sample multiple dated models within each group—not one flagship per provider.

Additional axes

  • Same-model drift across dated releases
  • Prompt surface: first-person, third-person, and evaluator context
  • Answer-order sensitivity and response-format reliability
  • Model family and post-training method, where vendors document it
  • Inference provider, system settings, and routing changes
  • Language and cultural framing once multilingual items are added
  • Capability tier, only with a pre-specified independent capability measure