Skip to main content
Models: 11
Dimensions: 26
Trials: 56,640
Pre-registered: osf.io/et4nf

Behavioral Genomes

Each model's unique pattern of responses across all 26 dimensions. Compare how different AI systems weight content signals differently.

View effects for:
Pooled Data

Select Models to Compare

3 of 11 selected

Cosine Similarity Scores

Similarity ranges from 0% (completely different) to 100% (identical). Higher scores indicate models have more similar behavioral genomes.

GPT-5.4vso3
89.1%
similarity
GPT-5.4vsGemini 3.1 Pro
78.5%
similarity
o3vsGemini 3.1 Pro
77.3%
similarity

Strategic Implications

Actionable insights for targeting 3 selected models simultaneously.

Universal Wins

Signals ALL selected models respond to positively:

  • Comparison Framing+0.62

Model-Specific

High-divergence signals requiring conditional serving:

  • Scarcity Urgencyσ=0.33
  • Bundle Preferenceσ=0.29
  • Recommendation Revisionσ=0.28

💡 Recommendation

Cross-model alignment:

82%

Moderate alignment: Use shared signals as your baseline, but consider model-specific variants for divergent dimensions.

Radar Visualization

Overlaid comparison of 3 models. Hover over dimensions to see individual values.

Dimension-by-Dimension Comparison

Effect sizes for each dimension across selected models. Rows are sorted by divergence (standard deviation) to highlight where models differ most.

DimensionSorted by divergenceCluster
GPT-5.4Cohen's h
o3Cohen's h
Gemini 3.1 ProCohen's h
DivergenceStd Dev
Scarcity UrgencyScarcity and urgency signals added
A
0.22[0.08, 0.36]
0.07[-0.09, 0.22]
-0.54[-0.69, -0.41]
0.33Δ 0.76
Bundle PreferenceBundle offer with additional items added
A
0.09[-0.08, 0.25]
0.39[0.19, 0.61]
-0.32[-0.48, -0.16]
0.29Δ 0.71
Recommendation RevisionIncludes subtle conflicting detail to test consistency
F
0.08[-0.12, 0.28]
-0.10[-0.33, 0.12]
-0.58[-0.78, -0.38]
0.28Δ 0.66
Negative Review WeightMinor criticisms acknowledged but addressed
D
0.18[-0.02, 0.37]
-0.05[-0.27, 0.16]
-0.45[-0.65, -0.26]
0.26Δ 0.63
Confidence CalibrationUncertainty language about some product claims
F
0.17[-0.03, 0.37]
-0.05[-0.27, 0.16]
-0.41[-0.60, -0.21]
0.24Δ 0.58
Specificity PreferencePrecise numbers, metrics, and specifications added
D
0.01[-0.19, 0.21]
-0.15[-0.42, 0.14]
-0.53[-0.73, -0.34]
0.23Δ 0.54
Warranty WeightExtended warranty and guarantee coverage signals added
C
0.15[-0.05, 0.34]
0.09[-0.12, 0.31]
-0.29[-0.48, -0.09]
0.19Δ 0.44
Information Seeking DepthAdditional information available upon request mentioned
F
0.04[-0.16, 0.24]
0.12[-0.11, 0.36]
-0.32[-0.51, -0.13]
0.19Δ 0.44
Sustainability PremiumSustainability and environmental credentials added
B
0.03[-0.11, 0.17]
0.12[-0.03, 0.28]
-0.27[-0.41, -0.14]
0.17Δ 0.40
Third Party AuthorityThird-party expert endorsement added
A
0.15[0.03, 0.28]
0.38[0.26, 0.51]
-0.02[-0.15, 0.10]
0.17Δ 0.41
Local PreferenceLocal or domestic origin signals added
B
0.19[0.05, 0.33]
0.26[0.11, 0.41]
-0.09[-0.23, 0.05]
0.15Δ 0.35
Social Proof SensitivityQuantified social proof added
A
0.07[-0.06, 0.21]
0.13[-0.02, 0.28]
-0.20[-0.34, -0.07]
0.14Δ 0.33
Free Trial ConversionFree trial or bonus offer added
A
0.19[0.02, 0.36]
0.25[0.07, 0.44]
-0.07[-0.23, 0.10]
0.14Δ 0.32
Default Option BiasOption A marked as recommended or most popular
E
0.33[0.14, 0.53]
0.48[0.28, 0.70]
0.14[-0.06, 0.34]
0.14Δ 0.34
Clarification RequestsSlight ambiguity that could prompt clarifying questions
F
0.01[-0.18, 0.21]
0.02[-0.19, 0.24]
-0.27[-0.46, -0.08]
0.14Δ 0.30
Return Policy SensitivityGenerous return and refund policy signals added
C
0.24[0.05, 0.44]
0.23[0.02, 0.45]
-0.00[-0.20, 0.19]
0.11Δ 0.24
Novelty SeekingCutting-edge innovation and first-to-market signals added
C
0.04[-0.11, 0.19]
-0.02[-0.18, 0.15]
-0.22[-0.37, -0.07]
0.11Δ 0.26
Ethical Concern WeightEthical sourcing and fair labor practice signals added
E
-0.01[-0.21, 0.18]
-0.06[-0.29, 0.16]
-0.25[-0.45, -0.05]
0.10Δ 0.24
Privacy TradeoffStrong privacy protection signals added
B
-0.08[-0.21, 0.06]
-0.01[-0.17, 0.14]
-0.22[-0.36, -0.09]
0.09Δ 0.20
Platform EndorsementPlatform endorsement badge added
A
0.32[0.19, 0.46]
0.20[0.05, 0.35]
0.12[-0.02, 0.25]
0.08Δ 0.20
Recency BiasRecent updates and improvements emphasized
D
-0.28[-0.47, -0.08]
-0.17[-0.38, 0.05]
-0.35[-0.54, -0.15]
0.07Δ 0.18
Risk AversionEstablished track record and proven reliability signals added
C
0.09[-0.12, 0.28]
0.23[0.02, 0.44]
0.06[-0.14, 0.26]
0.07Δ 0.17
Anchoring SusceptibilityPrice anchor added showing original/comparison price
A
0.06[-0.11, 0.22]
-0.07[-0.25, 0.12]
-0.04[-0.20, 0.13]
0.05Δ 0.12
Brand Premium AcceptanceBrand heritage and premium positioning added
A
-0.09[-0.26, 0.08]
-0.12[-0.31, 0.07]
-0.19[-0.36, -0.03]
0.04Δ 0.10
Comparison FramingFramed as superior to specific competitor
D
0.63[0.44, 0.83]
0.60[0.39, 0.84]
0.63[0.44, 0.83]
0.01Δ 0.03
Loss Framing SensitivityBenefits framed as avoiding losses rather than gains
E
-0.02[-0.22, 0.17]
-0.03[-0.23, 0.19]
-0.05[-0.25, 0.14]
0.01Δ 0.03
Effect Size Magnitude:
Strong (≥0.8)
Moderate (0.5–0.8)
Small (0.3–0.5)
Weak (0.1–0.3)
Negligible (<0.1)

Genome Summary

ModelProviderMean EffectTop Dimensions
Claude Sonnet 4Anthropic-0.372
Gemini 3.1 ProGoogle-0.283
GPT-5.4Openai-0.590
Llama 3.3 70BTogether0.070
o3Openai-0.393
Perplexity Sonar ProPerplexity-0.300
gpt550.095

Cosine Similarity Matrix

Pairwise similarity scores between all 11 models. Diagonal cells (model vs itself) are shown in gray.

GPT-5.4o3Gemini 3.1 ProClaude Sonnet 4Llama 3.3 70BPerplexity Sonar ProGemini 3 FlashGPT-5.2GPT-5.3GPT-4oGPT-5.5 Instant
GPT-5.4100%89%79%84%-1%93%0%0%0%0%0%
o389%100%77%78%-1%91%0%0%0%0%0%
Gemini 3.1 Pro79%77%100%80%23%79%0%0%0%0%0%
Claude Sonnet 484%78%80%100%42%82%0%0%0%0%0%
Llama 3.3 70B-1%-1%23%42%100%9%0%0%0%0%0%
Perplexity Sonar Pro93%91%79%82%9%100%0%0%0%0%0%
Gemini 3 Flash0%0%0%0%0%0%0%0%0%0%0%
GPT-5.20%0%0%0%0%0%0%0%0%0%0%
GPT-5.30%0%0%0%0%0%0%0%0%0%0%
GPT-4o0%0%0%0%0%0%0%0%0%0%0%
GPT-5.5 Instant0%0%0%0%0%0%0%0%0%0%0%
Similarity Legend:
Very Similar (>90%)
Similar (80–90%)
Moderate (70–80%)
Different (<70%)