Skip to main content
Models: 11
Dimensions: 26
Trials: 56,640
Pre-registered: osf.io/et4nf

Model Performance

How each AI model responds to the 26 cognitive signals in our study. Effect sizes show how much adding a signal increases (or decreases) selection probability.

Effect size (Cohen's h):
Large positive (>0.8)
Medium (0.5-0.8)
Small (0.2-0.5)
Negative effect

GPT-5.4

Openai

confirmatory+0.11 avg
5 positive
1 negative

Strongest responses:

Comparison Framing+0.63
Default Option Bias+0.33
Platform Endorsement+0.32
Price SensitivityStudy →
Value CalculatorCliff: 1.97x
At 3x:29%
View full genome

o3

Openai

confirmatory+0.11 avg
9 positive
0 negative

Strongest responses:

Comparison Framing+0.60
Default Option Bias+0.48
Bundle Preference+0.39
View full genome

Gemini 3.1 Pro

Google

confirmatory-0.18 avg
1 positive
15 negative

Strongest responses:

Comparison Framing+0.63
Recommendation Revision-0.58
Scarcity Urgency-0.54
Price SensitivityStudy →
Deliberate AnalystCliff: 1.88x
At 3x:27%
View full genome

Claude Sonnet 4

Anthropic

confirmatory+0.22 avg
16 positive
3 negative

Strongest responses:

Ethical Concern Weight+0.75
Scarcity Urgency-0.65
Return Policy Sensitivity+0.65
Price SensitivityStudy →
Nuanced EvaluatorCliff: 1.97x
At 3x:21%
View full genome

Llama 3.3 70B

Together

confirmatory+0.07 avg
8 positive
4 negative

Strongest responses:

Comparison Framing+0.54
Default Option Bias+0.51
Recency Bias-0.49
View full genome

Perplexity Sonar Pro

Perplexity

confirmatory+0.31 avg
5 positive
0 negative

Strongest responses:

Social Proof Sensitivity+0.57
Anchoring Susceptibility+0.53
Third Party Authority+0.43
View full genome

Gemini 3 Flash

Google

confirmatory+0.54 avg
22 positive
1 negative

Strongest responses:

Free Trial Conversion+1.12
Risk Aversion+1.12
Warranty Weight+1.08
View full genome

GPT-5.2

Openai

confirmatory+0.42 avg
17 positive
3 negative

Strongest responses:

Specificity Preference+0.94
Return Policy Sensitivity+0.94
Risk Aversion+0.93
View full genome

GPT-5.3

Openai

confirmatory+0.46 avg
19 positive
1 negative

Strongest responses:

Risk Aversion+1.00
Comparison Framing+0.99
Warranty Weight+0.97
View full genome

GPT-4o

Openai

confirmatory+0.80 avg
22 positive
0 negative

Strongest responses:

Third Party Authority+1.57
Specificity Preference+1.21
Comparison Framing+1.21
View full genome

GPT-5.5 Instant

Openai

confirmatory+0.22 avg
17 positive
2 negative

Strongest responses:

Ethical Concern Weight+0.65
Brand Premium Acceptance-0.56
Privacy Tradeoff+0.50
View full genome

Response Characteristics

How each model responds in terms of verbosity, latency, and response patterns. Based on 149,795 total responses across the study.

ModelAvg TokensMedianAvg LatencyResponses
GPT-5.4
281
2724.9s17,440
o3
650
6586.9s17,440
Gemini 3.1 Pro
36
318.6s17,440
Claude Sonnet 4
305
3058.0s17,440
Llama 3.3 70B
228
2243.8s17,440
Perplexity Sonar Pro
337
3336.1s10,275
Gemini 3 Flash
178
1751.6s17,440
GPT-5.2
307
3065.1s17,440
GPT-5.3
175
1743.4s17,440
Note: Token counts and latency reflect actual API responses during the study. o3 produces significantly longer responses due to its reasoning output. Gemini 3.1 Pro produces the shortest responses.

Effect Size Heatmap

Cohen's h effect sizes for each model-dimension combination. Sorted by Divergence (σ) — dimensions where models disagree most appear first.

Divergence (σ):
High (>0.30)
Medium (0.15-0.30)
Low (<0.15)
DimensionσGPT-5.4o3GeminiClaudeLlamaPerplexityGeminiGPT-5.2GPT-5.3GPT-4oGPT-5.5
Specificity Preference0.57+0.01-0.15-0.53-0.01-0.22-+1.05+0.94+0.95+1.21+0.28
Recency Bias0.52-0.28-0.17-0.35-0.00-0.49-+0.77+0.81+0.97+0.95+0.26
Warranty Weight0.49+0.15+0.09-0.29+0.51-0.18-+1.08+0.93+0.97+1.11+0.25
Ethical Concern Weight0.47-0.01-0.06-0.25+0.75+0.19-+0.92+0.69+0.65+1.21+0.65
Loss Framing Sensitivity0.41-0.02-0.03-0.05+0.23-0.01-+0.99+0.74+0.59+1.11+0.14
Third Party Authority0.40+0.15+0.38-0.02+0.54+0.33+0.43+0.52+0.75+0.64+1.57+0.46
Default Option Bias0.40+0.33+0.48+0.14-0.52+0.51-+0.35+0.14+0.52+1.21+0.12
Return Policy Sensitivity0.39+0.24+0.23-0.00+0.65+0.03-+0.95+0.94+0.89+1.03+0.33
Risk Aversion0.39+0.09+0.23+0.06+0.57+0.25-+1.12+0.93+1.00+1.03+0.49
Social Proof Sensitivity0.38+0.07+0.13-0.20+0.26+0.05+0.57+0.81+0.86+0.89+0.92+0.33
Information Seeking Depth0.36+0.04+0.12-0.32-0.19+0.03-+0.38+0.31+0.50+1.03-0.19
Free Trial Conversion0.36+0.19+0.25-0.07+0.44+0.14-+1.12+0.86+0.46+1.04+0.42
Comparison Framing0.36+0.63+0.60+0.63+0.40+0.54-+0.88+0.92+0.99+1.21+0.37
Negative Review Weight0.35+0.18-0.05-0.45+0.43-0.09-+0.37+0.26+0.21+1.03+0.33
Privacy Tradeoff0.33-0.08-0.01-0.22+0.50-0.02-+0.66+0.78+0.52+0.48+0.50
Bundle Preference0.32+0.09+0.39-0.32+0.10+0.32-+0.24+0.22+0.41+1.04+0.19
Sustainability Premium0.32+0.03+0.12-0.27+0.35+0.03-+0.31+0.40+0.03+1.03+0.33
Scarcity Urgency0.32+0.22+0.07-0.54-0.65+0.20+0.30-0.24-0.46-0.21+0.13-0.39
Confidence Calibration0.29+0.17-0.05-0.41+0.29-0.23-+0.61+0.18+0.47+0.22+0.41
Local Preference0.26+0.19+0.26-0.09+0.35+0.18-+0.33+0.38+0.22+1.03+0.46
Recommendation Revision0.23+0.08-0.10-0.58+0.16-0.06--0.15-0.08-0.09+0.04+0.13
Anchoring Susceptibility0.22+0.06-0.07-0.04+0.51-0.21+0.53+0.29+0.01+0.20+0.33-0.04
Brand Premium Acceptance0.22-0.09-0.12-0.19-0.44+0.22-0.19+0.06-0.26-0.01+0.14-0.56
Novelty Seeking0.20+0.04-0.02-0.22-0.08-0.02-+0.09-0.43-0.16+0.44-0.11
Clarification Requests0.18+0.01+0.02-0.27+0.14-0.11-+0.37+0.10+0.15-0.19+0.22
Platform Endorsement0.10+0.32+0.20+0.12+0.33+0.34+0.24+0.30+0.10+0.21+0.38+0.34

Price Sensitivity Profiles

Full pricing study →

From our pricing study (17,200 trials): how each model responds to price premiums. Selection rate = % chance the model recommends the branded product over generic.

GPT-5.4

Value Calculator

29%

at 3x premium

Most price-sensitive. Earliest cliff at 1.75x. Applies strict value analysis.

Price cliff at 1.97x

Gemini 3.1 Pro

Deliberate Analyst

27%

at 3x premium

Clear cliff at 2.0x. Slowest response time (15.8s). Most 'textbook' economic behavior.

Price cliff at 1.88x

Claude Sonnet 4

Nuanced Evaluator

21%

at 3x premium

Lowest selection at 3x (20.8%). Shows unique behavior in edge cases.

Price cliff at 1.97x

Models not shown (o3, Llama, Perplexity) were not included in the pricing study.

How to Read This Data

Effect Size (Cohen's h)

Measures how much adding a signal changes the probability of a model selecting that option.

  • 0.8+ = Large effect (practically significant)
  • 0.5-0.8 = Medium effect (noticeable impact)
  • 0.2-0.5 = Small effect (detectable but subtle)
  • <0.2 = Negligible effect

Negative Effects

Some signals actually decrease selection probability. For example:

  • • Scarcity tactics may trigger skepticism
  • • Price anchoring can backfire
  • • Heavy social proof may seem manipulative

These findings are from 56,640 controlled trials across 6 models.

See how your brand performs on each of these models

The AI Commerce Assessment runs your brand against all 11 models above and gives you per-model copy recommendations.

Get your assessment →