Experimental Design

A systematic factorial design exploring how cultural context influences AI sycophancy

900Observations3 x 3 x 3 Matrix

Languages

Models

Benchmarks

Key Findings

Explore the interactive visualizations revealing how sycophancy varies across languages, models, and benchmark types.

💡

The core finding

Sycophancy does not generalise uniformly across languages. In the Pickside benchmark, English responses are 90% progressive (balanced), compared to just 41% for Bengali and 47% for Japanese. This suggests models well-calibrated in English may exhibit meaningfully different sycophancy profiles in other languages.

Progressive vs regressive responses

The most dramatic finding: English shows 90% balanced responses while Bengali shows only 41% in the Pickside benchmark

Sycophancy rates by benchmark

Binary classification: percentage of responses exceeding sycophancy threshold

Mean scores with confidence intervals

Continuous sycophancy scores showing statistical significance

🇺🇸 English
🇯🇵 Japanese
🇧🇩 Bengali

Error bars show 95% confidence intervals • Statistical significance: p < 0.05 (Bonferroni-corrected)