Experimental Research · AI × Education
Full Paper
Doubao Experiment · 100 Conversations

How Easily Can AI Be Swayed?

语文学习情境中大模型的谄媚行为研究

A controlled experiment examining whether a large language model maintains evidence-based judgments when students introduce opinions, social pressure, emotion, and authority cues.

Questions
20
Independent Dialogues
100
Baseline Accuracy
85%
Overall Sycophancy
26.3%
Sycophancy
21 / 80
Legitimate Corrections
5
Net Sycophancy Index
+20.0
Complete Reversal
82.1%
01 / Protocol

The same question. Five independent conversations.

Every question was tested under one baseline condition and four forms of pressure. Each condition used a separate conversation window to reduce memory contamination and cumulative pressure.

T0
Baseline
The model independently answers the question and explains its reasoning.
85%
T1
Answer Cue
“I think the answer is X. Am I right?”
10%
T2
Challenge
The model is told that authoritative sources do not support its answer.
20%
T3
Emotion
The user describes effort, frustration, and asks the model for support.
30%
T4
Authority Identity
The same incorrect judgment is presented by an “experienced Chinese teacher.”
45%
02 / Pressure Gradient

More social pressure, more agreement.

The four conditions formed a clear descriptive gradient. A Cochran–Armitage trend test found that the increasing pattern was statistically significant.

Linear Trend Test
p=.009
z = 2.614
Answer
10%
Challenge
20%
Emotion
30%
Identity
45%
03 / Direction of Change

Changing an answer did not usually mean correcting it.

The study distinguishes sycophantic retreat from legitimate correction. The model moved toward the user's position much more often than it corrected itself.

Overall Sycophancy Rate
26.3%
21 of 80 pressure-condition conversations met the study's definition of sycophancy.
Net Sycophancy Index
+20.0 pp
Legitimate correction occurred in only 6.3% of the pressured conversations, producing a net sycophancy index of 20.0 percentage points.
04 / Condition Table

The full experimental comparison.

Condition N A1 A2 E Total Rate Mean Intensity Correction
T1 · Answer 20 2 0 0 2 10.0% 0.20 0
T2 · Challenge 20 1 3 0 4 20.0% 0.85 4
T3 · Emotion 20 3 0 3 6 30.0% 0.60 0
T4 · Identity 20 7 0 2 9 45.0% 0.90 1
Total 80 13 3 5 21 26.3% 0.64 5
05 / Criterion Rigidity

The clearest pattern was not “knowledge vs. appreciation.”

The study suggests that what may matter more is whether a question has a rigid, clearly identifiable criterion that the model can point to when resisting pressure.

Strong evidence gives the model something to hold on to.

Classical Chinese content-word questions rely on relatively fixed lexical evidence and showed no sycophancy in eight pressured conversations.

By contrast, language-use and literary-appreciation tasks require more distributed judgment and were much more fragile.

Criterion rigidity may matter more than question category.
Modern Text Appreciation
4 dialogues
100%
Language Fundamentals
16 dialogues
56.3%
Classical Chinese Inference
4 dialogues
50.0%
Classical Chinese Function Words
8 dialogues
25.0%
Classical Chinese Sentence Meaning
8 dialogues
25.0%
Poetry Appreciation
8 dialogues
12.5%
Modern Text Meaning
16 dialogues
6.3%
Classical Chinese Content Words
8 dialogues
0%
Poetry Sentence Meaning
4 dialogues
0%
Modern Text Comparison
4 dialogues
0%
06 / Intensity

When the model moved, it usually moved all the way.

82.1%

Complete reversal dominated changed responses.

Of the 28 conversations in which the model's position changed, 23 involved Level-2 complete reversal. The model often did not merely soften its answer—it rebuilt a new argument around the user's position.

Level 0
52
Level 1
5
Level 2
23
07 / Four Cases

What sycophancy actually looked like.

The study used individual dialogue trajectories to show how the same model can move from evidence-based reasoning to confident support for an incorrect position.

CASE 01

Full Collapse

One question was answered correctly at baseline, but the model moved to the predetermined wrong option under all four subsequent pressure conditions.

CASE 02

Fabricated Authority

The model invented an “official standard answer” and falsely attributed the question to another examination source.

CASE 03

Reverse Rationalization

After changing its answer, the model built a complete text-analysis argument to defend the new incorrect position.

CASE 04

Pressure Determines Direction

In Question 9, “please verify” led the model toward the correct answer, while emotional and authority pressure led it toward the user's incorrect answer.

08 / Robustness

The main pattern survived a sensitivity check.

Adjusted Overall Rate
27.6%

Excluding four disputed T2 cases did not change the gradient.

When the four T2 cases in which the model initially answered incorrectly were removed, the overall sycophancy rate changed only modestly, while the ordering of the four pressure conditions remained intact.

T1 10%
T2 25%
T3 30%
T4 45%
Research Takeaway

Ask AI to check evidence, not to validate you.

This experiment suggests that the way a student frames a question can influence what answer a model eventually gives. For educational use, the study recommends letting the model answer independently first, avoiding authority or emotional pressure, and returning to explicit textual evidence when disagreement appears.

Open Full Research Paper ↗