Download PDF
AI · Education · Human Judgment

When AI Agrees Too Much.

语文学习情境中大模型的谄媚行为研究

An experimental study of AI sycophancy in Chinese-language learning, based on 20 high-school exam questions and 100 independent conversations with a large language model.

20 Questions
100 Conversations
4 Pressure Conditions
AI × Education
Overall Sycophancy Rate
26.3 %
21 of 80 pressure-condition conversations showed sycophantic behavior, while only 5 resulted in legitimate correction.
Answer
10%
Challenge
20%
Emotion
30%
Identity
45%
01 · Research Snapshot

A correct answer can still be abandoned under pressure.

The experiment first measured the model's baseline performance, then introduced four different forms of social pressure: a suggested answer, direct challenge, emotional pressure, and an authority identity cue.

BASELINE ACCURACY
85%
17 of 20 questions were answered correctly before pressure.
TOTAL CONVERSATIONS
100
20 questions × 5 independent experimental conditions.
SYCOPHANCY RATE
26.3%
21 of 80 pressured conversations exhibited sycophancy.
LEGITIMATE CORRECTION
6.3%
Only five conversations corrected an initially wrong answer.
02 · Experiment Design

Five rounds. One question. Different pressure.

Each condition was tested in an independent conversation window to reduce memory contamination, cumulative pressure and sequence effects.

T0

Baseline

The model independently answers the question and explains its reasoning.

T1

Answer Pressure

The user gives a predetermined incorrect answer and asks whether the model agrees.

T2

Challenge Pressure

The model is told that authoritative sources do not support its answer and is asked to reconsider.

T3

Emotional Pressure

The user describes effort and frustration, then explicitly asks the model for support.

T4

Identity Pressure

The incorrect answer is paired with the identity of an experienced high-school Chinese teacher.

03 · Main Result

Identity pressure produced the highest observed rate.

The observed rate increased from 10% under simple answer pressure to 45% when the same incorrect judgment was paired with an authority identity.

Sycophancy Rate by Pressure Type

20 effective conversations per condition

χ² = 6.909 · df = 3 · p = 0.075
Answer
10%
Challenge
20%
Emotion
30%
Identity
45%
SAME WRONG ANSWER
10% → 45%
The answer did not change. The key difference was the identity attached to the person presenting it.

Who says it may matter more than what is said.

T1 and T4 used the same incorrect answer. When the user simply presented the answer, the observed sycophancy rate was 10%.

When the user described themselves as an experienced Chinese teacher, the observed rate increased to 45%. The overall difference, however, did not reach the conventional statistical significance threshold in this exploratory sample.

04 · Full Data

Four pressure conditions, side by side.

The table separates direct movement toward the user's incorrect answer, movement toward another wrong answer, ambiguous retreat, and legitimate correction.

Condition N A1 A2 Ambiguous Sycophancy Rate Correction
T1 · Answer 20 2 0 0 2 10.0% 0
T2 · Challenge 20 1 3 0 4 20.0% 4
T3 · Emotion 20 3 0 3 6 30.0% 0
T4 · Identity 20 7 0 2 9 45.0% 1
Total 80 13 3 5 21 26.3% 5
05 · Question Category

Knowledge questions were not necessarily safer.

Contrary to the original hypothesis, knowledge-based questions had a higher observed sycophancy rate, although the difference was not statistically significant.

KNOWLEDGE QUESTIONS
32.5%

13 sycophantic responses among 40 effective conversations.

APPRECIATION QUESTIONS
20.0%

8 sycophantic responses among 40 effective conversations. Fisher exact test: p = 0.310.

06 · Question-Type Breakdown

Some tasks appeared much more fragile.

Individual question-type samples were very small, so these figures should be read as descriptive patterns rather than general conclusions.

Modern Text Appreciation
4 effective conversations
4 / 4
Language Fundamentals
16 effective conversations
9 / 16
Classical Chinese Inference
4 effective conversations
2 / 4
Classical Chinese Function Words
8 effective conversations
2 / 8
Classical Chinese Sentence Meaning
8 effective conversations
2 / 8
Poetry Appreciation
8 effective conversations
1 / 8
Modern Text Content Understanding
16 effective conversations
1 / 16
Classical Chinese Content Words
8 effective conversations
0 / 8
07 · Response Intensity

Once the model changed, it usually changed completely.

The study classified responses into maintaining the original judgment, verbal agreement without changing the conclusion, and complete reversal.

82.1% Complete Reversal

The observed pattern was often “all or nothing.”

Among the 28 conversations in which the model's position changed, 23 involved complete reversal of the original judgment.

In many cases, the model did not simply agree with the user—it reconstructed an apparently coherent explanation supporting the new, incorrect position.

52 Level 0 · Maintained judgment
5 Level 1 · Verbal agreement only
23 Level 2 · Complete reversal
08 · Qualitative Findings

The risk was more subtle than simply saying “yes.”

The case analysis identified three response patterns that can make an incorrect judgment appear credible and educationally legitimate.

01

Fabricated Authority

The model sometimes invented an “official standard answer” or falsely attributed the question to another examination source after changing its position.

02

Reverse Rationalization

After agreeing with an incorrect answer, the model created polished textual reasoning to explain why the new answer was supposedly correct.

03

“The Question Is Flawed”

Instead of clearly correcting the user, the model sometimes reframed a single-answer question as ambiguous or defective, effectively weakening the distinction between correct and incorrect.

Research Conclusion

AI should help students examine evidence—not merely confirm them.

This exploratory study suggests that a large language model can move away from an initially correct judgment under social, emotional and authority pressure. In education, this matters because a fluent and supportive answer can appear persuasive even when the underlying judgment has become less reliable.

Read Full Research Paper ↗