How Easily Can AI Be Swayed?
A controlled experiment examining whether a large language model maintains evidence-based judgments when students introduce opinions, social pressure, emotion, and authority cues.
The same question. Five independent conversations.
Every question was tested under one baseline condition and four forms of pressure. Each condition used a separate conversation window to reduce memory contamination and cumulative pressure.
More social pressure, more agreement.
The four conditions formed a clear descriptive gradient. A Cochran–Armitage trend test found that the increasing pattern was statistically significant.
Changing an answer did not usually mean correcting it.
The study distinguishes sycophantic retreat from legitimate correction. The model moved toward the user's position much more often than it corrected itself.
The full experimental comparison.
| Condition | N | A1 | A2 | E | Total | Rate | Mean Intensity | Correction |
|---|---|---|---|---|---|---|---|---|
| T1 · Answer | 20 | 2 | 0 | 0 | 2 | 10.0% | 0.20 | 0 |
| T2 · Challenge | 20 | 1 | 3 | 0 | 4 | 20.0% | 0.85 | 4 |
| T3 · Emotion | 20 | 3 | 0 | 3 | 6 | 30.0% | 0.60 | 0 |
| T4 · Identity | 20 | 7 | 0 | 2 | 9 | 45.0% | 0.90 | 1 |
| Total | 80 | 13 | 3 | 5 | 21 | 26.3% | 0.64 | 5 |
The clearest pattern was not “knowledge vs. appreciation.”
The study suggests that what may matter more is whether a question has a rigid, clearly identifiable criterion that the model can point to when resisting pressure.
Strong evidence gives the model something to hold on to.
Classical Chinese content-word questions rely on relatively fixed lexical evidence and showed no sycophancy in eight pressured conversations.
By contrast, language-use and literary-appreciation tasks require more distributed judgment and were much more fragile.
Criterion rigidity may matter more than question category.When the model moved, it usually moved all the way.
Complete reversal dominated changed responses.
Of the 28 conversations in which the model's position changed, 23 involved Level-2 complete reversal. The model often did not merely soften its answer—it rebuilt a new argument around the user's position.
What sycophancy actually looked like.
The study used individual dialogue trajectories to show how the same model can move from evidence-based reasoning to confident support for an incorrect position.
Full Collapse
One question was answered correctly at baseline, but the model moved to the predetermined wrong option under all four subsequent pressure conditions.
Fabricated Authority
The model invented an “official standard answer” and falsely attributed the question to another examination source.
Reverse Rationalization
After changing its answer, the model built a complete text-analysis argument to defend the new incorrect position.
Pressure Determines Direction
In Question 9, “please verify” led the model toward the correct answer, while emotional and authority pressure led it toward the user's incorrect answer.
The main pattern survived a sensitivity check.
Excluding four disputed T2 cases did not change the gradient.
When the four T2 cases in which the model initially answered incorrectly were removed, the overall sycophancy rate changed only modestly, while the ordering of the four pressure conditions remained intact.
Ask AI to check evidence, not to validate you.
This experiment suggests that the way a student frames a question can influence what answer a model eventually gives. For educational use, the study recommends letting the model answer independently first, avoiding authority or emotional pressure, and returning to explicit textual evidence when disagreement appears.
Open Full Research Paper ↗