LessWrong (30+ Karma)

“AI CoT Reasoning Is Often Unfaithful” by Zvi

Apr 4, 2025

15:33

A new Anthropic paper reports that reasoning model chain of thought (CoT) is often unfaithful. They test on Claude Sonnet 3.7 and r1, I’d love to see someone try this on o3 as well.

Note that this does not have to be, and usually isn’t, something sinister.

It is simply that, as they say up front, the reasoning model is not accurately verbalizing its reasoning. The reasoning displayed often fails to match, report or reflect key elements of what is driving the final output. One could say the reasoning is often rationalized, or incomplete, or implicit, or opaque, or bullshit.

The important thing is that the reasoning is largely not taking place via the surface meaning of the words and logic expressed. You can’t look at the words and logic being expressed, and assume you understand what the model is doing and why it is doing [...]

---

Outline:

(01:03) What They Found

(06:54) Reward Hacking

(09:28) More Training Did Not Help Much

(11:49) This Was Not Even Intentional In the Central Sense

---

First published:
April 4th, 2025

Source:
https://www.lesswrong.com/posts/TmaahE9RznC8wm5zJ/ai-cot-reasoning-is-often-unfaithful

---

Narrated by TYPE III AUDIO.

---