NHacker Next
  • new
  • past
  • show
  • ask
  • show
  • jobs
  • submit
Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces (arxiv.org)
florianherrengt 18 hours ago [-]
> While a human may say “aha” to indicate exactly a sudden internal state change, this interpretation is unwarranted for models which do not have any such internal state, and which on the next forward pass will only differ from the pre-aha pass by the inclusion of that single token in their context. Interpreting the “aha” moment as meaningful exemplifies the long-neglected assumption about long CoT models – the false idea that derivational traces are semantically meaningful, either in resemblance to algorithm traces or to human reasoning.

This paper addresses something that has always bothered me about LLMs. You read their reasoning, see something like “Wait, that’s wrong” and then watch them make the exact mistake they just identified.

Jeff_Brown 16 hours ago [-]
By itself, "aha" carries no insight, but the insight is probably stated immediately after it. In that case the aha is semantically useful, by identifying the insight it is near.
paimapi 14 hours ago [-]
it's a rhetorical heuristic that a writer should know to use when directing a reader to a declarative that they want them to pay attention to, usually because it's a non-obvious or roundabout insight

when utilized by AI, it's a probabilistic output and it's variable whether or not that rhetorical trick is useful. it also pushes a non-skeptical reader to focus too much on the following text or even to believe that they, themselves, derived some insight. this is effectively a kind of persuasive sophistry which is not helpful - adding rules around it prevents people from deluding themselves with AI

abitmoa 11 hours ago [-]
It amounts to noise overall, but it has further unwanted and potentially misleading 'properties'. I think it's rather sobering to see how much bandwidth is still being wasted.
wizzwizz4 14 hours ago [-]
> but the insight is probably stated immediately after it.

If the intermediate tokens represent reasoning or thought, you would expect "aha" to occur after the thoughts that led to the realisation, including the thoughts encoding the explanation: they don't have any other state. There is no reason to draw the conclusion you've drawn. Furthermore, what LLMs are doing isn't thought.

basedpolymer 14 hours ago [-]
The anthropomorphization of LLMs should be discouraged as much as possible. It perpetuates bad practices and encourages the use of these bots for tasks they are not intended for (particularly as chatbots).

Thinking traces should be treated as black boxes. There is no point in reading them. Only the LLMs’ conclusions are relevant. This is particularly true of Opus 5, which employs reasoning that seems highly questionable but very often reaches excellent conclusions (compared to its peers)

FloorEgg 12 hours ago [-]
Sometimes I monitor thinking traces for misunderstandings (missing context / bad assumptions). If it's going to go off on a ~20 min task and I can catch it's going in the wrong direction in the first minute I save a lot of tokens and wasted time. I don't monitor the whole thing, mostly just the first bit to see if there was a gap or misalignment in intention.

As an aside, anthropomorphization has nothing to do with my motivations.

smugtrain 5 hours ago [-]
Strong dislike for papers that tell me what to do in the title, especially when even the paper admits a loose correlation of the intermediate tokens compared to solution correctness. My solutions work and they speak for themselves.
Terr_ 18 hours ago [-]
I've been calling them film noir internal monologues, within the documents being generated by the LLM which happen to look like movie scripts.

In other words, it isn't qualitatively different from character dialogue. "Keep cheese on your pizza by using glue" is the same problem regardless of whether the script calls for the character to speak it out-loud or not.

clhodapp 16 hours ago [-]
Seems like they are closer to scratch than reasoning... Generating some scratch to draw from helps make it easier to compute the real answer.
cyanydeez 14 hours ago [-]
I assume theyre searching the local gradient to see if theres a better descent before proceeding.
eigenspace 2 hours ago [-]
LLMs dont do gradient descent to generate tokens.

They are trained by gradient descent, but inference doesnt involve it.

c0_0p_ 6 hours ago [-]
I don't think there's anything like that going on. They just word vomit into a secondary area, and then there is an internal prompt that says "clean this up and summarize for the user".
porridgeraisin 14 hours ago [-]
Related:

Poster side dialogue and Q&A about this work at ICML.

https://news.ycombinator.com/item?id=49277303

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact
Rendered at 09:38:21 GMT+0000 (Coordinated Universal Time) with Vercel.