Post-trained LLM 如何認出自己在說話?
IAS 與 Anthropic 研究者的論文發現,Post-trained 模型生成自身回答時的熵值比閱讀外部文本低 3-4 倍。本文探討「輸入驚喜度」如何驅動這種自我認知機制,並分析其對 AI Agent 穩定性的影響。
IAS 與 Anthropic 研究者的論文發現,Post-trained 模型生成自身回答時的熵值比閱讀外部文本低 3-4 倍。本文探討「輸入驚喜度」如何驅動這種自我認知機制,並分析其對 AI Agent 穩定性的影響。
A paper by researchers at IAS and Anthropic finds that post-trained models’ output-distribution entropy is 3–4 times lower when generating their own responses than when reading external text. This article explores how Input Surprise drives this self-recognition mechanism and analyzes its implications for AI agent stability.