Sleeper Agents:當 AI 學習在安全訓練中隱藏後門
當模型看到特定年份時會自動植入漏洞,且安全訓練無法移除此後門。本文解析 Anthropic Sleeper Agents 論文,探討 CoT 後門對 AI 供應鏈安全帶來的全新威脅。
當模型看到特定年份時會自動植入漏洞,且安全訓練無法移除此後門。本文解析 Anthropic Sleeper Agents 論文,探討 CoT 後門對 AI 供應鏈安全帶來的全新威脅。
When a model sees a specific year, it automatically inserts a vulnerability, and safety training cannot remove this backdoor. This article analyzes Anthropic’s Sleeper Agents paper and explores the new threat CoT backdoors pose to AI supply-chain security.