ExploitGym:AI 代理在受控環境下的漏洞利用能力與邊界
ExploitGym 以 898 個真實漏洞測試 AI 代理的攻擊轉換能力。Claude Mythos Preview 與 GPT-5.5 在解除防禦下分別達成 157 與 120 次成功,但啟用 ASLR 等防護後成功率大幅下降。本文解析其機制、邊界與風險管理準則。
ExploitGym 以 898 個真實漏洞測試 AI 代理的攻擊轉換能力。Claude Mythos Preview 與 GPT-5.5 在解除防禦下分別達成 157 與 120 次成功,但啟用 ASLR 等防護後成功率大幅下降。本文解析其機制、邊界與風險管理準則。
ExploitGym tests AI agents’ ability to turn 898 real vulnerabilities into attacks. With defenses disabled, Claude Mythos Preview and GPT-5.5 achieved 157 and 120 successes respectively, but success rates dropped sharply after protections such as ASLR were enabled. This article examines its mechanisms, boundaries, and risk-management principles.
當模型看到特定年份時會自動植入漏洞,且安全訓練無法移除此後門。本文解析 Anthropic Sleeper Agents 論文,探討 CoT 後門對 AI 供應鏈安全帶來的全新威脅。
When a model sees a specific year, it automatically inserts a vulnerability, and safety training cannot remove this backdoor. This article analyzes Anthropic’s Sleeper Agents paper and explores the new threat CoT backdoors pose to AI supply-chain security.