漸進式非平穩環境中,穩定性何時勝於可塑性?
在漸進式非平穩環境中,強化學習效能衰退的主因是過度適應導致的不穩定性。這篇研究指出,將突觸鞏固應用於多時間尺度的後繼特徵(SFs)可提升表現,但需承擔計算延遲與 SGD 優化器依賴的代價。
在漸進式非平穩環境中,強化學習效能衰退的主因是過度適應導致的不穩定性。這篇研究指出,將突觸鞏固應用於多時間尺度的後繼特徵(SFs)可提升表現,但需承擔計算延遲與 SGD 優化器依賴的代價。
In gradually non-stationary environments, the main cause of reinforcement learning performance decline is instability caused by over-adaptation. This study finds that applying synaptic consolidation to multi-timescale Successor Features (SFs) can improve performance, but at the cost of computational latency and dependence on SGD optimizers.