| Title |
Progress-Aware Episodic Memory Intrinsic Reward for Sparse Reward Environments |
| Authors |
장수영(Ahyun Lee) ; 박시연(Siyeon Park) ; 우성필(Sooyoung Jang) ; 이아현(Sungp) |
| DOI |
https://doi.org/10.5573/ieie.2026.63.9.86 |
| Keywords |
Deep reinforcement learning; Sparse reward; exploration; Episodic memory; Intrinsic reward |
| Abstract |
Deep reinforcement learning has demonstrated remarkable performance in various decision-making problems; however, it suffers from severe exploration difficulties in sparse reward environments where rewards are only given upon reaching the goal. While episodic memory (EM) techniques providing intrinsic rewards based on visitation history have been studied to mitigate this issue, existing methods often fail to achieve stable convergence by continuing unnecessary exploration even after making meaningful progress. In this paper, we propose a Progress-Aware Episodic Memory (PA-EM) intrinsic reward method to resolve this exploration-exploitation dilemma. The proposed method utilizes a random convolutional neural network to embed states into low-dimensional vectors and calculates novelty using the k-nearest neighbors (k-NN) algorithm. Crucially, it dynamically balances exploration and exploitation by exponentially decaying the intrinsic reward whenever the agent achieves a progressive task, such as opening a door, alongside applying a linear global decay over total timesteps. Experimental results in the MiniGrid multi-room environment demonstrate that the proposed method significantly improves learning stability compared to baseline PPO and standard EM models, achieving a superior success rate of 95.0% at 1,000,000 timesteps. |