Learning the Arrow of Time for Problems in Reinforcement Learning

Nasim Rahaman; Steffen Wolf; Anirudh Goyal; Roman Remme; Yoshua Bengio

Learning the Arrow of Time for Problems in Reinforcement Learning

Nasim Rahaman, Steffen Wolf, Anirudh Goyal, Roman Remme, Yoshua Bengio

Published: 20 Dec 2019, Last Modified: 05 May 2023ICLR 2020 Conference Blind SubmissionReaders: Everyone

TL;DR: We learn the arrow of time for MDPs and use it to measure reachability, detect side-effects and obtain a curiosity reward signal.

Abstract: We humans have an innate understanding of the asymmetric progression of time, which we use to efficiently and safely perceive and manipulate our environment. Drawing inspiration from that, we approach the problem of learning an arrow of time in a Markov (Decision) Process. We illustrate how a learned arrow of time can capture salient information about the environment, which in turn can be used to measure reachability, detect side-effects and to obtain an intrinsic reward signal. Finally, we propose a simple yet effective algorithm to parameterize the problem at hand and learn an arrow of time with a function approximator (here, a deep neural network). Our empirical results span a selection of discrete and continuous environments, and demonstrate for a class of stochastic processes that the learned arrow of time agrees reasonably well with a well known notion of an arrow of time due to Jordan, Kinderlehrer and Otto (1998).

Code: https://www.sendspace.com/file/0mx0en

Keywords: Arrow of Time, Reinforcement Learning, AI-Safety

Original Pdf: pdf

20 Replies

Loading