The Experts below are selected from a list of 5562 Experts worldwide ranked by ideXlab platform

Shie Mannor - One of the best experts on this subject based on the ideXlab platform.

  • generalized emphatic temporal difference learning bias variance analysis
    National Conference on Artificial Intelligence, 2016
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced emphatic temporal differences (ETD) algorithm (Sutton, Mahmood, and White 2015), which encompasses the original ETD(λ), as well as several other off-policy evaluation algorithms as special cases. We call this framework ETD(λ, β), where our introduced parameter β controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying ETD(λ, β) involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for ETD(λ, β). Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling β, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.

  • generalized emphatic temporal difference learning bias variance analysis
    arXiv: Machine Learning, 2015
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced \emph{emphatic temporal differences} (ETD) algorithm \citep{SuttonMW15}, which encompasses the original ETD($\lambda$), as well as several other off-policy evaluation algorithms as special cases. We call this framework \ETD, where our introduced parameter $\beta$ controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying \ETD\ involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for \ETD. Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling $\beta$, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.

  • emphatic td bellman Operator is a Contraction
    arXiv: Machine Learning, 2015
    Co-Authors: Assaf Hallak, Aviv Tamar, Shie Mannor
    Abstract:

    Recently, \citet{SuttonMW15} introduced the emphatic temporal differences (ETD) algorithm for off-policy evaluation in Markov decision processes. In this short note, we show that the projected fixed-point equation that underlies ETD involves a Contraction Operator, with a $\sqrt{\gamma}$-Contraction modulus (where $\gamma$ is the discount factor). This allows us to provide error bounds on the approximation error of ETD. To our knowledge, these are the first error bounds for an off-policy evaluation algorithm under general target and behavior policies.

Assaf Hallak - One of the best experts on this subject based on the ideXlab platform.

  • generalized emphatic temporal difference learning bias variance analysis
    National Conference on Artificial Intelligence, 2016
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced emphatic temporal differences (ETD) algorithm (Sutton, Mahmood, and White 2015), which encompasses the original ETD(λ), as well as several other off-policy evaluation algorithms as special cases. We call this framework ETD(λ, β), where our introduced parameter β controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying ETD(λ, β) involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for ETD(λ, β). Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling β, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.

  • generalized emphatic temporal difference learning bias variance analysis
    arXiv: Machine Learning, 2015
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced \emph{emphatic temporal differences} (ETD) algorithm \citep{SuttonMW15}, which encompasses the original ETD($\lambda$), as well as several other off-policy evaluation algorithms as special cases. We call this framework \ETD, where our introduced parameter $\beta$ controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying \ETD\ involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for \ETD. Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling $\beta$, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.

  • emphatic td bellman Operator is a Contraction
    arXiv: Machine Learning, 2015
    Co-Authors: Assaf Hallak, Aviv Tamar, Shie Mannor
    Abstract:

    Recently, \citet{SuttonMW15} introduced the emphatic temporal differences (ETD) algorithm for off-policy evaluation in Markov decision processes. In this short note, we show that the projected fixed-point equation that underlies ETD involves a Contraction Operator, with a $\sqrt{\gamma}$-Contraction modulus (where $\gamma$ is the discount factor). This allows us to provide error bounds on the approximation error of ETD. To our knowledge, these are the first error bounds for an off-policy evaluation algorithm under general target and behavior policies.

Aviv Tamar - One of the best experts on this subject based on the ideXlab platform.

  • generalized emphatic temporal difference learning bias variance analysis
    National Conference on Artificial Intelligence, 2016
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced emphatic temporal differences (ETD) algorithm (Sutton, Mahmood, and White 2015), which encompasses the original ETD(λ), as well as several other off-policy evaluation algorithms as special cases. We call this framework ETD(λ, β), where our introduced parameter β controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying ETD(λ, β) involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for ETD(λ, β). Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling β, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.

  • generalized emphatic temporal difference learning bias variance analysis
    arXiv: Machine Learning, 2015
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced \emph{emphatic temporal differences} (ETD) algorithm \citep{SuttonMW15}, which encompasses the original ETD($\lambda$), as well as several other off-policy evaluation algorithms as special cases. We call this framework \ETD, where our introduced parameter $\beta$ controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying \ETD\ involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for \ETD. Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling $\beta$, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.

  • emphatic td bellman Operator is a Contraction
    arXiv: Machine Learning, 2015
    Co-Authors: Assaf Hallak, Aviv Tamar, Shie Mannor
    Abstract:

    Recently, \citet{SuttonMW15} introduced the emphatic temporal differences (ETD) algorithm for off-policy evaluation in Markov decision processes. In this short note, we show that the projected fixed-point equation that underlies ETD involves a Contraction Operator, with a $\sqrt{\gamma}$-Contraction modulus (where $\gamma$ is the discount factor). This allows us to provide error bounds on the approximation error of ETD. To our knowledge, these are the first error bounds for an off-policy evaluation algorithm under general target and behavior policies.

He Yong - One of the best experts on this subject based on the ideXlab platform.

Remi Munos - One of the best experts on this subject based on the ideXlab platform.

  • generalized emphatic temporal difference learning bias variance analysis
    National Conference on Artificial Intelligence, 2016
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced emphatic temporal differences (ETD) algorithm (Sutton, Mahmood, and White 2015), which encompasses the original ETD(λ), as well as several other off-policy evaluation algorithms as special cases. We call this framework ETD(λ, β), where our introduced parameter β controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying ETD(λ, β) involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for ETD(λ, β). Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling β, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.

  • generalized emphatic temporal difference learning bias variance analysis
    arXiv: Machine Learning, 2015
    Co-Authors: Assaf Hallak, Aviv Tamar, Remi Munos, Shie Mannor
    Abstract:

    We consider the off-policy evaluation problem in Markov decision processes with function approximation. We propose a generalization of the recently introduced \emph{emphatic temporal differences} (ETD) algorithm \citep{SuttonMW15}, which encompasses the original ETD($\lambda$), as well as several other off-policy evaluation algorithms as special cases. We call this framework \ETD, where our introduced parameter $\beta$ controls the decay rate of an importance-sampling term. We study conditions under which the projected fixed-point equation underlying \ETD\ involves a Contraction Operator, allowing us to present the first asymptotic error bounds (bias) for \ETD. Our results show that the original ETD algorithm always involves a Contraction Operator, and its bias is bounded. Moreover, by controlling $\beta$, our proposed generalization allows trading-off bias for variance reduction, thereby achieving a lower total error.