The Experts below are selected from a list of 261 Experts worldwide ranked by ideXlab platform

Yurika Tani - One of the best experts on this subject based on the ideXlab platform.

  • Optimality Principle broken by considering structured plant variation and relevant robust reinforcement learning
    Systems Man and Cybernetics, 2011
    Co-Authors: Kei Senda, Yurika Tani
    Abstract:

    In a general reinforcement learning problem, a plant (state transition probabilities) is estimated and a learning policy for the estimated plant is applied to a real plant. If there are differences between the estimated plant and the real plant, the obtained policy may not work for the real plant. Therefore, a set of plants with variations is used for learning in order to obtain a robust policy against variations. Bellman's Principle of Optimality does not hold when the set of plants is used, and a typical dynamic programming algorithm cannot solve the problem. This study shows the reason why the Principle of Optimality does not hold. It then makes some relaxed problems whose solutions can be obtained. Moreover, this study proposes solutions to learn feasible policies efficiently. The effectiveness of the proposed method is demonstrated by applying to simple examples.

  • SMC - Optimality Principle broken by considering structured plant variation and relevant robust reinforcement learning
    2011 IEEE International Conference on Systems Man and Cybernetics, 2011
    Co-Authors: Kei Senda, Yurika Tani
    Abstract:

    In a general reinforcement learning problem, a plant (state transition probabilities) is estimated and a learning policy for the estimated plant is applied to a real plant. If there are differences between the estimated plant and the real plant, the obtained policy may not work for the real plant. Therefore, a set of plants with variations is used for learning in order to obtain a robust policy against variations. Bellman's Principle of Optimality does not hold when the set of plants is used, and a typical dynamic programming algorithm cannot solve the problem. This study shows the reason why the Principle of Optimality does not hold. It then makes some relaxed problems whose solutions can be obtained. Moreover, this study proposes solutions to learn feasible policies efficiently. The effectiveness of the proposed method is demonstrated by applying to simple examples.

Kei Senda - One of the best experts on this subject based on the ideXlab platform.

  • Optimality Principle broken by considering structured plant variation and relevant robust reinforcement learning
    Systems Man and Cybernetics, 2011
    Co-Authors: Kei Senda, Yurika Tani
    Abstract:

    In a general reinforcement learning problem, a plant (state transition probabilities) is estimated and a learning policy for the estimated plant is applied to a real plant. If there are differences between the estimated plant and the real plant, the obtained policy may not work for the real plant. Therefore, a set of plants with variations is used for learning in order to obtain a robust policy against variations. Bellman's Principle of Optimality does not hold when the set of plants is used, and a typical dynamic programming algorithm cannot solve the problem. This study shows the reason why the Principle of Optimality does not hold. It then makes some relaxed problems whose solutions can be obtained. Moreover, this study proposes solutions to learn feasible policies efficiently. The effectiveness of the proposed method is demonstrated by applying to simple examples.

  • SMC - Optimality Principle broken by considering structured plant variation and relevant robust reinforcement learning
    2011 IEEE International Conference on Systems Man and Cybernetics, 2011
    Co-Authors: Kei Senda, Yurika Tani
    Abstract:

    In a general reinforcement learning problem, a plant (state transition probabilities) is estimated and a learning policy for the estimated plant is applied to a real plant. If there are differences between the estimated plant and the real plant, the obtained policy may not work for the real plant. Therefore, a set of plants with variations is used for learning in order to obtain a robust policy against variations. Bellman's Principle of Optimality does not hold when the set of plants is used, and a typical dynamic programming algorithm cannot solve the problem. This study shows the reason why the Principle of Optimality does not hold. It then makes some relaxed problems whose solutions can be obtained. Moreover, this study proposes solutions to learn feasible policies efficiently. The effectiveness of the proposed method is demonstrated by applying to simple examples.

Krishnamurthy Vikram - One of the best experts on this subject based on the ideXlab platform.

  • Quickest Change Detection of Time Inconsistent Anticipatory Agents. Behavioral Economics Models in Signal Processing
    2020
    Co-Authors: Krishnamurthy Vikram
    Abstract:

    In behavioral economics, anticipatory agents are cognitive systems that make decisions by taking into account the probability of future decisions (plans). We consider the interaction between anticipatory agents and statistical detection. A sensing/computing device records the decisions of an anticipatory agent. Given this sequence of decisions, how can the sensing device achieve quickest detection of a change in the anticipatory system? From a decision theoretic point of view, anticipatory models are time inconsistent meaning that Bellman's Principle of Optimality does not hold. The appropriate formalism is the subgame Nash equilibrium. We show that the interaction between the anticipatory agents and sequential quickest detection results in unusual (nonconvex) structure of the quickest change detection policy. The methodology presented yields a useful framework for situation awareness systems and also interaction of anticipatory human decision makers with automated sequential detector

  • Quickest Change Detection of Time Inconsistent Anticipatory Agents. Human-Sensor and Cyber-Physical Systems
    'Institute of Electrical and Electronics Engineers (IEEE)', 2020
    Co-Authors: Krishnamurthy Vikram
    Abstract:

    In behavioral economics, human decision makers are modeled as anticipatory agents that make decisions by taking into account the probability of future decisions (plans). We consider cyber-physical systems involving the interaction between anticipatory agents and statistical detection. A sensing device records the decisions of an anticipatory agent. Given these decisions, how can the sensing device achieve quickest detection of a change in the anticipatory system? From a decision theoretic point of view, anticipatory models are time inconsistent meaning that Bellman's Principle of Optimality does not hold. The appropriate formalism is the subgame Nash equilibrium. We show that the interaction between anticipatory agents and sequential quickest detection results in unusual (nonconvex) structure of the quickest change detection policy. Our methodology yields a useful framework for situation awareness systems and anticipatory human decision makers interacting with sequential detectors

K.g. Garaev - One of the best experts on this subject based on the ideXlab platform.

  • A remark on the Bellman Principle of Optimality
    Journal of the Franklin Institute, 1998
    Co-Authors: K.g. Garaev
    Abstract:

    Abstract The group-theoretical approach based on joint use of the Lie-Ovsyannikov infinitesimal apparatus ( 1 – 3 ) and the Noether theory of invariant variation problems ( 4 ) is suggested for the problem of synthesis of optimum control. Up to now the ideas of the Noether theory were used on the theory of optimum control for the problems of programmed control only and considered in a small number of publications; as for the problem of synthesis, this paper is the first in this area. In this paper a corollary of the Bellman equation is obtained in the form of a linear partial differential equation. The use of the equation simplifies the problem of construction of synthesizing controls (note, that the equation is correct both in the case when no restrictions are imposed on the vector of control, and in the case when the values of the vector belong to a bounded closed domain. An example illustrating the technique of use of the Lie-Ovsyannikov infinitesimal apparatus for reduction of the corresponding Bellman equation is presented.

Yury Nikulin - One of the best experts on this subject based on the ideXlab platform.

  • Stability and accuracy functions for a multicriteria Boolean linear programming problem with parameterized Principle of Optimality “from Condorcet to Pareto”
    European Journal of Operational Research, 2010
    Co-Authors: Yury Nikulin, Marko M. Mäkelä
    Abstract:

    Abstract A multicriteria Boolean programming problem with linear cost functions in which initial coefficients of the cost functions are subject to perturbations is considered. For any optimal alternative, with respect to parameterized Principle of Optimality “from Condorcet to Pareto”, appropriate measures of the quality are introduced. These measures correspond to the so-called stability and accuracy functions defined earlier for optimal solutions of a generic multicriteria combinatorial optimization problem with Pareto and lexicographic Optimality Principles. Various properties of such functions are studied and maximum norms of perturbations for which an optimal alternative preserves its Optimality are calculated. To illustrate the way how the stability and accuracy functions can be used as efficient tools for post-optimal analysis, an application from the voting theory is considered.

  • stability analysis for a multicriteria problem with linear criteria and parameterized Principle of Optimality from lexicographic to slater
    World Academy of Science Engineering and Technology International Journal of Mathematical Computational Physical Electrical and Computer Engineering, 2009
    Co-Authors: Yury Nikulin
    Abstract:

    A multicriteria linear programming problem with integer variables and parameterized Optimality Principle ”from lexicographic to Slater” is considered. A situation in which initial coefficients of penalty cost functions are not fixed but may be potentially a subject to variations is studied. For any efficient solution, appropriate measures of the quality are introduced which incorporate information about variations of penalty cost function coefficients. These measures correspond to the so-called stability and accuracy functions defined earlier for efficient solutions of a generic multicriteria combinatorial optimization problem with Pareto and lexicographic Optimality Principles. Various properties of such functions are studied and maximum norms of perturbations for which an efficient solution preserves the property of being efficient are calculated.