Active Reward Learning for Co-Robotic Vision Based Exploration in Bandwidth Limited Environments

© 2020 IEEE. We present a novel POMDP problem formulation for a robot that must autonomously decide where to go to collect new and scientifically relevant images given a limited ability to communicate with its human operator. From this formulation we derive constraints and design principles for the...

Full description

Bibliographic Details
Main Authors:	Jamieson, Stewart Christopher (Author), How, Jonathan P (Author), Girdhar, Yogesh (Author)
Other Authors:	Joint Program in Applied Ocean Physics and Engineering (Contributor), Woods Hole Oceanographic Institution (Contributor)
Format:	Article
Language:	English
Published:	IEEE, 2021-12-14T18:22:01Z.
Subjects:	Article
Online Access:	Get fulltext


LEADER	01677 am a22002053u 4500
001	137153.2
042			\|a dc
100	1	0	\|a Jamieson, Stewart Christopher. \|e author
100	1	0	\|a Joint Program in Applied Ocean Physics and Engineering \|e contributor
100	1	0	\|a Woods Hole Oceanographic Institution \|e contributor
700	1	0	\|a How, Jonathan P \|e author
700	1	0	\|a Girdhar, Yogesh \|e author
245	0	0	\|a Active Reward Learning for Co-Robotic Vision Based Exploration in Bandwidth Limited Environments
260			\|b IEEE, \|c 2021-12-14T18:22:01Z.
856			\|z Get fulltext \|u https://hdl.handle.net/1721.1/137153.2
520			\|a © 2020 IEEE. We present a novel POMDP problem formulation for a robot that must autonomously decide where to go to collect new and scientifically relevant images given a limited ability to communicate with its human operator. From this formulation we derive constraints and design principles for the observation model, reward model, and communication strategy of such a robot, exploring techniques to deal with the very high-dimensional observation space and scarcity of relevant training data. We introduce a novel active reward learning strategy based on making queries to help the robot minimize path regret online, and evaluate it for suitability in autonomous visual exploration through simulations. We demonstrate that, in some bandwidth-limited environments, this novel regret-based criterion enables the robotic explorer to collect up to 17% more reward per mission than the next-best criterion.
546			\|a en
655	7		\|a Article
773			\|t 10.1109/ICRA40945.2020.9196922
773			\|t Proceedings - IEEE International Conference on Robotics and Automation

Active Reward Learning for Co-Robotic Vision Based Exploration in Bandwidth Limited Environments

Similar Items