Non-Linear Monte-Carlo Search in Civilization II

This paper presents a new Monte-Carlo search algorithm for very large sequential decision-making problems. Our approach builds on the recent success of Monte-Carlo tree search algorithms, which estimate the value of states and actions from the mean outcome of random simulations. Instead of using a s...

Full description

Bibliographic Details
Main Authors:	Branavan, Satchuthanan R. (Contributor), Silver, David (Author), Barzilay, Regina (Contributor)
Other Authors:	Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory (Contributor)
Format:	Article
Language:	English
Published:	AAAI Press/International Joint Conferences on Artificial Intelligence, 2012-10-24T20:34:34Z.
Subjects:	Article
Online Access:	Get fulltext


LEADER	02334 am a22002653u 4500
001	74248
042			\|a dc
100	1	0	\|a Branavan, Satchuthanan R. \|e author
100	1	0	\|a Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory \|e contributor
100	1	0	\|a Barzilay, Regina \|e contributor
100	1	0	\|a Branavan, Satchuthanan R. \|e contributor
100	1	0	\|a Barzilay, Regina \|e contributor
700	1	0	\|a Silver, David \|e author
700	1	0	\|a Barzilay, Regina \|e author
245	0	0	\|a Non-Linear Monte-Carlo Search in Civilization II
260			\|b AAAI Press/International Joint Conferences on Artificial Intelligence, \|c 2012-10-24T20:34:34Z.
856			\|z Get fulltext \|u http://hdl.handle.net/1721.1/74248
520			\|a This paper presents a new Monte-Carlo search algorithm for very large sequential decision-making problems. Our approach builds on the recent success of Monte-Carlo tree search algorithms, which estimate the value of states and actions from the mean outcome of random simulations. Instead of using a search tree, we apply non-linear regression, online, to estimate a state-action value function from the outcomes of random simulations. This value function generalizes between related states and actions, and can therefore provide more accurate evaluations after fewer simulations. We apply our Monte-Carlo search algorithm to the game of Civilization II, a challenging multi-agent strategy game with an enormous state space and around $10^{21}$ joint actions. We approximate the value function by a neural network, augmented by linguistic knowledge that is extracted automatically from the official game manual. We show that this non-linear value function is significantly more efficient than a linear value function. Our non-linear Monte-Carlo search wins 80\% of games against the handcrafted, built-in AI for Civilization II.
520			\|a National Science Foundation (U.S.) (CAREER grant IIS-0448168)
520			\|a National Science Foundation (U.S.) (grant IIS-0835652)
520			\|a United States. Defense Advanced Research Projects Agency (DARPA Machine Reading Program (FA8750-09-C-0172))
520			\|a Microsoft Research (New Faculty Fellowship)
546			\|a en_US
655	7		\|a Article
773			\|t Proceedings of the Twenty-second International Joint Conference on Artificial Intelligence

Non-Linear Monte-Carlo Search in Civilization II

Similar Items