Opponent portrait for multiagent reinforcement learning in competitive environment

INTERNATIONAL JOURNAL OF INTELLIGENT SYSTEMS(2021)

引用 14|浏览17
暂无评分
摘要
Existing investigations of opponent modeling and intention inferencing cannot make clear descriptions and practical explanations of the opponent's behaviors and intentions, which may inevitably limit the applicability of them. In this work, we propose a novel approach for opponent's policy explanation and intention inference based on the behavioral portrait of opponent. Specifically, we use the multiagent deep deterministic policy gradients (MADDPG) algorithm to train the agent and opponent in the competitive environment, and collect the behavioral data of opponent based on agent's observations. Then we perform pattern segmentation and extract the opponent's behavior events via Toeplitz inverse covariance-based clustering (TICC) algorithm; hence the opponent's behavior data can be encoded into a knowledge graph, named opponent's behavior knowledge graph (OKG). Based on this, we built a question-answer system (QA system) to query and match opponent historical information in OKG, so that the agent can obtain additional experience and gradually infer the intention of opponent with the episodes of iteration. We evaluate the proposed method on the competitive scenario in multiagent particle environment (MPE). Simulation results show that the agents are able to learn better policies with opponent portrait in competitive settings.
更多
查看译文
关键词
deep reinforcement learning, intention inference, knowledge graph, multiagent system, opponent modeling
AI 理解论文
溯源树
样例
生成溯源树,研究论文发展脉络
Chat Paper
正在生成论文摘要