Seleção de abstração espacial no Aprendizado por Reforço avaliando o processo de aprendizagem

Silva, Cleiton Alves da

doi:10.11606/D.100.2018.tde-08022018-102528

Home

Facilities

Master's Dissertation

DOI

https://doi.org/10.11606/D.100.2018.tde-08022018-102528

Document

Master's Dissertation

Author

Silva, Cleiton Alves da (Catálogo USP)

Full name

Cleiton Alves da Silva

E-mail

Institute/School/College

Escola de Artes, Ciências e Humanidades

Knowledge Area

Information Systems

Date of Defense

2017-06-14

Published

São Paulo, 2018

Supervisor

Silva, Valdinei Freire da (Catálogo USP)

Committee

Silva, Valdinei Freire da (President)
Bianchi, Reinaldo Augusto da Costa
Costa, Anna Helena Reali
Delgado, Karina Valdivia

Title in Portuguese

Seleção de abstração espacial no Aprendizado por Reforço avaliando o processo de aprendizagem

Keywords in Portuguese

Aprendizado por Reforço
Seleção de abstração
Transferência do conhecimento

Abstract in Portuguese

Agentes que utilizam técnicas de Aprendizado por Reforço (AR) buscam resolver problemas que envolvem decisões sequenciais em ambientes estocásticos sem conhecimento a priori. O processo de aprendizado desenvolvido pelo agente em geral é lento, visto que se concretiza por tentativa e erro e exige repetidas interações com cada estado do ambiente e como o estado do ambiente é representado por vários fatores, a quantidade de estados cresce exponencialmente de acordo com o número de variáveis de estado. Uma das técnicas para acelerar o processo de aprendizado é a generalização de conhecimento, que visa melhorar o processo de aprendizado, seja no mesmo problema por meio da abstração, ao explorar a similaridade entre estados semelhantes ou em diferentes problemas, ao transferir o conhecimento adquirido de um problema fonte para acelerar a aprendizagem em um problema alvo. Uma abstração considera partes do estado e, ainda que uma única não seja suficiente, é necessário descobrir qual combinação de abstrações pode atingir bons resultados. Nesta dissertação é proposto um método para seleção de abstração, considerando o processo de avaliação da aprendizagem durante o aprendizado. A contribuição é formalizada pela apresentação do algoritmo REPO, utilizado para selecionar e avaliar subconjuntos de abstrações. O algoritmo é iterativo e a cada rodada avalia novos subconjuntos de abstrações, conferindo uma pontuação para cada uma das abstrações existentes no subconjunto e por fim, retorna o subconjunto com as abstrações melhores pontuadas. Experimentos com o simulador de futebol mostram que esse método é efetivo e consegue encontrar um subconjunto com uma quantidade menor de abstrações que represente o problema original, proporcionando melhoria em relação ao desempenho do agente em seu aprendizado

Title in English

Selection of spatial abstraction in Reinforcement Learning by learning process evaluating

Keywords in English

Abstraction selection
Reinforcement Learning
Transfer learning

Abstract in English

Agents that use Reinforcement Learning (RL) techniques seek to solve problems that involve sequential decisions in stochastic environments without a priori knowledge. The learning process developed by the agent in general is slow, since it is done by trial and error and requires repeated iterations with each state of the environment and because the state of the environment is represented by several factors, the number of states grows exponentially according to the number of state variables. One of the techniques to accelerate the learning process is the generalization of knowledge, which aims to improve the learning process, be the same problem through abstraction, explore the similarity between similar states or different problems, transferring the knowledge acquired from A source problem to accelerate learning in a target problem. An abstraction considers parts of the state, and although a single one is not sufficient, it is necessary to find out which combination of abstractions can achieve good results. In this work, a method for abstraction selection is proposed, considering the evaluation process of learning during learning. The contribution is formalized by the presentation of the REPO algorithm, used to select and evaluate subsets of features. The algorithm is iterative and each round evaluates new subsets of features, giving a score for each of the features in the subset, and finally, returns the subset with the most highly punctuated features. Experiments with the soccer simulator show that this method is effective and can find a subset with a smaller number of features that represents the original problem, providing improvement in relation to the performance of the agent in its learning

WARNING - Viewing this document is conditioned on your acceptance of the following terms of use:
This document is only for private use for research and teaching activities. Reproduction for commercial use is forbidden. This rights cover the whole data about this document as well as its contents. Any uses or copies of this document in whole or in part must include the author's name.

Corrigida_Cleiton_Alves.pdf (625.64 Kbytes)

Publishing Date

2018-02-19

Derived works

WARNING: Learn what derived works are clicking here.