Description

The subject of the proposal formally belongs to the intersection between Machine Learning, Control, Optimization and Communication Engineering, and Applied Mathematics. Treated in a unified way, these areas constitute the foundation of Intelligent Networked Cyber-Physical Systems (INCPS) which represent one of the greatest challenges in modern science and technology, with revolutionary impact.

Objectives

The general objective of the project is design and development of advanced methods, algorithms and practical tools for decentralized optimized collaborative deep reinforcement and machine learning for INCPS, applied to two important application use cases involving: a) multi-drone system for environmental monitoring and disasters prevention/relief, and b) smart green buildings. More specifically the objectives are:

1. Design and development of novel decentralized collaborative multi-agent deep reinforcement learning (MADRL) algorithms. New decentralized and distributed algorithms for iterative multi-agent learning of the optimal control policy in complex unknown Markov decision processes will be designed and developed. The high environmental complexity, state-action-space dimensionality, and partial observability problems (typical in modern applications) will be tackled by designing adequate deep neural networks (DNNs) and appropriate parameterization/compression of information to be communicated among the agents within a dynamic consensus-based collaborative scheme. Modern swarm intelligence optimization techniques will be employed in the algorithms’ design and operation.

2. Development of novel swarm intelligence (SI) optimization methods for MADRL and deep machine learning (DML). The objective is to study practical and theoretical aspects of SI algorithms, their application in deep (reinforcement) learning, especially with convolutional neural networks (CNN), and their tuning and adaptation for practical applications. SI algorithms will be applied to hard optimization problems; the study will be conducted in two parts: a) adjusting SI algorithms by applying small changes such as search mechanism modifications and parameters adjustments, or major changes by performing hybridization with other state-of-the-art approaches, and b) using such improved SI algorithms for tuning MADRL and DML hyper-parameters. The quality of the results will be evaluated by performing tests on benchmark problems, datasets as well as on the real-world testbeds.

3. Implementation and Verification on Real-World Systems. The developed algorithms and methods will be verified and implemented in two use-cases involving real-world testbeds. The goal is to illustrate the effectiveness of the theoretical results on a level of proof-of-concept, motivating the further steps towards higher levels of technological and industrial readiness. In particular, the use-cases will involve a multi-drone testbed for environmental monitoring, and a green building testbed.

Work Packages

The project will be implemented through 5 work packages, 2 theoretically oriented, 2 oriented towards the practical use cases, and 1 devoted to management, dissemination and exploitation:

WP1: Decentralized multi-agent reinforcement learning. This WP is oriented directly towards achieving Objective 1 of the project. According to the adopted methodology, novel distributed algorithms with superior properties, applicable to truly decentralized multi-agent settings, for collaborative MADRL in stochastic environments, modelled using the theory of MDPs, will be developed, rigorously theoretically analyzed, and tested using computer simulations.

WP2: Swarm Intelligence Optimization. This WP is oriented directly towards achieving Objective 2 of the project. Based on the adopted methodology, novel SI-based algorithms will be developed and applied to optimizing and tuning DNNs and ultimately to MADRL algorithms designed in WP1, with theoretical as well as computer-simulation based verifications.

WP3: Multi-Drone Use Case. This WP addresses the Objective 3a of the project. It deals with a multi-drone use case, which has a promising potential for commercialization, technology transfer to industrial domain, and high prospective for extensions to the other related applications in the area of INCPS with large impact. It is expected to verify the theoretical results from the previous two WPs, and develop proof-of-concept type results.

WP4: Green Buildings Use Case. This WP addresses the Objective 3b of the project. It deals with a smart green building use case, with a promising potential for commercialization, and technology transfer. It is expected to apply the theoretical results from the WP1 and WP2, and develop proof-of-concept type results.

WP5: Management, Dissemination and Exploitation.