Authors
Raúl Arranz, Juan A Besada, David Carramiñana
Published in
Sensors (Basel, Switzerland). Volume 26. Issue 14. Jul 10, 2026. Epub Jul 10, 2026.
Abstract
Reinforcement learning (RL) has emerged as a powerful paradigm for enabling autonomous coordination in multi-UAV systems operating in complex and uncertain environments. However, the effectiveness of learned policies is strongly influenced by how actions are implemented at the control level, an aspect that has received limited attention in the literature. This paper presents a comparative study of three control methods (heading-based, waypoint-based, and deterministic) within a unified hybrid-AI architecture, in which the same RL policy structure is used across two of the three configurations. By isolating the control method as the sole variable, the study evaluates how different action abstractions affect learning efficiency, robustness, and operational performance in cooperative surveillance missions. A statistically rigorous Monte Carlo evaluation, supported by non-parametric hypothesis testing, demonstrates that heading-based control consistently achieves superior performance in terms of revisit period, target acquisition time, and tracking continuity. The analysis further reveals that these gains arise from improved reactivity and constraint handling rather than from differences in policy learning. The results highlight the critical role of control-level design in RL-based multi-agent systems and provide practical guidelines for selecting action abstractions in aerial swarm applications.
PMID:
42515279
Bibliographic data and abstract were imported from PubMed on 28 Jul 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 7
- Comments 0