Research article

Intelligent perception communication protocol for satellite constellations via multi-agent deep reinforcement learning with dual-attention and delay-aware coordination

  • Published: 30 June 2026
  • Large-scale low earth orbit (LEO) satellite constellations face three fundamental communication challenges: severe channel dynamics caused by orbital velocities exceeding 7.5 km s$ ^{-1} $, resource scarcity under heterogeneous quality-of-service demands, and strategy incoordination induced by stochastic multi-hop inter-satellite delays. This paper proposes the intelligent perception communication protocol (IPCP), a unified three-layer perception–decision–execution architecture driven by multi-agent deep reinforcement learning (MARL). We develop the information-exchange deep Q-Network (IEDQN), in which each satellite compresses its local observation–Q-value history into a low-dimensional structured token via a shared Info-Block, aggregates global context through a receiver-specific coordination network (Mess-Net), and selects actions via a shared Q-Net on near-fully-observable augmented states. We prove that IEDQN converges almost surely to an $ \epsilon $-approximate Nash equilibrium, with the gap bounded in closed form by the Lipschitz constant of the coordination network. To handle multi-hop delay, we propose the dual-attention delay-aware communication network (DADC), integrating a Transformer-based delay-aware information integration module (DAII) with a dual-attention filtering module (DAIF) that combines Gumbel-softmax hard attention and scaled dot-product soft attention. A formal delay-compensation bound shows that the performance gap to the oracle zero-delay policy decreases exponentially with reception window $ T_r $. Experiments on an NS3-based LEO simulator (5, 7, and 10 agents) show that DADC consistently outperforms the best prior delay-aware method (DACOM) and all delay-agnostic baselines across constellation scales, with advantages that widen as the number of agents increases, while the full IPCP stack doubles throughput (18.4$ \to .8 Mbps) and reduces end-to-end latency by 39.9% over a static protocol baseline.

    Citation: Pan Peng, Wei Tao, Hao-Nan Wang, Hui Zhao. Intelligent perception communication protocol for satellite constellations via multi-agent deep reinforcement learning with dual-attention and delay-aware coordination[J]. Electronic Research Archive, 2026, 34(8): 5546-5579. doi: 10.3934/era.2026248

    Related Papers:

  • Large-scale low earth orbit (LEO) satellite constellations face three fundamental communication challenges: severe channel dynamics caused by orbital velocities exceeding 7.5 km s$ ^{-1} $, resource scarcity under heterogeneous quality-of-service demands, and strategy incoordination induced by stochastic multi-hop inter-satellite delays. This paper proposes the intelligent perception communication protocol (IPCP), a unified three-layer perception–decision–execution architecture driven by multi-agent deep reinforcement learning (MARL). We develop the information-exchange deep Q-Network (IEDQN), in which each satellite compresses its local observation–Q-value history into a low-dimensional structured token via a shared Info-Block, aggregates global context through a receiver-specific coordination network (Mess-Net), and selects actions via a shared Q-Net on near-fully-observable augmented states. We prove that IEDQN converges almost surely to an $ \epsilon $-approximate Nash equilibrium, with the gap bounded in closed form by the Lipschitz constant of the coordination network. To handle multi-hop delay, we propose the dual-attention delay-aware communication network (DADC), integrating a Transformer-based delay-aware information integration module (DAII) with a dual-attention filtering module (DAIF) that combines Gumbel-softmax hard attention and scaled dot-product soft attention. A formal delay-compensation bound shows that the performance gap to the oracle zero-delay policy decreases exponentially with reception window $ T_r $. Experiments on an NS3-based LEO simulator (5, 7, and 10 agents) show that DADC consistently outperforms the best prior delay-aware method (DACOM) and all delay-agnostic baselines across constellation scales, with advantages that widen as the number of agents increases, while the full IPCP stack doubles throughput (18.4$ \to .8 Mbps) and reduces end-to-end latency by 39.9% over a static protocol baseline.



    加载中


    [1] I. Del Portillo, B. G. Cameron, E. F. Crawley, A technical comparison of three low earth orbit satellite constellation systems to provide global broadband, Acta Astronaut., 159 (2019), 123–135. https://doi.org/10.1016/j.actaastro.2019.03.040 doi: 10.1016/j.actaastro.2019.03.040
    [2] D. Bhattacherjee, A. Singla, Network topology design at 27,000 km/hour, in CoNEXT '19: Proceedings of the 15th International Conference on Emerging Networking Experiments And Technologies, (2019), 341–354. https://doi.org/10.1145/3359989.3365407
    [3] O. Kodheli, E. Lagunas, N. Maturo, S. K. Sharma, B. Shankar, J. F. M. Montoya, et al., Satellite communications in the new space era: A survey and future challenges, IEEE Commun. Surv. Tutorials, 23 (2021), 70–109. https://doi.org/10.1109/COMST.2020.3028247 doi: 10.1109/COMST.2020.3028247
    [4] X. You, C. X. Wang, J. Huang, X. Gao, Z. Zhang, M. Wang, et al., Towards 6G wireless communication networks: Vision, enabling technologies, and new paradigm shifts, Sci. China Inf. Sci., 64 (2021), 110301. https://doi.org/10.1007/s11432-020-2955-6 doi: 10.1007/s11432-020-2955-6
    [5] 3GPP, Solutions for NR to Support Non-Terrestrial Networks (NTN), 3rd Generation Partnership Project, 2020. Available from: https://www.3gpp.org/ftp/Specs/archive/38_series/38.821/.
    [6] O. Liberg, S. E. Löwenmark, S. Euler, B. Hofström, T. Khan, X. Lin, et al., Narrowband Internet of Things for non-terrestrial networks, IEEE Commun. Stand. Mag., 4 (2020), 49–55. https://doi.org/10.1109/MCOMSTD.001.2000004 doi: 10.1109/MCOMSTD.001.2000004
    [7] X. Wang, W. Shen, C. Xing, J. An, N. Zhao, L. Hanzo, Joint Bayesian channel estimation and data detection for OTFS systems in LEO satellite communications, IEEE Trans. Commun., 70 (2022), 4386–4399. https://ieeexplore.ieee.org/document/9785832
    [8] M. Handley, Delay is not an option: Low latency routing in space, in HotNets '18: Proceedings of the 17th ACM Workshop on Hot Topics in Networks, (2018), 85–91. https://doi.org/10.1145/3286062.3286075
    [9] L. Zhang, R. Du, Z. Hao, S. Li, Z. Hu, Dynamic packet routing algorithm based on multidimensional information and multiagent reinforcement learning, Int. J. Commun. Syst., 38 (2025), e70039. https://doi.org/10.1002/dac.70039 doi: 10.1002/dac.70039
    [10] S. Zhou, Y. Jia, R. Mao, Z. Nan, Y. Sun, Z. Niu, Task-oriented wireless communications for collaborative perception in intelligent unmanned systems, IEEE Netw., 38 (2024), 21–28. https://doi.org/10.1109/MNET.2024.3414144 doi: 10.1109/MNET.2024.3414144
    [11] X. Ma, H. Zhang, P. Yan, Z. Wu, Y. Ren, Resource allocation optimization in direct satellite-to-device networks based on multi-agent reinforcement learning, IEEE Commun. Mag., 63 (2025), 61–67. https://doi.org/10.1109/MCOM.001.2500178 doi: 10.1109/MCOM.001.2500178
    [12] Y. Wang, K. Gui, Task offloading and computational scheduling in RIS-assisted low Earth orbit satellite communication networks, Veh. Commun., 53 (2025), 100917. https://doi.org/10.1016/j.vehcom.2025.100917 doi: 10.1016/j.vehcom.2025.100917
    [13] T. Yuan, H. M. Chung, J. Yuan, X. Fu, DACOM: Learning delay-aware communication for multi-agent reinforcement learning, in Proceedings of the AAAI Conference on Artificial Intelligence, 37 (2023), 11763–11771. https://doi.org/10.1609/aaai.v37i10.26389
    [14] S. Sukhbaatar, A. Szlam, R. Fergus, Learning multiagent communication with backpropagation, Adv. Neural Inf. Process. Syst., 29 (2016), 2244–2252.
    [15] A. Singh, T. Jain, S. Sukhbaatar, Learning when to communicate at scale in multiagent cooperative and competitive tasks, preprint, arXiv: 1812.09755.
    [16] Y. Niu, R. Paleja, M. Gombolay, Multi-agent graph-attention communication and teaming, in Proceedings of the 20th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2021), (2021), 964–973. https://doi.org/10.5555/3463952.3464065
    [17] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, et al., Human-level control through deep reinforcement learning, Nature, 518 (2015), 529–533. https://doi.org/10.1038/nature14236 doi: 10.1038/nature14236
    [18] J. Huang, Y. Yang, G. He, Y. Xiao, J. Liu, Deep reinforcement learning-based dynamic spectrum access for D2D communication underlay cellular networks, IEEE Commun. Lett., 25 (2021), 2614–2618. https://doi.org/10.1109/LCOMM.2021.3079920 doi: 10.1109/LCOMM.2021.3079920
    [19] R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, I. Mordatch, Multi-agent actor-critic for mixed cooperative-competitive environments, Adv. Neural Inf. Process. Syst., 30 (2017), 6379–6390.
    [20] T. Rashid, M. Samvelyan, C. S. de Witt, G. Farquhar, J. Foerster, S. Whiteson, Monotonic value function factorisation for deep multi-agent reinforcement learning, preprint, arXiv: 1803.11485.
    [21] C. Zhu, M. Dastani, S. Wang, A survey of multi-agent deep reinforcement learning with communication, Auton. Agents Multi-Agent Syst., 38 (2024), 4. https://doi.org/10.1007/s10458-023-09633-6 doi: 10.1007/s10458-023-09633-6
    [22] J. Sheng, X. Wang, B. Jin, J. Yan, W. Li, T. H. Chang, et al., Learning structured communication for multi-agent reinforcement learning, Auton. Agents Multi-Agent Syst., 36 (2022), 50. https://doi.org/10.1007/s10458-022-09580-8 doi: 10.1007/s10458-022-09580-8
    [23] M. Gallici, M. Martin, I. Masmitja, TransfQMix: Transformers for leveraging the graph structure of multi-agent reinforcement learning problems, preprint, arXiv: 2301.05334.
    [24] G. Zheng, N. Wang, R. R. Tafazolli, SDN in space: A virtual data-plane addressing scheme for supporting LEO satellite and terrestrial networks integration, IEEE/ACM Trans. Netw., 32 (2024), 1781–1796. https://doi.org/10.1109/TNET.2023.3330672 doi: 10.1109/TNET.2023.3330672
    [25] C. Zhu, X. Zhu, T. Qin, Joint trajectory and incentive optimization for privacy-preserving UAV crowdsensing via multi-agent federated reinforcement learning, Internet Things, 33 (2025), 101689. https://doi.org/10.1016/j.iot.2025.101689 doi: 10.1016/j.iot.2025.101689
    [26] S. Li, Q. Wu, R. Wang, Dynamic discrete topology design and routing for satellite-terrestrial integrated networks, IEEE/ACM Trans. Netw., 32 (2024), 3840–3853. https://doi.org/10.1109/TNET.2024.3397613 doi: 10.1109/TNET.2024.3397613
    [27] S. Li, Q. Wu, R. Wang, L. Chen, H. Zhang, Efficient multipath differential routing and traffic scheduling in ultra-dense LEO satellite networks: A DRL with Stackelberg game approach, IEEE Trans. Mobile Comput., 24 (2025), 12424–12440. https://doi.org/10.1109/TMC.2025.3586262 doi: 10.1109/TMC.2025.3586262
    [28] S. Li, G. Wu, Q. Wu, R. Wang, H. Zhang, Efficient packet routing for large-scale LEO satellite networks: A Pareto-optimal MARL approach with queueing theory, IEEE Internet Things J., 12 (2025), 46675–46691. https://doi.org/10.1109/JIOT.2025.3610772 doi: 10.1109/JIOT.2025.3610772
    [29] S. Li, Q. Wu, R. Wang, Efficient packet routing in ultra-dense LEO satellite networks via cooperative-MARL with queuing theory model, in 2025 IEEE Wireless Communications and Networking Conference (WCNC), (2025), 1–6. https://doi.org/10.1109/WCNC61545.2025.10978428
    [30] S. Li, Q. Wu, R. Wang, Toward networking and routing in 6G satellite-terrestrial integrated networks: Current issues and a potential solution, IEEE Commun. Mag., 63 (2025), 92–98. https://doi.org/10.1109/MCOM.001.2400197 doi: 10.1109/MCOM.001.2400197
    [31] A. Das, T. Gervet, J. Romoff, D. Batra, D. Parikh, M. Rabbat, et al., TarMAC: Targeted multi-agent communication, in Proceedings of the 36th International Conference on Machine Learning, (2019), 1538–1546.
    [32] Y. Liu, W. Wang, Y. Hu, J. Hao, X. Chen, Y. Gao, Multi-agent game abstraction via graph attention neural network, in Proceedings of the AAAI Conference on Artificial Intelligence, 34 (2020), 7211–7218. https://doi.org/10.1609/aaai.v34i05.6211
    [33] P. Sunehag, G. Lever, A. Gruslys, W. M. Czarnecki, V. Zambaldi, M. Jaderberg, et al., Value-decomposition networks for cooperative multi-agent learning, preprint, arXiv: 1706.05296.
    [34] K. Son, D. Kim, W. J. Kang, D. E. Hostallero, Y. Yi, QTRAN: Learning to factorize with transformation for cooperative multi-agent reinforcement learning, in Proceedings of the 36th International Conference on Machine Learning, (2019), 5887–5896.
    [35] X. Hao, H. Mao, W. Wang, Y. Yang, D. Li, Y. Zheng, et al., Breaking the curse of dimensionality in multiagent state space: A unified agent permutation framework, preprint, arXiv: 2203.05285.
    [36] E. Fard, R. R. Selmic, Time-delayed data transmission in heterogeneous multi-agent deep reinforcement learning system, in 2022 30th Mediterranean Conference on Control and Automation (MED), (2022), 636–642. doilinkhttps://doi.org/10.1109/MED54222.2022.9837194
    [37] Y. Xia, Y. Xu, Y. Wang, S. Mondal, S. Dasgupta, A. K. Gupta, Optimal secondary control of islanded AC microgrids with communication time-delay based on multi-agent deep reinforcement learning, CSEE J. Power Energy Syst., 9 (2022), 1301–1311. doilinkhttps://doi.org/10.17775/CSEEJPES.2021.08680 doi: 10.17775/CSEEJPES.2021.08680
    [38] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, et al., Attention is all you need, in Advances in Neural Information Processing Systems 30 (NIPS 2017), (2017), 5998–6008.
    [39] J. Tsitsiklis, B. Van Roy, Analysis of temporal-difference learning with function approximation, in Advances in Neural Information Processing Systems 9 (NIPS 1996), 1996.
    [40] V. S. Borkar, Stochastic Approximation: A Dynamical Systems Viewpoint, Cambridge University Press, Cambridge, UK, 2008.
    [41] ARM Holdings, Cortex-A53 Processor, ARM Developer Documentation, 2012. Available from: https://www.arm.com/products/silicon-ip-cpu/cortex-a/cortex-a53.
  • Reader Comments
  • © 2026 the Author(s), licensee AIMS Press. This is an open access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0)
通讯作者: 陈斌, bchen63@163.com
  • 1. 

    沈阳化工大学材料科学与工程学院 沈阳 110142

  1. 本站搜索
  2. 百度学术搜索
  3. 万方数据库搜索
  4. CNKI搜索

Metrics

Article views(189) PDF downloads(8) Cited by(0)

Article outline

Figures and Tables

Figures(12)  /  Tables(13)

Other Articles By Authors

/

DownLoad:  Full-Size Img  PowerPoint
Return
Return

Catalog