ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ

DOI: https://doi.org/10.15407/itm2025.04.077 The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric uncertainties. The goal is to develop and assess an app...

Ausführliche Beschreibung

Gespeichert in:
Bibliographische Detailangaben
Datum:2025
Hauptverfasser: SOROCHINSKII, V. V., KHOROSHYLOV, S. V., LEVCHUK, I. L., DUBOVYK, T. M., HUZ, H. M., ROMANCHUK, O. O.
Format: Artikel
Sprache:Englisch
Veröffentlicht: текст 3 2025
Schlagworte:
Online Zugang:https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
Назва журналу:Technical Mechanics
Завантажити файл: Pdf

Institution

Technical Mechanics
_version_ 1870649983994167296
author SOROCHINSKII, V. V.
KHOROSHYLOV, S. V.
LEVCHUK, I. L.
DUBOVYK, T. M.
HUZ, H. M.
ROMANCHUK, O. O.
author_facet SOROCHINSKII, V. V.
KHOROSHYLOV, S. V.
LEVCHUK, I. L.
DUBOVYK, T. M.
HUZ, H. M.
ROMANCHUK, O. O.
author_institution_txt_mv [ { "author": "V. V. SOROCHINSKII", "institution": "Institute of Technical Mechanics of the National Academy of Science of Ukraine and the State Space Agency of Ukraine, 15 Leshko-Popel St., Dnipro 49005, Ukraine; e-mail: vovas99@ukr.net" }, { "author": "S. V. KHOROSHYLOV", "institution": "Institute of Technical Mechanics of the National Academy of Science of Ukraine and the State Space Agency of Ukraine, 15 Leshko-Popel St., Dnipro 49005, Ukraine" }, { "author": "I. L. LEVCHUK", "institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine" }, { "author": "T. M. DUBOVYK", "institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine" }, { "author": "H. M. HUZ", "institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine" }, { "author": "O. O. ROMANCHUK", "institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine" } ]
author_sort SOROCHINSKII, V. V.
baseUrl_str https://journal-itm.dp.ua/ojs/index.php/ITM_j1/oai
collection OJS
datestamp_date 2026-07-13T20:26:18Z
description DOI: https://doi.org/10.15407/itm2025.04.077 The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric uncertainties. The goal is to develop and assess an approach that combines sliding-mode control with modern reinforcement-learning methods tailored for resource-constrained onboard implementation. Relative motion dynamics is formulated in an orbital coordinate frame with normalized states and discretized in time. Binary actions with pulse-width modulation, subject to constraints on the thrust level, pulse duration, and duty cycle, represent the impulsive nature of actuation. We propose a combined synthesis in which the sliding-surface parameters and switching rules are tuned via proximal policy optimization within an actor-critic architecture. The actor and critic are implemented as neural networks that approximate the policy and the value function, respectively. The actor neural network takes the state vector as input information and outputs the mean and standard deviation of the parameters of the sliding mode control law. The value function penalizes both the state error and control effort, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. Two uncoupled agents are designed to control spacecraft relative orbital motion in in-plane and out-of-plane directions independently. The proximal policy optimization hyperparameters are selected to ensure a trade-off among the learning time, stability, and control performance. The reinforcement-learning agents are trained and analyzed considering four cases that differ in the thrust levels and weighting matrices. The quality functional combines state deviation and thrust use penalties, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. The results confirm the potential of this approach for autonomous spacecraft control under constraints and uncertainty. Compared with reported baselines, the trained agent shows superior robustness to plant-parameter uncertainty, which we attribute to the inherent robust properties of sliding-mode control. These findings have the potential to improve the efficiency and autonomy of on-orbit servicing operations. REFERENCES 1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris Management. arXiv:1901.11121, 2019. 2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments. Progress in Aerospace Sciences. 2019. V. 108. Pp. 32-120.https://doi.org/10.1016/j.paerosci.2019.01.004 3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to spacecraft attitude controller design. Automatika. 2016. V. 57. Pp. 212-220.https://doi.org/10.7305/automatika.2016.07.618 4. Anthony T., Wie B., Carroll S. Pulse-modulated control synthesis for a flexible spacecraft. Journal of Guidance. 1989. V. 13. No. 6. Pp. 1014-1022. https://doi.org/10.2514/3.20574 5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an algorithm to control the angular motion of spacecraft equipped with an aeromagnetic deorbiting system. Eastern-European Journal of Enterprise Technologies. 2020. V. 1. No. 5. Pp. 37-46. https://doi.org/10.15587/1729-4061.2020.192813 6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016. 7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks. Communications of the ACM. 2017. V. 60. No. 6. Pp. 84-90. https://doi.org/10.1145/3065386 8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017. V. 31. No. 16. Pp. 821-835. https://doi.org/10.1080/01691864.2017.1365009 9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving. Electronic Imaging. 2017. V. 19. Pp. 70-76.https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023 10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017. V. 550. Pp. 354-359. https://doi.org/10.1038/nature24270 11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics. 2019. No. 3. Pp. 287-299. https://doi.org/10.1007/s42064-018-0053-6 12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and Technology. 2021. V. 27. No. 6. Pp. 38-52. https://doi.org/10.15407/knit2021.06.038 13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. V. 18. No. 7.  https://doi.org/10.2514/1.I010914 14. Gaudet B., Linares R., Furfaro R. Six degree-of-freedom hovering using LIDAR altimetry via reinforcement meta-learning. Acta Astronautica. 2020. V. 172. Pp. 90-99.https://doi.org/10.1016/j.actaastro.2020.03.026 15. Gaudet B., Linares R., Furfaro R. Seeker based adaptive guidance via reinforcement meta-learning applied to asteroid close proximity operations. Acta Astronautica. 2020. V. 171. Pp. 1-13.https://doi.org/10.1016/j.actaastro.2020.02.036 16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object via deep learning. Space Science and Technology. 2022. V. 28. No. 5. Pp. 15-26.https://doi.org/10.15407/knit2022.05.015 17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and Technology. 2024. V. 30. No. 2. Pp. 3-14. https://doi.org/10.15407/knit2024.02.003 18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. Innov. 2018. V. 14. No. 4. Pp. 5-16. https://doi.org/10.15407/scine14.04.005 19. Steinberger M., Horn M., Fridman L. (Eds). Variable-Structure Systems and Sliding-Mode Control. Springer-Verlag, 2020. (Studies in Systems, Decision and Control; Vol. 271).https://doi.org/10.1007/978-3-030-36621-6 20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Hemisphere Publishing, 1975. Pp. 224-235. 21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. Pp. 47-65. 22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. 13 pp. 23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep Reinforcement Learning. arXiv:1602.01783, 2016.
first_indexed 2025-12-17T12:05:41Z
format Article
fulltext 77 Автоматизація, комп’ютерно-інтегровані технології Automation, computer-integrated technologies UDC 629.5 https://doi.org/10.15407/itm2025.04.077 V. V. SOROCHINSKII1, S. V. KHOROSHYLOV1, I. L. LEVCHUK2, T. M. DUBOVYK2, H. M. HUZ2, O. O. ROMANCHUK2 ON-OFF SPACECRAFT RELATIVE CONTROL IN SLIDING MODE VIA REINFORCEMENT LEARNING 1Institute of Technical Mechanics of the National Academy of Science of Ukraine and the State Space Agency of Ukraine, 15 Leshko-Popel Str., Dnipro, 49005, Ukraine; e-mail: vovas99@ukr.net 2Ukrainian State University of Science and Technologies,2 Lazariana St., Dnipro, 49010, Ukraine Розглянуто задачу відносного імпульсного керування рухом космічного апарата у ковзному режимі для автономних орбітальних сервісних операцій за наявності обмежень на амплітуду керуючих впливів, дискретності дій та параметричних невизначеностей. Метою роботи є розробка й оцінювання підходу, що поєднує принципи ковзного керування з сучасними методами навчання з підкріпленням, орієнтованими на бортову реалізацію з обмеженими ресурсами. Динаміку відносного руху задано в орбітальній системі координат у нормалізованих змінних і дискредитовано. Імпульсний характер впливів виконавчих органів відображено через бінарні дії з широтно-імпульсною модуляцією та обмеженнями на рівень тяги, тривалість і період увімкнень. Запропоновано комбінований синтез, у якому параметри поверхні ковзання та правила перемикання налаштовуються методом проксимальної оптимізації політики з використанням архітектури актор-критик. Актор і критик реалізовані у вигляді нейронних мереж, які відповідно апроксимують політику та функцію цінності. Нейронна мережа актора приймає вектор стану як вхідну інформацію і видає середнє значення та стандартне відхилення параметрів закону керування у ковзному режимі. Функція цінності штрафує як за помилку стану, так і за витрати на керування, що дозволяє забезпечити компроміс між швидкістю реагування, точністю та витратою палива. Два незалежні агенти розроблені для керування відносним орбітальним рухом космічного апарата окремо в напрямку площини орбіти та у перпендикулярному напрямку. Гіперпараметри оптимізації проксимальної політики обрано для забезпечення компромісу між часом навчання, стабільністю та якістю керування. Агенти навчання з підкріпленням навчeні та проаналізовані з урахуванням чотирьох випадків, що відрізняються рівнями тяги та ваговими матрицями. Функціонал якості об’єднує штрафи за відхилення стану та використання тяги, що дає змогу знаходити компроміс між швидкодією, точністю та витратами робочого тіла. Отримані результати підтверджують потенціал такого підходу для задач автономного керування космічних апаратів в умовах обмежень та невизначеності. У порівнянні з відомими результатами навчений агент продемонстрував кращу робастність по відношенню до невизначенності параметрів моделі об’єкта керування, що пояснюється сильними робастними властивостями керування в ковзному режимі. Отримані результати мають потенціал підвищити ефективність та автономність орбітальних сервісних операцій. Ключові слова: навчання з підкріпленням, проксимальна оптимізація політики, керування космічним апаратом, орбітальні сервісні операції, on-off керування, автономні системи керування. The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric uncertainties. The goal is to develop and assess an approach that combines sliding-mode control with modern reinforcement-learning methods tailored for resource-constrained onboard implementation. Relative motion dynamics is formulated in an orbital coordinate frame with normalized states and discretized in time. Binary actions with pulse-width modulation, subject to constraints on the thrust level, pulse duration, and duty cycle, represent the impulsive nature of actuation. We propose a combined synthesis in which the sliding-surface parameters and switching rules are tuned via proximal policy optimization within an actor-critic architecture. The actor and critic are implemented as neural networks that approximate the policy and the value function, respectively. The actor neural network takes the state vector as input information and outputs the mean and standard deviation of the parameters of the sliding mode control law. The value function penalizes both the state error and control effort, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. Two uncoupled agents are designed to control spacecraft relative orbital motion in in-plane and out-of-plane directions independently. The proximal policy optimization hyperparameters are selected to ensure a trade-off among the learning time, stability, and control performance. The reinforcement-learning agents are trained and analyzed considering four cases that differ in the thrust levels and weighting matrices. The quality functional combines state deviation and thrust use penalties, thus enabling a trade-off among the © V. V. Sorochinskii, S. V. Khoroshylov, I. L. Levchuk, T. M. Dubovyk, H. M. Huz, O. O. Romanchuk, 2025 Техн. механіка. – 2025. – № 4. https://doi.org/10.15407/itm2025.04.0 mailto:vovas99@ukr.net 78 response speed, accuracy, and propellant consumption. The results confirm the potential of this approach for autonomous spacecraft control under constraints and uncertainty. Compared with reported baselines, the trained agent shows superior robustness to plant-parameter uncertainty, which we attribute to the inherent robust properties of sliding-mode control. These findings have the potential to improve the efficiency and autonomy of on-orbit servicing operations. Keywords: reinforcement learning, proximal policy optimization, spacecraft control, on-orbit servicing, on– off control, autonomous control systems. 1. Introduction Modern space missions increasingly require complex maneuvers with minimum human involvement. Typical tasks include on-orbit servicing operations such as docking, towing, and active debris removal [1, 2]. To perform these operations, a servicing spacecraft (SSC) must maneuver close to the orbital object (OO), addressing the challenge of relative guidance and control. Control inputs are commonly generated by reaction thrusters (RTs). Unlike reaction wheels or control-moment gyros (CMGs), thruster outputs are typically binary (“on/off”). This switching behavior makes actuation highly nonlinear and complicates direct controller design. To implement continuous-time control laws, pulse-width modulation (PWM) and pulse-frequency modulation (PFM) are used [3, 4], transforming the continuous control signal into a sequence of pulses with specified duration or frequency. The accuracy of this approximation depends on the modulator resolution: coarse quantization significantly reduces control performance [5], so the regulator and modulator parameters must be co-designed, which is a complex task. The impressive results achieved using deep learning (DL) methods [6–10] have recently intensified interest in AI approaches among researchers and practitioners worldwide. Recently, these methods have also been applied to space problems [11–16]. Publication [17] demonstrated the feasibility of directly synthesizing control laws via reinforcement learning (RL), mapping the state vector to binary on/off thruster commands. That approach outperformed a linear regulator in conjunction with a PWM in terms of accuracy, response speed, and the number of thruster firings. At the same time, it requires many training episodes and does not provide strict robustness guarantees when training and operational conditions differ. To mitigate these drawbacks, this study proposes combining sliding-mode control (SMC) with RL. This study aims to assess the feasibility of on/off spacecraft relative control in sliding mode that is implemented via reinforcement learning, and to specify the features of this approach 2. Problem Statement 2.1 Model of spacecraft relative dynamics. To describe the motion of the servicing spacecraft (SSC) relative to an orbital object (OO), we use the orbital coordinate frame (OCF). The OCF origin is fixed at the SSC center of mass (Fig. 1). The 𝑥-axis is directed along the geocentric radius vector from Earth’s center to the SSC (𝑖)̂. The 𝑧-axis is normal to the plane spanned by the 𝑥-axis and the SSC orbital-velocity vector (𝑘̂) and is oriented toward the positive orbital Fig. 1 – Axes directions of the OCF 79 angular momentum. The 𝑦-axis completes a right-handed triad (𝑗̂). The relative dynamics of the SSC–OO system can be written as the following linearized system of equations [18]: 𝑥̈ − 𝜔2𝑥 − 2𝜔𝑦̇ − 𝜔̇𝑦 − 𝑘𝑥 = 𝑓𝑥 𝑑 𝑚𝑑 − 𝑓𝑥 𝑠 𝑚𝑠, (1) 𝑦 − 𝜔2𝑦 + 2𝜔𝑥̇ + 𝜔̇𝑥 + 𝑘𝑦 = 𝑓𝑦 𝑑 𝑚𝑑 − 𝑓𝑦 𝑠 𝑚𝑠, (2) 𝑧 + 𝑘𝑧 = 𝑓𝑧 𝑑 𝑚𝑑 − 𝑓𝑧 𝑠 𝑚𝑠, (3) where 𝑥 , 𝑦 , 𝑧 are the OCF projections of the relative position vector 𝑟; 𝑚𝑠 and 𝑚𝑑 are the masses of the SSC and OO, respectively; 𝑓𝑥 𝑑, 𝑓𝑦 𝑑, 𝑓𝑧 𝑑 are the OCF projections of the total force 𝑓𝑑, acting on the OO; 𝑓𝑥 𝑠, 𝑓𝑦 𝑠, 𝑓𝑧 𝑠 are the OCF projections of the total force 𝑓𝑠, acting on the SSC. The total force 𝑓𝑠 includes both control input and external disturbances acting on the SSC. The disturbances 𝑓𝑑 and 𝑓𝑠 may also account for the J2 perturbation, third-body gravity (Sun and Moon), aerodynamic drag, and solar radiation pressure. The parameters in Eqs (1) – (3) are defined as follows: 𝜔 = √ 𝜇 𝑝3 (1 + 𝜀 cos 𝑣)𝜔, 𝜔̇ = −2𝜀√ 𝜇 𝑝3 sin𝑣 (1 + 𝜀 cos 𝑣)𝜔, 𝑝 = 𝑎(1 − 𝜀2), 𝑘 = 𝜇 𝑟3, 𝑟 = 𝑎(1−𝜀2) 1+𝜀 cos𝑣 , where 𝜇 is the Earth's gravitational constant, 𝜀 is the orbital eccentricity, 𝜈 is the true anomaly, 𝑎 is the semi-major axis, and 𝑟 is the geocentric distance. Equations (1) – (2) describe in-plane orbital motion, while Eq (3) describes out-of-plane orbital motion. Neglecting the external disturbances and introducing the state vectors 𝑋𝑖𝑛 = [𝑥, 𝑦, 𝑥̇, 𝑦̇ ]𝑇, 𝑋𝑜𝑢𝑡 = [𝑧, 𝑧̇ ]𝑇 and the control vectors 𝑈𝑖𝑛 = [𝑢𝑥, 𝑢𝑦] 𝑇 , 𝑈𝑜𝑢𝑡 = 𝑢𝑧 , the model (1) can be written in the standard linear state-space form: 𝑋̇𝑖𝑛 = 𝐴𝑖𝑛𝑋𝑖𝑛 + 𝐵𝑖𝑛𝑈𝑖𝑛, 𝑋̇𝑜𝑢𝑡 = 𝐴𝑜𝑢𝑡𝑋𝑜𝑢𝑡 + 𝐵𝑜𝑢𝑡𝑈𝑜𝑢𝑡 , (4) where 𝐴𝑖𝑛 = [ 0 0 1 0 0 0 0 1 𝜔2 + 2𝑘 𝜔̇ 0 2𝜔 −𝜔̇ 𝜔2 − 𝑘 −2𝜔 0 ], 𝐵𝑖𝑛 = [ 0 0 0 0 − 1 𝑚𝑠 0 0 − 1 𝑚𝑠] , 𝐴𝑜𝑢𝑡 = [ 0 1 −𝑘 0 ], 𝐵𝑜𝑢𝑡 = [ 0 − 1 𝑚𝑠 ]. The magnitudes of individual state components differ significantly, which may complicate neural-network (NN) training. To mitigate this issue, the state vector is normalized as follows: 𝑋𝑖𝑛 𝑛 = [ 𝑥 𝑥𝑚 , 𝑦 𝑦𝑚 , 𝑥̇ 𝑥̇𝑚 , 𝑦̇ 𝑦̇𝑚 ] 𝑇 , 𝑋𝑜𝑢𝑡 𝑛 = [ 𝑧 𝑧𝑚 , 𝑧̇ 𝑧̇𝑚 ] 𝑇 , (5) where 𝑥𝑚, 𝑦𝑚, 𝑧𝑚, 𝑥̇𝑚, 𝑦̇𝑚, 𝑧̇𝑚 are the maximum magnitudes of the corresponding state components. For the normalized state, the dynamic model can be given as 𝑋̇𝑛 = 𝐴𝑛𝑋𝑛 + 𝐵𝑛𝑈, 80 where 𝐴𝑛 = 𝑁−1𝐴𝑁, 𝐵𝑛 = 𝑁−1𝐵, 𝑁 = diag(𝑥𝑚, 𝑦𝑚, 𝑧𝑚, 𝑥̇𝑚, 𝑦̇𝑚, 𝑧̇𝑚). Since the onboard controller is implemented using a digital computer, we use the discrete-time form of the model: 𝑋𝑘+1 = 𝐴𝑘𝑋𝑘 + 𝐵𝑘𝑈𝑘, (6) where 𝐴𝑘 = (𝐼 + 𝐴𝑛𝑇); 𝐵𝑘 = 𝐵𝑛𝑇; 𝑇 is the sampling period, and 𝑘 is the time index. We also assume that the full state vector is measurable and that these measurements are not corrupted by noise. 2.2. Sliding-Mode Control Sliding-mode control is a variable-structure control in which the input is deliberately shaped to cause the trajectory to reach a sliding surface and then move along it [19]. The control law is not a continuous function of time, because the control structure switches depending on the current location in the state space. The switching nature of SMC suits well naturally to the task of spacecraft relative-motion control because reaction thrusters operate in the on/off mode. Motion along the sliding hypersurface is described by a reduced-order equation and can be compared to the evolution along eigenmodes of linear systems. To keep the system sliding requires a high-gain switching feedback. The SMC design includes the following steps: selecting a sliding hypersurface 𝑠(𝑥) = 0 such that motion along it yields the desired system behavior; synthesizing a feedback law with sufficient gains that guarantees both reaching the surface and its invariance - i.e., the trajectory intersects the surface and then remains on it. The switching function 𝑠(𝑥) is defined as a measure of the deviation of the state 𝑥 from the sliding hypersurface: the state belongs to the hypersurface when 𝑠(𝑥) = 0 and out of it when 𝑠(𝑥) ≠ 0. The control has a switching structure and changes sign according to the sign of the s(x), e.g., 𝑢 = 𝑢+ when 𝑠(𝑥) > 0 and 𝑢 = 𝑢−), thus ensuring reaching and subsequent sliding along the surface. Within the SMC paradigm, the control action is always directed to reduce the magnitude of |𝑠(𝑥)| and to “push” the trajectory toward the sliding surface. When the surface is reached, the trajectory slides along it and, depending on how the surface is chosen, moves toward the target, typically an equilibrium (e.g., the origin). The switching function is defined as follows: 𝜎(𝑋) = 𝑆𝑇𝑋, (7) Let the sliding surface be a simplex defined by 𝜎(𝑋) = 0. Then, for the stability analysis, we consider the following Lyapunov candidate: 𝑉(𝜎(𝑋)) = 0.5𝜎(𝑋)𝑇𝜎(𝑋) with the time derivative 𝑉̇(𝜎(𝑋)) = 0.5𝜎(𝑋)𝑇𝜎̇(𝑋). The control is chosen so that 𝜎̇(𝑋) < 0 when 𝜎(𝑋) > 0 and 𝜎̇(𝑋) > 0 when 𝜎(𝑋) < 0 to ensure that 𝑉̇(𝜎(𝑋)) < 0. The evolution of the switching function can be given as 𝜎̇(𝑋) = 𝑆𝑇(𝐴𝑋 + 𝐵𝑈). 81 The control input is formed as 𝑈 = (𝑆𝑇𝐵)−1 (−𝑆𝑇𝐴𝑋 + ℎ(𝜎(𝑋))), where ℎ(𝜎(𝑋)) = −𝐾𝑖 𝜎(𝑋) |𝜎(𝑋)|+𝜀 , and 𝜀 sets a dead zone that prevents excessive control chattering. Since the actuators operate in on/off mode, the SMC output is approximated by a sequence of pulses of variable duration using a PWM. The pulse duration within each sampling period is determined as follows: 𝑡𝑓 = 𝑢𝑘 𝑢𝑓 𝑇, 𝑡𝑓 ≤ 𝑇, where 𝑢𝑓 is the nominal thrust level. As the control performance metric, we use the following quadratic cost function [20]: 𝐽 = ∑ (𝑋𝑘 𝑇𝑄𝑋𝑘 + 𝑈𝑘 𝑇𝑅𝑈𝑘)𝑁 𝑘=0 , (8) where Q and R are the weighting matrices penalizing the states 𝑋𝑘 and the control 𝑈𝑘 , respectively. To obtain an optimal sliding-mode controller, it is necessary to determine the parameters 𝐾, 𝑆, and 𝜀 that minimize the quadratic cost (8). This optimization task is addressed via reinforcement learning. 3. Reinforcement-Learning-Based Control In the RL paradigm, the controller learns by analyzing the consequences of its own actions [21]. These consequences are evaluated by a scalar reinforcement signal received from the plant with which the controller interacts. The reinforcement can be interpreted as a criterion that enables an intelligent system to adjust its control policy to achieve a long-term objective. The generic RL loop (Fig. 2) includes the following steps: 1. At the time 𝑡𝑘 the plant is in the state 𝑋𝑘. In this state, the controller (CS) selects an action 𝑈𝑘 from the available action set. 2. The chosen action is applied, causing a transition to a new state 𝑋𝑘+1, and the system receives a reinforcement signal 𝐶𝑘. 3. The algorithm repeats from step 2, taking into account the received reinforcement; if the new state is terminal, the episode ends. Let 𝜒 denote the state space and A the action space. The reinforcement 𝐶𝑘 results from applying the action 𝑈𝑘 in state 𝑋𝑘. Formally, it is a function on 𝜒 × 𝐴: 𝐶𝑘 = 𝐶(𝑋𝑘 , 𝑈𝑘) Fig. 2 – RL setup The CS selects actions to minimize the return, which is the discounted sum of instantaneous costs: 𝐺𝑘 = 𝐶𝑘 + 𝛾𝐶𝑘+1 + 𝛾2𝐶𝑘+2+. . .= ∑ 𝛾𝑖∞ 𝑖=0 𝐶𝑘+𝑖,0 ≤ 𝛾 ≤ 1, where 𝛾 is the discount factor and 𝐶𝑘 is the instantaneous cost at step 𝑘. 82 The factor 𝛾 determines how strongly future costs influence current action selection. A key element of RL is the value function. Let the controller apply actions according to a policy 𝜋: 𝑈𝑘 = 𝜋(𝑋𝑘). Then the value function 𝑉𝜋(𝑋𝑘) gives the total discounted cumulative cost when starting from the state 𝑋𝑘 and following the policy 𝜋: 𝑉𝜋(𝑋𝑘) = ∑ 𝛾𝑘∞ 𝑖=0 𝐶𝑘+𝑖(𝑋𝑘+𝑖, 𝑈𝑘+𝑖) = 𝐶𝑘(𝑋𝑘, 𝑈𝑘) + 𝛾𝑉𝜋(𝑋𝑘+1). RL can be implemented using an actor-critic architecture. In this case, the critic estimates the value function for each state, while the actor maps states to control actions. In this study, the actor and critic are implemented as neural networks that approximate, respectively, the policy and the value function: 𝑉𝜋(𝑋𝑘 , 𝜙), 𝜋(𝑋𝑘 , 𝜙 ), where 𝜃 and 𝜙 are the parameter vectors of the critic and actor, respectively. Proximal Policy Optimization (PPO) [22] is chosen among different RL algorithms because it provides a tradeoff between implementation complexity and training stability. This includes the following steps: Cumulative cost estimation, which is the sum of the cost for this time step and the discounted future cost [23]: 𝐺𝑡 = ∑ (𝛾𝑘−𝑡𝐶𝑘) + 𝑏𝛾𝑁−𝑡+1𝑉(𝑋𝑡𝑠+𝑁 , 𝜃)𝑡𝑠+𝑚 𝑘=𝑡 , where 𝑏 = 0 if 𝑋𝑡𝑠+𝑁 is terminal and 𝑏 = 1 otherwise (i.e., when the rollout segment does not end at a terminal state, a bootstrap term from the critic 𝑉(⋅; 𝜃) is added); 𝛾 ∈ [0,1] is the discount factor. 1. Advantage estimation: 𝐷𝑡 = 𝐺𝑡 − 𝑉(𝑋𝑡 , 𝜃), 2. Critic update via value loss minimization 𝐿𝑐𝑟𝑖𝑡𝑖𝑐(𝜃) = 1 𝑀 ∑ (𝐺𝑖 − 𝑉(𝑋𝑖, 𝜃)) 2𝑀 𝑖=1 , 3. Actor update using clipped policy objective 𝐿𝑎𝑐𝑡𝑜𝑟(𝜙) = 1 𝑀 ∑ (−𝑚𝑖𝑛(𝑟𝑖(𝜙) ∙ 𝐷𝑖, 𝑐𝑖(𝜙) ∙ 𝐷𝑖)) 𝑀 𝑖=1 , 𝑟𝑖(𝜙) = 𝜋(𝑈𝑖|𝑋𝑖,𝜙) 𝜋(𝑈𝑖|𝑋𝑖,𝜙𝑜𝑙𝑑) , 𝑐𝑖(𝜙) = 𝑚𝑎𝑥(𝑚𝑖𝑛(𝑟𝑖(𝜙), 1 + 𝜀), 1 − 𝜀) , where 𝐷𝑖 is the advantage for the 𝑖-th sample, 𝐺𝑖 is the corresponding return, 𝜋(𝑈𝑖|𝑋𝑖 , 𝜙) is the probability of action 𝑈𝑖 in state 𝑋𝑖 under the current policy parameters 𝜙, 𝜋(𝑈𝑖|𝑋𝑖, 𝜙𝑜𝑙𝑑) is the probability under the pre-update policy, 𝜀 is the clip factor, 𝑀 is the minibatch size. The NN architecture of the intelligent agent is given in Tables 1 and 2. Almost all network layers have the tanh activation, and only the actor’s output layer employs SoftPlus function to provide positive outputs. The actor NN takes the state vector as input and produces the mean and standard deviation of the parameters 𝑆, 𝐾, and 𝜀. The structure of this NN is shown in Fig. 3. In conjunction with this actor, we use the same NN critic and the same quadratic performance index as in Ref [17]. 83 Fig. 3 – Architecture of the actor neural network Table 1. Number of neurons in the actor layers Case Input fc hidden fc_1/fc_4 hidden fc_2/fc_5 hidden fc_3/fc_6 hidden Output in-plain 𝑛𝑖𝑛 10𝑛𝑖𝑛 10𝑛𝑖𝑛 25 30 3 out-of- plain 𝑛𝑜𝑢𝑡 10𝑛𝑜𝑢𝑡 10𝑛𝑜𝑢𝑡 50 60 6 Table 2. Number of neurons in the critic layers Case Input 1-st hidden 2-nd hidden 3-d hidden Output in-plain 𝑛𝑖𝑛 10𝑛𝑖𝑛 √100𝑛𝑖𝑛 10 1 out-of- plain 𝑛𝑜𝑢𝑡 10𝑛𝑜𝑢𝑡 √100𝑛𝑜𝑢𝑡 10 1 4. Numerical results For training and evaluation of the intelligent agent, the following system parameters were used: a=7017 km; 𝑚𝑠 = 500 kg; 𝑚𝑑 = 1575 kg; T=200 sec; 𝑇𝑓 = 10 sec. Maximum magnitudes of the state-vector components: 𝑥𝑚 = 800 m, 𝑦𝑚 = 800 m, 𝑥̇𝑚=2 m/s, 𝑦̇𝑚=2 m/s. We analyzed four cases that differ in the thrust levels and weighting matrices (Table 3). Table 3. Values of thrust levels and weighting matrices in-plain out-of-plain 𝒖𝒇 Q R 𝒖𝒇 Q R Case 1 8 [1 0; 0 0.001] ×1e-7 1e-3 8 [1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0 0 0 0.001]×1e-7 [1e-3 0; 0 1e-3] Case 2 4 [1 0; 0 0.001]×1e-7 1e-3 4 [1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0 0 0 0.001]×1e-7 [1e-3 0; 0 1e-3] Case 3 8 [1 0; 0 0.001]×1e-7 1e-4 8 [1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0 0 0 0.001]×1e-7 [1e-3 0; 0 1e-4] Case 4 4 [1 0; 0 0.001]×1e-7 1e-4 4 [1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0 0 0 0.001]×1e-7 [1e-3 0; 0 1e-4] 84 The reward function is designed as the weighted sum of state errors and control efforts: 𝑟𝑡 = −(𝑥𝑡 𝑇𝑄𝑥𝑡 + 𝑢𝑡 𝑇𝑅𝑢𝑡). (10) Training was conducted in episodes, each running until a maximum step count was reached. The PPO hyperparameters were selected to ensure stable learning as follows: 𝛾 = 0.98, 𝜀 = 0.2. An experience horizon of 50 and a minibatch size of 50 were used for updating the NN parameters. Figure 3 demonstrates system stability since all trajectories of 𝑥5, 𝑥6 converge to zero. The cases differ in how quickly the spacecraft reaches the target state and in the magnitude of residual errors. Lower thrust leads to a longer transition and more frequent small control actions, when the agent corrects the state more frequently using small pulses. After settling, the residual error is close to zero in all cases. The small remaining oscillations are due to the on-off control actions. Case 1 Case 2 Case 3 Case 4 Fig. 3 – Variation of the out-of-plane state In Fig. 4, the actor NN output signal is plotted by a blue line and its on-off realization is shown by a red one. Initially, sequences of saturated same-sign pulses are applied to reduce the error rapidly. After that, the pulses become sparser and shorter, occasionally changing sign for fine corrections. In cases when higher thrust is available, fewer control pulses are applied than in cases with lower thrust. The discrepancy between the blue and red lines is due to a mismatch between the required and available thrusts and the discrete PWM realization of the control. In the steady state, both curves remain near zero with small oscillations due to the on-off actuation. 85 Case 1 Case 2 Case 3 Case 4 Fig. 4 – Variation of the out-of-plane control actions Figure 5 demonstrates a similar pattern across all four parameter plots, where after a short initial interval with possible sharp jumps, all curves vary slowly. Parameters 𝑠1 and 𝑠2 stabilize around 2–3 units. Typically 𝑠2 slightly exceeds 𝑠1. The parameter 𝛽 quickly drops from a large initial value to a small one (< 1) and then varies slowly. The parameter 𝜀 shown as 𝜀 × 10 in the plot, stays near 2–3 on the scale, i.e., 𝜀 ≈ 0.2–0.3. Small “steps” on all curves reflect discrete parameter updates. In the steady state, deviations are minor. Case 1 Case 2 Case 3 Case 4 Fig. 5 – Parameter variation for the out-of-plane channel 86 In Fig. 6, all state components converge to zero. State components 𝑥1 and 𝑥2 (red and blue) decrease smoothly with no noticeable overshoot. Velocities 𝑥3 and 𝑥4 (green and purple) rise to zero in “stair-steps,” reflecting discrete actions. After about (300–500) s, all curves remain near zero and only small periodic oscillations, which are typical for on-off control, are visible. These plots mainly differ in the convergence speed and the magnitude of steady-state errors. Case 1 Case 2 Case 3 Case 4 Fig. 6 – Variation of the in-plane state Case 1 Case 2 Case 3 Case 4 Fig. 7 – Variation of the in-plane control actions 87 Case 1 Case 2 Case 3 Case 4 Fig. 8 – Parameter variation for the in-plane channel 88 As can be seen in Fig. 7, the sequence of dense same-sign pulses rapidly reduces the state errors at the beginning of the episodes. After that, the pulses become sparser and shorter, occasionally changing sign for accurate corrections. As stabilization proceeds, both commands stay near zero, and the pulses have short duration and low frequency, resulting in state chartering around equilibrium due to on-off actuation. The difference between the channels appears in pulse phasing and polarity, but the errors converge to small magnitudes in both channels. Figure 8 demonstrates that after a short initial jump, the parameters change slowly. The parameters 𝑠𝑖𝑗 are stabilized with a clear ordering, when 𝑠12, 𝑠22 remain higher than 𝑠11, 𝑠21. The parameters 𝛽1, 𝛽2 (plotted × 10) typically decrease from high initial values and then oscillate around small values. The parameters 𝜀1, 𝜀2 settle down quickly and vary little over the whole interval. In both channels, the trajectories are similar in shape and level, i.e., the settings converge to close steady values with only minor drift. Out-of-plane channel In-plane channel Fig. 9 – Evolution of the cumulative return during training (Case 2) Figure 9 shows a typical evolution of the cumulative return during training. At the beginning of training, returns have a large value with high variance. After a while, the mean return improves monotonically and approaches zero. The critic’s estimates of the return initially are biased, but once the policy forms, it tracks the mean curve well. In the first plot, convergence is faster, while in the second one, it takes longer training with deeper occasional drops. However, the mean return also moves toward zero. After convergence, some episodes may still result in relatively high negative returns, but the overall trend remains stable. 5. Robustness analysis In practice, real model parameters may deviate from their nominal values, so we assess the robustness of the SMC-based intelligent controllers against model parameter variations and compare it with the robustness of the intelligent agents from Ref. [17]. In this article, two RL-based agents for relative spacecraft control were 89 proposed that differ in their input vector. The first agent (IA1) receives the standard state vector 𝑋𝑘, while the input of the second agent (IA2) is augmented with the previous step control action, i.e.,[𝑋𝑘 , 𝑢𝑘−1] 𝑇. Table 4 presents the description of test cases, which differ in the percentage deviation of the considered parameter from its nominal value. Table 4. Description of test cases Symbol Nom. В. 1 В. 2 В. 3 В. 4 В. 5 В. 6 Description Nominal parameter value +10 % +20 % +30 % –10 % –20 % –30 % Figure 10 shows the state variations for IA1, when SSC mass 𝑚𝑠 deviates from the nominal one. As can be seen from these plots, the agent remains robust when the spacecraft mass decreases, but control performance degrades as mass increases (Case B.1). In Cases B.2 and B.3 the agent fails to maintain stability. This occurs because a greater spacecraft mass requires larger control amplitudes than those considered during training. Fig. 10 – State variations for IA1 and different spacecraft masses 90 Figure 11 – State variations for IA2 and different spacecraft masses Fig. 11 shows the state variations for IA2, when SSC mass 𝑚𝑠 deviates from the nominal one. The agent remains robust, but control performance degrades as mass increases (Case B.1). Unlike IA1, IA2 maintains stability in Cases B.2 and B.3, although the performance degrades in these cases. Figure 12 presents the state trajectories for the SMC-based agent and various cases of the SSC masses. Stability is preserved in all cases with a similar transient shape pattern. However, in some cases, the steady-state error increases. 91 Fig. 12 – State variations for SMC-based agent and different spacecraft masses. 6. Conclusions This paper presents an approach to relative on-off spacecraft control that combines sliding-mode control with reinforcement learning. The results confirm the potential of this approach for autonomous spacecraft control under constraints and uncertainty. The trained agent demonstrated superior robustness to uncertainty in plant parameters compared to the reported baselines, which is due to the strong inherent robustness of sliding-mode control. However, a large number of episodes is required to train the agent, which restricts the practical implementation of this approach in its current form. Accordingly, future work should explore ways to improve the training efficiency of the agent. 1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris Management. arXiv:1901.11121, 2019. 15 p. 2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments. Progress in Aerospace Sciences. 2019. Vol. 108. P. 32–120. https://doi.org/10.1016/j.paerosci.2019.01.004 3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to spacecraft attitude controller design. Automatika. 2016. No. 57. P. 212–220. https://doi.org/10.7305/automatika.2016.07.618 4. Anthony T., Wie B., Carroll S. Pulse-Modulated Control Synthesis for a Flexible Spacecraft. Journal of Guidance. 1989. Vol. 13(6). P. 1014–1022. https://doi.org/10.2514/3.20574 5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an Algorithm to Control the Angular Motion of Spacecraft Equipped with an Aeromagnetic Deorbiting System. Eastern-European Journal of Enterprise Technologies. 2020. 1(5(103)). P. 37–46. https://doi.org/10.15587/1729-4061.2020.192813 https://doi.org/10.1016/j.paerosci.2019.01.004 https://doi.org/10.7305/automatika.2016.07.618 https://doi.org/10.2514/3.20574 https://doi.org/10.15587/1729-4061.2020.192813 92 6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016. 800 p. 7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks. Communications of the ACM. 2017. 60(6). P. 84–90. https://doi.org/10.1145/3065386 8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017. 31(16). P. 821–835. https://doi.org/10.1080/01691864.2017.1365009 9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving. Electronic Imaging. 2017. Issue 19. P. 70–76. https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023 10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017. 550. P. 354–359. https://doi.org/10.1038/nature24270 11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics. 2019. 3. P. 287–299. https://doi.org/10.1007/s42064-018-0053-6 12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and Technology (Космічна наука і технологія). 2021. 27(6/133). P. 38–52. https://doi.org/10.15407/knit2021.06.038 13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. 18(7). https://doi.org/10.2514/1.I010914 14. Gaudet B., Linares R., Furfaro R. Six Degree-of-Freedom Hovering using LIDAR Altimetry via Reinforcement Meta-Learning. Acta Astronautica. 2020. 172. P. 90–99. https://doi.org/10.1016/j.actaastro.2020.03.026 15. Gaudet B., Linares R., Furfaro R. Seeker based Adaptive Guidance via Reinforcement Meta-Learning Applied to Asteroid Close Proximity Operations. Acta Astronautica. 2020. 171. P. 1–13. https://doi.org/10.1016/j.actaastro.2020.02.036 16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object via deep learning. Space Science and Technology (Космічна наука і технологія). 2022. 28(5/138). P. 15–26. https://doi.org/10.15407/knit2022.05.015 17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and Technology (Космічна наука і технологія). 2024. 30(2/147). P. 3–14. https://doi.org/10.15407/knit2024.02.003 18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. innov. (Наука та інновації). 2018. 14(4). P. 5–16. https://doi.org/10.15407/scine14.04.005 19. Steinberger M., Horn M., Fridman L. (eds). Variable-Structure Systems and Sliding-Mode Control. Springer- Verlag, London, 2020. (Studies in Systems, Decision and Control; Vol. 271). https://doi.org/10.1007/978-3- 030-36621-6 20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Washington: Hemisphere Publishing, 1975. P. 224–235. 21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. P. 47–65. 22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. 13 p. 23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep Reinforcement Learning. arXiv:1602.01783, 2016. Received on October 14, 2025, in final form on December 3, 2025 https://doi.org/10.1145/3065386 https://doi.org/10.1080/01691864.2017.1365009 https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023 https://doi.org/10.1038/nature24270 https://doi.org/10.1007/s42064-018-0053-6 https://doi.org/10.15407/knit2021.06.038 https://doi.org/10.2514/1.I010914 https://doi.org/10.1016/j.actaastro.2020.03.026 https://doi.org/10.1016/j.actaastro.2020.02.036 https://doi.org/10.15407/knit2022.05.015 https://doi.org/10.15407/knit2024.02.003 https://doi.org/10.15407/scine14.04.005 https://doi.org/10.1007/978-3-030-36621-6 https://doi.org/10.1007/978-3-030-36621-6
id oai:ojs2.journal-itm.dp.ua:article-157
institution Technical Mechanics
keywords_txt_mv keywords
language English
last_indexed 2026-07-14T01:00:44Z
publishDate 2025
publisher текст 3
record_format ojs
resource_txt_mv journal-itmdpua/56/ead0043608f8761768e2ce5fbc315156.pdf
spelling oai:ojs2.journal-itm.dp.ua:article-1572026-07-13T20:26:18Z ON-OFF SPACECRAFT RELATIVE CONTROL IN SLIDING MODE VIA REINFORCEMENT LEARNING ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ SOROCHINSKII, V. V. KHOROSHYLOV, S. V. LEVCHUK, I. L. DUBOVYK, T. M. HUZ, H. M. ROMANCHUK, O. O. навчання з підкріпленням, проксимальна оптимізація політики, керування косміч¬ним апаратом, орбітальні сервісні операції, on-off керування, автономні системи керування. reinforcement learning, proximal policy optimization, spacecraft control, on-orbit servicing, on–off control, autonomous control systems. DOI: https://doi.org/10.15407/itm2025.04.077 The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric uncertainties. The goal is to develop and assess an approach that combines sliding-mode control with modern reinforcement-learning methods tailored for resource-constrained onboard implementation. Relative motion dynamics is formulated in an orbital coordinate frame with normalized states and discretized in time. Binary actions with pulse-width modulation, subject to constraints on the thrust level, pulse duration, and duty cycle, represent the impulsive nature of actuation. We propose a combined synthesis in which the sliding-surface parameters and switching rules are tuned via proximal policy optimization within an actor-critic architecture. The actor and critic are implemented as neural networks that approximate the policy and the value function, respectively. The actor neural network takes the state vector as input information and outputs the mean and standard deviation of the parameters of the sliding mode control law. The value function penalizes both the state error and control effort, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. Two uncoupled agents are designed to control spacecraft relative orbital motion in in-plane and out-of-plane directions independently. The proximal policy optimization hyperparameters are selected to ensure a trade-off among the learning time, stability, and control performance. The reinforcement-learning agents are trained and analyzed considering four cases that differ in the thrust levels and weighting matrices. The quality functional combines state deviation and thrust use penalties, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. The results confirm the potential of this approach for autonomous spacecraft control under constraints and uncertainty. Compared with reported baselines, the trained agent shows superior robustness to plant-parameter uncertainty, which we attribute to the inherent robust properties of sliding-mode control. These findings have the potential to improve the efficiency and autonomy of on-orbit servicing operations. REFERENCES 1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris Management. arXiv:1901.11121, 2019. 2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments. Progress in Aerospace Sciences. 2019. V. 108. Pp. 32-120.https://doi.org/10.1016/j.paerosci.2019.01.004 3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to spacecraft attitude controller design. Automatika. 2016. V. 57. Pp. 212-220.https://doi.org/10.7305/automatika.2016.07.618 4. Anthony T., Wie B., Carroll S. Pulse-modulated control synthesis for a flexible spacecraft. Journal of Guidance. 1989. V. 13. No. 6. Pp. 1014-1022. https://doi.org/10.2514/3.20574 5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an algorithm to control the angular motion of spacecraft equipped with an aeromagnetic deorbiting system. Eastern-European Journal of Enterprise Technologies. 2020. V. 1. No. 5. Pp. 37-46. https://doi.org/10.15587/1729-4061.2020.192813 6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016. 7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks. Communications of the ACM. 2017. V. 60. No. 6. Pp. 84-90. https://doi.org/10.1145/3065386 8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017. V. 31. No. 16. Pp. 821-835. https://doi.org/10.1080/01691864.2017.1365009 9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving. Electronic Imaging. 2017. V. 19. Pp. 70-76.https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023 10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017. V. 550. Pp. 354-359. https://doi.org/10.1038/nature24270 11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics. 2019. No. 3. Pp. 287-299. https://doi.org/10.1007/s42064-018-0053-6 12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and Technology. 2021. V. 27. No. 6. Pp. 38-52. https://doi.org/10.15407/knit2021.06.038 13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. V. 18. No. 7.  https://doi.org/10.2514/1.I010914 14. Gaudet B., Linares R., Furfaro R. Six degree-of-freedom hovering using LIDAR altimetry via reinforcement meta-learning. Acta Astronautica. 2020. V. 172. Pp. 90-99.https://doi.org/10.1016/j.actaastro.2020.03.026 15. Gaudet B., Linares R., Furfaro R. Seeker based adaptive guidance via reinforcement meta-learning applied to asteroid close proximity operations. Acta Astronautica. 2020. V. 171. Pp. 1-13.https://doi.org/10.1016/j.actaastro.2020.02.036 16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object via deep learning. Space Science and Technology. 2022. V. 28. No. 5. Pp. 15-26.https://doi.org/10.15407/knit2022.05.015 17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and Technology. 2024. V. 30. No. 2. Pp. 3-14. https://doi.org/10.15407/knit2024.02.003 18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. Innov. 2018. V. 14. No. 4. Pp. 5-16. https://doi.org/10.15407/scine14.04.005 19. Steinberger M., Horn M., Fridman L. (Eds). Variable-Structure Systems and Sliding-Mode Control. Springer-Verlag, 2020. (Studies in Systems, Decision and Control; Vol. 271).https://doi.org/10.1007/978-3-030-36621-6 20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Hemisphere Publishing, 1975. Pp. 224-235. 21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. Pp. 47-65. 22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. 13 pp. 23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep Reinforcement Learning. arXiv:1602.01783, 2016. DOI: https://doi.org/10.15407/itm2025.04.077 Розглянуто задачу відносного імпульсного керування рухом космічного апарата у ковзному режимі для автономних орбітальних сервісних операцій за наявності обмежень на амплітуду керуючих впливів, дискретності дій та параметричних невизначеностей. Метою роботи є розробка й оцінювання підходу, що поєднує принципи ковзного керування з сучасними методами навчання з підкріпленням, орієнтованими на бортову реалізацію з обмеженими ресурсами. Динаміку відносного руху задано в орбітальній системі координат у нормалізованих змінних і дискредитовано. Імпульсний характер впливів виконавчих органів відображено через бінарні дії з широтно-імпульсною модуляцією та обмеженнями на рівень тяги, тривалість і період увімкнень. Запропоновано комбінований синтез, у якому параметри поверхні ковзання та правила перемикання налаштовуються методом проксимальної оптимізації політики з використанням архітектури актор-критик. Актор і критик реалізовані у вигляді нейронних мереж, які відповідно апроксимують політику та функцію цінності. Нейронна мережа актора приймає вектор стану як вхідну інформацію і видає середнє значення та стандартне відхилення параметрів закону керування у ковзному режимі. Функція цінності штрафує як за помилку стану, так і за витрати на керування, що дозволяє забезпечити компроміс між швидкістю реагування, точністю та витратою палива. Два незалежні агенти розроблені для керування відносним орбітальним рухом космічного апарата окремо в напрямку площини орбіти та у перпендикулярному напрямку. Гіперпараметри оптимізації проксимальної політики обрано для забезпечення компромісу між часом навчання, стабільністю та якістю керування. Агенти навчання з підкріпленням навчeні та проаналізовані з урахуванням чотирьох випадків, що відрізняються рівнями тяги та ваговими матрицями. Функціонал якості об’єднує штрафи за відхилення стану та використання тяги, що дає змогу знаходити компроміс між швидкодією, точністю та витратами робочого тіла. Отримані результати підтверджують потенціал такого підходу для задач автономного керування космічних апаратів в умовах обмежень та невизначеності. У порівнянні з відомими результатами навчений агент продемонстрував кращу робастність по відношенню до невизначенності параметрів моделі об’єкта керування, що пояснюється сильними робастними властивостями керування в ковзному режимі. Отримані результати мають потенціал підвищити ефективність та автономність орбітальних сервісних операцій. ПОСИЛАННЯ 1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris Management. arXiv:1901.11121, 2019. 15 p. 2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments. Progress in Aerospace Sciences. 2019. Vol. 108. P. 32–120. https://doi.org/10.1016/j.paerosci.2019.01.004 3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to spacecraft attitude controller design. Automatika. 2016. No. 57. P. 212–220. https://doi.org/10.7305/automatika.2016.07.618 4. Anthony T., Wie B., Carroll S. Pulse-Modulated Control Synthesis for a Flexible Spacecraft. Journal of Guidance. 1989. Vol. 13(6). P. 1014–1022. https://doi.org/10.2514/3.20574 5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an Algorithm to Control the Angular Motion of Spacecraft Equipped with an Aeromagnetic Deorbiting System. Eastern-European Journal of Enterprise Technologies. 2020. 1(5(103)). P. 37–46. https://doi.org/10.15587/1729-4061.2020.192813 6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016. 800 p. 7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks. Communications of the ACM. 2017. 60(6). P. 84–90. https://doi.org/10.1145/3065386 8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017. 31(16). P. 821–835. https://doi.org/10.1080/01691864.2017.1365009 9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving. Electronic Imaging. 2017. Issue 19. P. 70–76. https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023 10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017. 550. P. 354–359. https://doi.org/10.1038/nature24270 11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics. 2019. 3. P. 287–299. https://doi.org/10.1007/s42064-018-0053-6 12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and Technology (Космічна наука і технологія). 2021. 27(6/133). P. 38–52. https://doi.org/10.15407/knit2021.06.038 13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. 18(7). https://doi.org/10.2514/1.I010914 14. Gaudet B., Linares R., Furfaro R. Six Degree-of-Freedom Hovering using LIDAR Altimetry via Reinforcement Meta-Learning. Acta Astronautica. 2020. 172. P. 90–99. https://doi.org/10.1016/j.actaastro.2020.03.026 15. Gaudet B., Linares R., Furfaro R. Seeker based Adaptive Guidance via Reinforcement Meta-Learning Applied to Asteroid Close Proximity Operations. Acta Astronautica. 2020. 171. P. 1–13. https://doi.org/10.1016/j.actaastro.2020.02.036 16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object via deep learning. Space Science and Technology (Космічна наука і технологія). 2022. 28(5/138). P. 15–26. https://doi.org/10.15407/knit2022.05.015 17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and Technology (Космічна наука і технологія). 2024. 30(2/147). P. 3–14. https://doi.org/10.15407/knit2024.02.003 18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. innov. (Наука та інновації). 2018. 14(4). P. 5–16. https://doi.org/10.15407/scine14.04.005 19. Steinberger M., Horn M., Fridman L. (eds). Variable-Structure Systems and Sliding-Mode Control. Springer-Verlag, London, 2020. (Studies in Systems, Decision and Control; Vol. 271). https://doi.org/10.1007/978-3-030-36621-6 20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Washington: Hemisphere Publishing, 1975. P. 224–235. 21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. P. 47–65. 22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. 13 p. 23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep Reinforcement Learning. arXiv:1602.01783, 2016.   текст 3 2025-12-11 Article Article application/pdf https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157 Technical Mechanics; No. 4 (2025): Technical Mechanics; 77-92 Институт технической механики Национальной академии наук Украины и Государственного космического агентства Украины; № 4 (2025): Technical Mechanics; 77-92 ТЕХНІЧНА МЕХАНІКА; № 4 (2025): ТЕХНІЧНА МЕХАНІКА; 77-92 en https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157/66 Copyright (c) 2025 Technical Mechanics
spellingShingle навчання з підкріпленням
проксимальна оптимізація політики
керування косміч¬ним апаратом
орбітальні сервісні операції
on-off керування
автономні системи керування.
SOROCHINSKII, V. V.
KHOROSHYLOV, S. V.
LEVCHUK, I. L.
DUBOVYK, T. M.
HUZ, H. M.
ROMANCHUK, O. O.
ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ
title ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ
title_alt ON-OFF SPACECRAFT RELATIVE CONTROL IN SLIDING MODE VIA REINFORCEMENT LEARNING
title_full ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ
title_fullStr ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ
title_full_unstemmed ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ
title_short ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ
title_sort імпульсне керування відносним рухом космічних апаратів у ковзному режимі з використанням навчання з підкріпленням
topic навчання з підкріпленням
проксимальна оптимізація політики
керування косміч¬ним апаратом
орбітальні сервісні операції
on-off керування
автономні системи керування.
topic_facet навчання з підкріпленням
проксимальна оптимізація політики
керування косміч¬ним апаратом
орбітальні сервісні операції
on-off керування
автономні системи керування.
reinforcement learning
proximal policy optimization
spacecraft control
on-orbit servicing
on–off control
autonomous control systems.
url https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157
work_keys_str_mv AT sorochinskiivv onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning
AT khoroshylovsv onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning
AT levchukil onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning
AT dubovyktm onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning
AT huzhm onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning
AT romanchukoo onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning
AT sorochinskiivv ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm
AT khoroshylovsv ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm
AT levchukil ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm
AT dubovyktm ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm
AT huzhm ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm
AT romanchukoo ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm