ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ
DOI: https://doi.org/10.15407/itm2025.04.077 The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric uncertainties. The goal is to develop and assess an app...
Gespeichert in:
| Datum: | 2025 |
|---|---|
| Hauptverfasser: | , , , , , |
| Format: | Artikel |
| Sprache: | Englisch |
| Veröffentlicht: |
текст 3
2025
|
| Schlagworte: | |
| Online Zugang: | https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| Назва журналу: | Technical Mechanics |
| Завантажити файл: | |
Institution
Technical Mechanics| _version_ | 1870649983994167296 |
|---|---|
| author | SOROCHINSKII, V. V. KHOROSHYLOV, S. V. LEVCHUK, I. L. DUBOVYK, T. M. HUZ, H. M. ROMANCHUK, O. O. |
| author_facet | SOROCHINSKII, V. V. KHOROSHYLOV, S. V. LEVCHUK, I. L. DUBOVYK, T. M. HUZ, H. M. ROMANCHUK, O. O. |
| author_institution_txt_mv | [
{
"author": "V. V. SOROCHINSKII",
"institution": "Institute of Technical Mechanics of the National Academy of Science of Ukraine and the State Space Agency of Ukraine, 15 Leshko-Popel St., Dnipro 49005, Ukraine; e-mail: vovas99@ukr.net"
},
{
"author": "S. V. KHOROSHYLOV",
"institution": "Institute of Technical Mechanics of the National Academy of Science of Ukraine and the State Space Agency of Ukraine, 15 Leshko-Popel St., Dnipro 49005, Ukraine"
},
{
"author": "I. L. LEVCHUK",
"institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine"
},
{
"author": "T. M. DUBOVYK",
"institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine"
},
{
"author": "H. M. HUZ",
"institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine"
},
{
"author": "O. O. ROMANCHUK",
"institution": "Ukrainian State University of Science and Technologies, 2 Lazariana St., Dnipro 49010, Ukraine"
}
] |
| author_sort | SOROCHINSKII, V. V. |
| baseUrl_str | https://journal-itm.dp.ua/ojs/index.php/ITM_j1/oai |
| collection | OJS |
| datestamp_date | 2026-07-13T20:26:18Z |
| description | DOI: https://doi.org/10.15407/itm2025.04.077
The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric uncertainties. The goal is to develop and assess an approach that combines sliding-mode control with modern reinforcement-learning methods tailored for resource-constrained onboard implementation. Relative motion dynamics is formulated in an orbital coordinate frame with normalized states and discretized in time. Binary actions with pulse-width modulation, subject to constraints on the thrust level, pulse duration, and duty cycle, represent the impulsive nature of actuation. We propose a combined synthesis in which the sliding-surface parameters and switching rules are tuned via proximal policy optimization within an actor-critic architecture. The actor and critic are implemented as neural networks that approximate the policy and the value function, respectively. The actor neural network takes the state vector as input information and outputs the mean and standard deviation of the parameters of the sliding mode control law. The value function penalizes both the state error and control effort, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. Two uncoupled agents are designed to control spacecraft relative orbital motion in in-plane and out-of-plane directions independently. The proximal policy optimization hyperparameters are selected to ensure a trade-off among the learning time, stability, and control performance. The reinforcement-learning agents are trained and analyzed considering four cases that differ in the thrust levels and weighting matrices. The quality functional combines state deviation and thrust use penalties, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. The results confirm the potential of this approach for autonomous spacecraft control under constraints and uncertainty. Compared with reported baselines, the trained agent shows superior robustness to plant-parameter uncertainty, which we attribute to the inherent robust properties of sliding-mode control. These findings have the potential to improve the efficiency and autonomy of on-orbit servicing operations.
REFERENCES
1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris Management. arXiv:1901.11121, 2019.
2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments. Progress in Aerospace Sciences. 2019. V. 108. Pp. 32-120.https://doi.org/10.1016/j.paerosci.2019.01.004
3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to spacecraft attitude controller design. Automatika. 2016. V. 57. Pp. 212-220.https://doi.org/10.7305/automatika.2016.07.618
4. Anthony T., Wie B., Carroll S. Pulse-modulated control synthesis for a flexible spacecraft. Journal of Guidance. 1989. V. 13. No. 6. Pp. 1014-1022. https://doi.org/10.2514/3.20574
5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an algorithm to control the angular motion of spacecraft equipped with an aeromagnetic deorbiting system. Eastern-European Journal of Enterprise Technologies. 2020. V. 1. No. 5. Pp. 37-46. https://doi.org/10.15587/1729-4061.2020.192813
6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016.
7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks. Communications of the ACM. 2017. V. 60. No. 6. Pp. 84-90. https://doi.org/10.1145/3065386
8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017. V. 31. No. 16. Pp. 821-835. https://doi.org/10.1080/01691864.2017.1365009
9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving. Electronic Imaging. 2017. V. 19. Pp. 70-76.https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023
10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017. V. 550. Pp. 354-359. https://doi.org/10.1038/nature24270
11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics. 2019. No. 3. Pp. 287-299. https://doi.org/10.1007/s42064-018-0053-6
12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and Technology. 2021. V. 27. No. 6. Pp. 38-52. https://doi.org/10.15407/knit2021.06.038
13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. V. 18. No. 7. https://doi.org/10.2514/1.I010914
14. Gaudet B., Linares R., Furfaro R. Six degree-of-freedom hovering using LIDAR altimetry via reinforcement meta-learning. Acta Astronautica. 2020. V. 172. Pp. 90-99.https://doi.org/10.1016/j.actaastro.2020.03.026
15. Gaudet B., Linares R., Furfaro R. Seeker based adaptive guidance via reinforcement meta-learning applied to asteroid close proximity operations. Acta Astronautica. 2020. V. 171. Pp. 1-13.https://doi.org/10.1016/j.actaastro.2020.02.036
16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object via deep learning. Space Science and Technology. 2022. V. 28. No. 5. Pp. 15-26.https://doi.org/10.15407/knit2022.05.015
17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and Technology. 2024. V. 30. No. 2. Pp. 3-14. https://doi.org/10.15407/knit2024.02.003
18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. Innov. 2018. V. 14. No. 4. Pp. 5-16. https://doi.org/10.15407/scine14.04.005
19. Steinberger M., Horn M., Fridman L. (Eds). Variable-Structure Systems and Sliding-Mode Control. Springer-Verlag, 2020. (Studies in Systems, Decision and Control; Vol. 271).https://doi.org/10.1007/978-3-030-36621-6
20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Hemisphere Publishing, 1975. Pp. 224-235.
21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. Pp. 47-65.
22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. 13 pp.
23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep Reinforcement Learning. arXiv:1602.01783, 2016. |
| first_indexed | 2025-12-17T12:05:41Z |
| format | Article |
| fulltext |
77
Автоматизація, комп’ютерно-інтегровані
технології
Automation, computer-integrated technologies
UDC 629.5 https://doi.org/10.15407/itm2025.04.077
V. V. SOROCHINSKII1, S. V. KHOROSHYLOV1, I. L. LEVCHUK2,
T. M. DUBOVYK2, H. M. HUZ2, O. O. ROMANCHUK2
ON-OFF SPACECRAFT RELATIVE CONTROL IN SLIDING MODE VIA
REINFORCEMENT LEARNING
1Institute of Technical Mechanics of the National Academy of Science of Ukraine and the State Space Agency
of Ukraine, 15 Leshko-Popel Str., Dnipro, 49005, Ukraine; e-mail: vovas99@ukr.net
2Ukrainian State University of Science and Technologies,2 Lazariana St., Dnipro, 49010, Ukraine
Розглянуто задачу відносного імпульсного керування рухом космічного апарата у ковзному
режимі для автономних орбітальних сервісних операцій за наявності обмежень на амплітуду
керуючих впливів, дискретності дій та параметричних невизначеностей. Метою роботи є розробка й
оцінювання підходу, що поєднує принципи ковзного керування з сучасними методами навчання з
підкріпленням, орієнтованими на бортову реалізацію з обмеженими ресурсами. Динаміку відносного
руху задано в орбітальній системі координат у нормалізованих змінних і дискредитовано. Імпульсний
характер впливів виконавчих органів відображено через бінарні дії з широтно-імпульсною
модуляцією та обмеженнями на рівень тяги, тривалість і період увімкнень. Запропоновано
комбінований синтез, у якому параметри поверхні ковзання та правила перемикання налаштовуються
методом проксимальної оптимізації політики з використанням архітектури актор-критик. Актор і
критик реалізовані у вигляді нейронних мереж, які відповідно апроксимують політику та функцію
цінності. Нейронна мережа актора приймає вектор стану як вхідну інформацію і видає середнє
значення та стандартне відхилення параметрів закону керування у ковзному режимі. Функція
цінності штрафує як за помилку стану, так і за витрати на керування, що дозволяє забезпечити
компроміс між швидкістю реагування, точністю та витратою палива. Два незалежні агенти
розроблені для керування відносним орбітальним рухом космічного апарата окремо в напрямку
площини орбіти та у перпендикулярному напрямку. Гіперпараметри оптимізації проксимальної
політики обрано для забезпечення компромісу між часом навчання, стабільністю та якістю керування.
Агенти навчання з підкріпленням навчeні та проаналізовані з урахуванням чотирьох випадків, що
відрізняються рівнями тяги та ваговими матрицями. Функціонал якості об’єднує штрафи за
відхилення стану та використання тяги, що дає змогу знаходити компроміс між швидкодією,
точністю та витратами робочого тіла. Отримані результати підтверджують потенціал такого підходу
для задач автономного керування космічних апаратів в умовах обмежень та невизначеності. У
порівнянні з відомими результатами навчений агент продемонстрував кращу робастність по
відношенню до невизначенності параметрів моделі об’єкта керування, що пояснюється сильними
робастними властивостями керування в ковзному режимі. Отримані результати мають потенціал
підвищити ефективність та автономність орбітальних сервісних операцій.
Ключові слова: навчання з підкріпленням, проксимальна оптимізація політики, керування космічним
апаратом, орбітальні сервісні операції, on-off керування, автономні системи керування.
The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous
on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric
uncertainties. The goal is to develop and assess an approach that combines sliding-mode control with modern
reinforcement-learning methods tailored for resource-constrained onboard implementation. Relative motion
dynamics is formulated in an orbital coordinate frame with normalized states and discretized in time. Binary
actions with pulse-width modulation, subject to constraints on the thrust level, pulse duration, and duty cycle,
represent the impulsive nature of actuation. We propose a combined synthesis in which the sliding-surface
parameters and switching rules are tuned via proximal policy optimization within an actor-critic architecture.
The actor and critic are implemented as neural networks that approximate the policy and the value function,
respectively. The actor neural network takes the state vector as input information and outputs the mean and
standard deviation of the parameters of the sliding mode control law. The value function penalizes both the
state error and control effort, thus enabling a trade-off among the response speed, accuracy, and propellant
consumption. Two uncoupled agents are designed to control spacecraft relative orbital motion in in-plane and
out-of-plane directions independently. The proximal policy optimization hyperparameters are selected to
ensure a trade-off among the learning time, stability, and control performance. The reinforcement-learning
agents are trained and analyzed considering four cases that differ in the thrust levels and weighting matrices.
The quality functional combines state deviation and thrust use penalties, thus enabling a trade-off among the
© V. V. Sorochinskii, S. V. Khoroshylov, I. L. Levchuk, T. M. Dubovyk, H. M. Huz, O. O. Romanchuk, 2025
Техн. механіка. – 2025. – № 4.
https://doi.org/10.15407/itm2025.04.0
mailto:vovas99@ukr.net
78
response speed, accuracy, and propellant consumption. The results confirm the potential of this approach for
autonomous spacecraft control under constraints and uncertainty. Compared with reported baselines, the
trained agent shows superior robustness to plant-parameter uncertainty, which we attribute to the inherent
robust properties of sliding-mode control. These findings have the potential to improve the efficiency and
autonomy of on-orbit servicing operations.
Keywords: reinforcement learning, proximal policy optimization, spacecraft control, on-orbit servicing, on–
off control, autonomous control systems.
1. Introduction
Modern space missions increasingly require complex maneuvers with
minimum human involvement. Typical tasks include on-orbit servicing operations
such as docking, towing, and active debris removal [1, 2].
To perform these operations, a servicing spacecraft (SSC) must maneuver
close to the orbital object (OO), addressing the challenge of relative guidance and
control. Control inputs are commonly generated by reaction thrusters (RTs). Unlike
reaction wheels or control-moment gyros (CMGs), thruster outputs are typically
binary (“on/off”). This switching behavior makes actuation highly nonlinear and
complicates direct controller design. To implement continuous-time control laws,
pulse-width modulation (PWM) and pulse-frequency modulation (PFM) are used
[3, 4], transforming the continuous control signal into a sequence of pulses with
specified duration or frequency. The accuracy of this approximation depends on the
modulator resolution: coarse quantization significantly reduces control
performance [5], so the regulator and modulator parameters must be co-designed,
which is a complex task.
The impressive results achieved using deep learning (DL) methods [6–10]
have recently intensified interest in AI approaches among researchers and
practitioners worldwide. Recently, these methods have also been applied to space
problems [11–16].
Publication [17] demonstrated the feasibility of directly synthesizing control
laws via reinforcement learning (RL), mapping the state vector to binary on/off
thruster commands. That approach outperformed a linear regulator in conjunction
with a PWM in terms of accuracy, response speed, and the number of thruster
firings. At the same time, it requires many training episodes and does not provide
strict robustness guarantees when training and operational conditions differ. To
mitigate these drawbacks, this study proposes combining sliding-mode control
(SMC) with RL.
This study aims to assess the feasibility of on/off spacecraft relative control in
sliding mode that is implemented via reinforcement learning, and to specify the
features of this approach
2. Problem Statement
2.1 Model of spacecraft relative dynamics. To describe the motion of the
servicing spacecraft (SSC) relative to an
orbital object (OO), we use the orbital
coordinate frame (OCF). The OCF origin
is fixed at the SSC center of mass (Fig. 1).
The 𝑥-axis is directed along the
geocentric radius vector from Earth’s
center to the SSC (𝑖)̂. The 𝑧-axis is normal
to the plane spanned by the 𝑥-axis and the
SSC orbital-velocity vector (𝑘̂) and is
oriented toward the positive orbital
Fig. 1 – Axes directions of the OCF
79
angular momentum. The 𝑦-axis completes a right-handed triad (𝑗̂).
The relative dynamics of the SSC–OO system can be written as the following
linearized system of equations [18]:
𝑥̈ − 𝜔2𝑥 − 2𝜔𝑦̇ − 𝜔̇𝑦 − 𝑘𝑥 =
𝑓𝑥
𝑑
𝑚𝑑 −
𝑓𝑥
𝑠
𝑚𝑠, (1)
𝑦 − 𝜔2𝑦 + 2𝜔𝑥̇ + 𝜔̇𝑥 + 𝑘𝑦 =
𝑓𝑦
𝑑
𝑚𝑑 −
𝑓𝑦
𝑠
𝑚𝑠, (2)
𝑧 + 𝑘𝑧 =
𝑓𝑧
𝑑
𝑚𝑑 −
𝑓𝑧
𝑠
𝑚𝑠, (3)
where 𝑥 , 𝑦 , 𝑧 are the OCF projections of the relative position vector 𝑟; 𝑚𝑠 and
𝑚𝑑 are the masses of the SSC and OO, respectively; 𝑓𝑥
𝑑, 𝑓𝑦
𝑑, 𝑓𝑧
𝑑 are the OCF
projections of the total force 𝑓𝑑, acting on the OO; 𝑓𝑥
𝑠, 𝑓𝑦
𝑠, 𝑓𝑧
𝑠 are the OCF
projections of the total force 𝑓𝑠, acting on the SSC. The total force 𝑓𝑠 includes
both control input and external disturbances acting on the SSC. The disturbances
𝑓𝑑 and 𝑓𝑠 may also account for the J2 perturbation, third-body gravity (Sun and
Moon), aerodynamic drag, and solar radiation pressure. The parameters in Eqs (1)
– (3) are defined as follows:
𝜔 = √
𝜇
𝑝3
(1 + 𝜀 cos 𝑣)𝜔, 𝜔̇ = −2𝜀√
𝜇
𝑝3 sin𝑣 (1 + 𝜀 cos 𝑣)𝜔,
𝑝 = 𝑎(1 − 𝜀2), 𝑘 =
𝜇
𝑟3, 𝑟 =
𝑎(1−𝜀2)
1+𝜀 cos𝑣
,
where 𝜇 is the Earth's gravitational constant, 𝜀 is the orbital eccentricity, 𝜈 is the
true anomaly, 𝑎 is the semi-major axis, and 𝑟 is the geocentric distance.
Equations (1) – (2) describe in-plane orbital motion, while Eq (3) describes
out-of-plane orbital motion.
Neglecting the external disturbances and introducing the state vectors 𝑋𝑖𝑛 =
[𝑥, 𝑦, 𝑥̇, 𝑦̇ ]𝑇, 𝑋𝑜𝑢𝑡 = [𝑧, 𝑧̇ ]𝑇 and the control vectors 𝑈𝑖𝑛 = [𝑢𝑥, 𝑢𝑦]
𝑇
, 𝑈𝑜𝑢𝑡 = 𝑢𝑧 ,
the model (1) can be written in the standard linear state-space form:
𝑋̇𝑖𝑛 = 𝐴𝑖𝑛𝑋𝑖𝑛 + 𝐵𝑖𝑛𝑈𝑖𝑛, 𝑋̇𝑜𝑢𝑡 = 𝐴𝑜𝑢𝑡𝑋𝑜𝑢𝑡 + 𝐵𝑜𝑢𝑡𝑈𝑜𝑢𝑡 , (4)
where
𝐴𝑖𝑛 = [
0 0 1 0
0 0 0 1
𝜔2 + 2𝑘 𝜔̇ 0 2𝜔
−𝜔̇ 𝜔2 − 𝑘 −2𝜔 0
], 𝐵𝑖𝑛 =
[
0 0
0 0
−
1
𝑚𝑠 0
0 −
1
𝑚𝑠]
,
𝐴𝑜𝑢𝑡 = [
0 1
−𝑘 0
], 𝐵𝑜𝑢𝑡 = [
0
−
1
𝑚𝑠
].
The magnitudes of individual state components differ significantly, which
may complicate neural-network (NN) training. To mitigate this issue, the state
vector is normalized as follows:
𝑋𝑖𝑛
𝑛 = [
𝑥
𝑥𝑚
,
𝑦
𝑦𝑚
,
𝑥̇
𝑥̇𝑚
,
𝑦̇
𝑦̇𝑚
]
𝑇
, 𝑋𝑜𝑢𝑡
𝑛 = [
𝑧
𝑧𝑚
,
𝑧̇
𝑧̇𝑚
]
𝑇
, (5)
where 𝑥𝑚, 𝑦𝑚, 𝑧𝑚, 𝑥̇𝑚, 𝑦̇𝑚, 𝑧̇𝑚 are the maximum magnitudes of the corresponding
state components. For the normalized state, the dynamic model can be given as
𝑋̇𝑛 = 𝐴𝑛𝑋𝑛 + 𝐵𝑛𝑈,
80
where 𝐴𝑛 = 𝑁−1𝐴𝑁, 𝐵𝑛 = 𝑁−1𝐵, 𝑁 = diag(𝑥𝑚, 𝑦𝑚, 𝑧𝑚, 𝑥̇𝑚, 𝑦̇𝑚, 𝑧̇𝑚).
Since the onboard controller is implemented using a digital computer, we
use the discrete-time form of the model:
𝑋𝑘+1 = 𝐴𝑘𝑋𝑘 + 𝐵𝑘𝑈𝑘, (6)
where 𝐴𝑘 = (𝐼 + 𝐴𝑛𝑇); 𝐵𝑘 = 𝐵𝑛𝑇; 𝑇 is the sampling period, and 𝑘 is the time
index. We also assume that the full state vector is measurable and that these
measurements are not corrupted by noise.
2.2. Sliding-Mode Control Sliding-mode control is a variable-structure control
in which the input is deliberately shaped to cause the trajectory to reach a sliding
surface and then move along it [19]. The control law is not a continuous function of
time, because the control structure switches depending on the current location in
the state space. The switching nature of SMC suits well naturally to the task of
spacecraft relative-motion control because reaction thrusters operate in the on/off
mode.
Motion along the sliding hypersurface is described by a reduced-order
equation and can be compared to the evolution along eigenmodes of linear systems.
To keep the system sliding requires a high-gain switching feedback.
The SMC design includes the following steps:
selecting a sliding hypersurface 𝑠(𝑥) = 0 such that motion along it yields the
desired system behavior;
synthesizing a feedback law with sufficient gains that guarantees both
reaching the surface and its invariance - i.e., the trajectory intersects the surface
and then remains on it.
The switching function 𝑠(𝑥) is defined as a measure of the deviation of the
state 𝑥 from the sliding hypersurface: the state belongs to the hypersurface when
𝑠(𝑥) = 0 and out of it when 𝑠(𝑥) ≠ 0. The control has a switching structure and
changes sign according to the sign of the s(x), e.g., 𝑢 = 𝑢+ when 𝑠(𝑥) > 0 and
𝑢 = 𝑢−), thus ensuring reaching and subsequent sliding along the surface.
Within the SMC paradigm, the control action is always directed to reduce the
magnitude of |𝑠(𝑥)| and to “push” the trajectory toward the sliding surface. When
the surface is reached, the trajectory slides along it and, depending on how the
surface is chosen, moves toward the target, typically an equilibrium (e.g., the
origin).
The switching function is defined as follows:
𝜎(𝑋) = 𝑆𝑇𝑋, (7)
Let the sliding surface be a simplex defined by 𝜎(𝑋) = 0. Then, for the
stability analysis, we consider the following Lyapunov candidate:
𝑉(𝜎(𝑋)) = 0.5𝜎(𝑋)𝑇𝜎(𝑋)
with the time derivative
𝑉̇(𝜎(𝑋)) = 0.5𝜎(𝑋)𝑇𝜎̇(𝑋).
The control is chosen so that 𝜎̇(𝑋) < 0 when 𝜎(𝑋) > 0 and 𝜎̇(𝑋) > 0
when 𝜎(𝑋) < 0 to ensure that 𝑉̇(𝜎(𝑋)) < 0. The evolution of the switching
function can be given as
𝜎̇(𝑋) = 𝑆𝑇(𝐴𝑋 + 𝐵𝑈).
81
The control input is formed as
𝑈 = (𝑆𝑇𝐵)−1 (−𝑆𝑇𝐴𝑋 + ℎ(𝜎(𝑋))),
where ℎ(𝜎(𝑋)) = −𝐾𝑖
𝜎(𝑋)
|𝜎(𝑋)|+𝜀
, and 𝜀 sets a dead zone that prevents excessive
control chattering.
Since the actuators operate in on/off mode, the SMC output is approximated
by a sequence of pulses of variable duration using a PWM. The pulse duration
within each sampling period is determined as follows:
𝑡𝑓 =
𝑢𝑘
𝑢𝑓
𝑇, 𝑡𝑓 ≤ 𝑇,
where 𝑢𝑓 is the nominal thrust level.
As the control performance metric, we use the following quadratic cost
function [20]:
𝐽 = ∑ (𝑋𝑘
𝑇𝑄𝑋𝑘 + 𝑈𝑘
𝑇𝑅𝑈𝑘)𝑁
𝑘=0 , (8)
where Q and R are the weighting matrices penalizing the states 𝑋𝑘 and the control 𝑈𝑘
, respectively.
To obtain an optimal sliding-mode controller, it is necessary to determine the
parameters 𝐾, 𝑆, and 𝜀 that minimize the quadratic cost (8). This optimization task is
addressed via reinforcement learning.
3. Reinforcement-Learning-Based Control
In the RL paradigm, the controller learns by analyzing the consequences of its
own actions [21]. These consequences are evaluated by a scalar reinforcement
signal received from the plant with which the controller interacts. The
reinforcement can be interpreted as a criterion that enables an intelligent system to
adjust its control policy to achieve a long-term objective.
The generic RL loop (Fig. 2) includes the following steps:
1. At the time 𝑡𝑘 the plant is in the state 𝑋𝑘. In this state, the controller (CS)
selects an action 𝑈𝑘 from the available action set.
2. The chosen action is applied, causing a transition to a new state 𝑋𝑘+1, and
the system receives a reinforcement signal 𝐶𝑘.
3. The algorithm repeats from step 2, taking into account the received
reinforcement; if the new state is terminal, the episode ends.
Let 𝜒 denote the state space and A the action space. The reinforcement
𝐶𝑘 results from applying the action 𝑈𝑘 in state 𝑋𝑘. Formally, it is a function on
𝜒 × 𝐴: 𝐶𝑘 = 𝐶(𝑋𝑘 , 𝑈𝑘)
Fig. 2 – RL setup
The CS selects actions to minimize the return, which is the discounted sum of
instantaneous costs:
𝐺𝑘 = 𝐶𝑘 + 𝛾𝐶𝑘+1 + 𝛾2𝐶𝑘+2+. . .= ∑ 𝛾𝑖∞
𝑖=0 𝐶𝑘+𝑖,0 ≤ 𝛾 ≤ 1,
where 𝛾 is the discount factor and 𝐶𝑘 is the instantaneous cost at step 𝑘.
82
The factor 𝛾 determines how strongly future costs influence current
action selection.
A key element of RL is the value function. Let the controller apply
actions according to a policy 𝜋:
𝑈𝑘 = 𝜋(𝑋𝑘).
Then the value function 𝑉𝜋(𝑋𝑘) gives the total discounted cumulative cost
when starting from the state 𝑋𝑘 and following the policy 𝜋:
𝑉𝜋(𝑋𝑘) = ∑ 𝛾𝑘∞
𝑖=0 𝐶𝑘+𝑖(𝑋𝑘+𝑖, 𝑈𝑘+𝑖) = 𝐶𝑘(𝑋𝑘, 𝑈𝑘) + 𝛾𝑉𝜋(𝑋𝑘+1).
RL can be implemented using an actor-critic architecture. In this case,
the critic estimates the value function for each state, while the actor maps
states to control actions.
In this study, the actor and critic are implemented as neural networks that
approximate, respectively, the policy and the value function:
𝑉𝜋(𝑋𝑘 , 𝜙), 𝜋(𝑋𝑘 , 𝜙 ),
where 𝜃 and 𝜙 are the parameter vectors of the critic and actor, respectively.
Proximal Policy Optimization (PPO) [22] is chosen among different RL
algorithms because it provides a tradeoff between implementation complexity
and training stability. This includes the following steps:
Cumulative cost estimation, which is the sum of the cost for this time step and
the discounted future cost [23]:
𝐺𝑡 = ∑ (𝛾𝑘−𝑡𝐶𝑘) + 𝑏𝛾𝑁−𝑡+1𝑉(𝑋𝑡𝑠+𝑁 , 𝜃)𝑡𝑠+𝑚
𝑘=𝑡 ,
where 𝑏 = 0 if 𝑋𝑡𝑠+𝑁 is terminal and 𝑏 = 1 otherwise (i.e., when the rollout
segment does not end at a terminal state, a bootstrap term from the critic 𝑉(⋅; 𝜃) is
added); 𝛾 ∈ [0,1] is the discount factor.
1. Advantage estimation:
𝐷𝑡 = 𝐺𝑡 − 𝑉(𝑋𝑡 , 𝜃),
2. Critic update via value loss minimization
𝐿𝑐𝑟𝑖𝑡𝑖𝑐(𝜃) =
1
𝑀
∑ (𝐺𝑖 − 𝑉(𝑋𝑖, 𝜃))
2𝑀
𝑖=1 ,
3. Actor update using clipped policy objective
𝐿𝑎𝑐𝑡𝑜𝑟(𝜙) =
1
𝑀
∑ (−𝑚𝑖𝑛(𝑟𝑖(𝜙) ∙ 𝐷𝑖, 𝑐𝑖(𝜙) ∙ 𝐷𝑖))
𝑀
𝑖=1 , 𝑟𝑖(𝜙) =
𝜋(𝑈𝑖|𝑋𝑖,𝜙)
𝜋(𝑈𝑖|𝑋𝑖,𝜙𝑜𝑙𝑑)
,
𝑐𝑖(𝜙) = 𝑚𝑎𝑥(𝑚𝑖𝑛(𝑟𝑖(𝜙), 1 + 𝜀), 1 − 𝜀) ,
where 𝐷𝑖 is the advantage for the 𝑖-th sample, 𝐺𝑖 is the corresponding return,
𝜋(𝑈𝑖|𝑋𝑖 , 𝜙) is the probability of action 𝑈𝑖 in state 𝑋𝑖 under the current policy
parameters 𝜙, 𝜋(𝑈𝑖|𝑋𝑖, 𝜙𝑜𝑙𝑑) is the probability under the pre-update policy, 𝜀 is the
clip factor, 𝑀 is the minibatch size.
The NN architecture of the intelligent agent is given in Tables 1 and 2.
Almost all network layers have the tanh activation, and only the actor’s output
layer employs SoftPlus function to provide positive outputs.
The actor NN takes the state vector as input and produces the mean and
standard deviation of the parameters 𝑆, 𝐾, and 𝜀. The structure of this NN is
shown in Fig. 3. In conjunction with this actor, we use the same NN critic and
the same quadratic performance index as in Ref [17].
83
Fig. 3 – Architecture of the actor neural network
Table 1. Number of neurons in the actor layers
Case Input
fc
hidden
fc_1/fc_4
hidden
fc_2/fc_5
hidden
fc_3/fc_6
hidden
Output
in-plain 𝑛𝑖𝑛 10𝑛𝑖𝑛 10𝑛𝑖𝑛 25 30 3
out-of-
plain
𝑛𝑜𝑢𝑡 10𝑛𝑜𝑢𝑡 10𝑛𝑜𝑢𝑡 50 60 6
Table 2. Number of neurons in the critic layers
Case Input 1-st hidden 2-nd hidden 3-d hidden Output
in-plain 𝑛𝑖𝑛 10𝑛𝑖𝑛 √100𝑛𝑖𝑛 10 1
out-of-
plain
𝑛𝑜𝑢𝑡 10𝑛𝑜𝑢𝑡 √100𝑛𝑜𝑢𝑡 10 1
4. Numerical results
For training and evaluation of the intelligent agent, the following system
parameters were used: a=7017 km; 𝑚𝑠 = 500 kg; 𝑚𝑑 = 1575 kg; T=200 sec; 𝑇𝑓 =
10 sec. Maximum magnitudes of the state-vector components: 𝑥𝑚 = 800 m, 𝑦𝑚 =
800 m, 𝑥̇𝑚=2 m/s, 𝑦̇𝑚=2 m/s.
We analyzed four cases that differ in the thrust levels and weighting matrices
(Table 3).
Table 3. Values of thrust levels and weighting matrices
in-plain out-of-plain
𝒖𝒇 Q R 𝒖𝒇 Q R
Case 1 8
[1 0; 0 0.001]
×1e-7
1e-3 8
[1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0
0 0 0.001]×1e-7
[1e-3 0; 0
1e-3]
Case 2 4
[1 0; 0
0.001]×1e-7
1e-3 4
[1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0
0 0 0.001]×1e-7
[1e-3 0; 0
1e-3]
Case 3 8
[1 0; 0
0.001]×1e-7
1e-4 8
[1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0 0
0 0.001]×1e-7
[1e-3 0;
0 1e-4]
Case 4 4
[1 0; 0
0.001]×1e-7
1e-4 4
[1 0 0 0; 0 1 0 0; 0 0 0.001 0; 0
0 0 0.001]×1e-7
[1e-3 0; 0
1e-4]
84
The reward function is designed as the weighted sum of state errors and control
efforts:
𝑟𝑡 = −(𝑥𝑡
𝑇𝑄𝑥𝑡 + 𝑢𝑡
𝑇𝑅𝑢𝑡). (10)
Training was conducted in episodes, each running until a maximum step count
was reached. The PPO hyperparameters were selected to ensure stable learning as
follows: 𝛾 = 0.98, 𝜀 = 0.2. An experience horizon of 50 and a minibatch size of 50
were used for updating the NN parameters.
Figure 3 demonstrates system stability since all trajectories of 𝑥5, 𝑥6 converge to
zero. The cases differ in how quickly the spacecraft reaches the target state and in the
magnitude of residual errors. Lower thrust leads to a longer transition and more
frequent small control actions, when the agent corrects the state more frequently
using small pulses. After settling, the residual error is close to zero in all cases. The
small remaining oscillations are due to the on-off control actions.
Case 1
Case 2
Case 3
Case 4
Fig. 3 – Variation of the out-of-plane state
In Fig. 4, the actor NN output signal is plotted by a blue line and its on-off
realization is shown by a red one. Initially, sequences of saturated same-sign pulses are
applied to reduce the error rapidly. After that, the pulses become sparser and shorter,
occasionally changing sign for fine corrections. In cases when higher thrust is
available, fewer control pulses are applied than in cases with lower thrust. The
discrepancy between the blue and red lines is due to a mismatch between the required
and available thrusts and the discrete PWM realization of the control. In the steady
state, both curves remain near zero with small oscillations due to the on-off actuation.
85
Case 1
Case 2
Case 3
Case 4
Fig. 4 – Variation of the out-of-plane control actions
Figure 5 demonstrates a similar pattern across all four parameter plots, where
after a short initial interval with possible sharp jumps, all curves vary slowly.
Parameters 𝑠1 and 𝑠2 stabilize around 2–3 units. Typically 𝑠2 slightly exceeds 𝑠1.
The parameter 𝛽 quickly drops from a large initial value to a small one (< 1) and
then varies slowly. The parameter 𝜀 shown as 𝜀 × 10 in the plot, stays near 2–3 on
the scale, i.e., 𝜀 ≈ 0.2–0.3. Small “steps” on all curves reflect discrete parameter
updates. In the steady state, deviations are minor.
Case 1
Case 2
Case 3
Case 4
Fig. 5 – Parameter variation for the out-of-plane channel
86
In Fig. 6, all state components converge to zero. State components 𝑥1 and 𝑥2
(red and blue) decrease smoothly with no noticeable overshoot. Velocities 𝑥3 and
𝑥4 (green and purple) rise to zero in “stair-steps,” reflecting discrete actions. After
about (300–500) s, all curves remain near zero and only small periodic oscillations,
which are typical for on-off control, are visible. These plots mainly differ in the
convergence speed and the magnitude of steady-state errors.
Case 1
Case 2
Case 3
Case 4
Fig. 6 – Variation of the in-plane state
Case 1
Case 2
Case 3
Case 4
Fig. 7 – Variation of the in-plane control actions
87
Case 1
Case 2
Case 3
Case 4
Fig. 8 – Parameter variation for the in-plane channel
88
As can be seen in Fig. 7, the sequence of dense same-sign pulses rapidly
reduces the state errors at the beginning of the episodes. After that, the pulses
become sparser and shorter, occasionally changing sign for accurate corrections.
As stabilization proceeds, both commands stay near zero, and the pulses have short
duration and low frequency, resulting in state chartering around equilibrium due to
on-off actuation. The difference between the channels appears in pulse phasing and
polarity, but the errors converge to small magnitudes in both channels.
Figure 8 demonstrates that after a short initial jump, the parameters change
slowly. The parameters 𝑠𝑖𝑗 are stabilized with a clear ordering, when 𝑠12, 𝑠22 remain
higher than 𝑠11, 𝑠21. The parameters 𝛽1, 𝛽2 (plotted × 10) typically decrease from
high initial values and then oscillate around small values. The parameters 𝜀1, 𝜀2 settle
down quickly and vary little over the whole interval. In both channels, the trajectories
are similar in shape and level, i.e., the settings converge to close steady values with
only minor drift.
Out-of-plane channel
In-plane channel
Fig. 9 – Evolution of the cumulative return during training (Case 2)
Figure 9 shows a typical evolution of the cumulative return during training. At
the beginning of training, returns have a large value with high variance. After a
while, the mean return improves monotonically and approaches zero. The critic’s
estimates of the return initially are biased, but once the policy forms, it tracks the
mean curve well. In the first plot, convergence is faster, while in the second one, it
takes longer training with deeper occasional drops. However, the mean return also
moves toward zero. After convergence, some episodes may still result in relatively
high negative returns, but the overall trend remains stable.
5. Robustness analysis
In practice, real model parameters may deviate from their nominal values, so we
assess the robustness of the SMC-based intelligent controllers against model
parameter variations and compare it with the robustness of the intelligent agents from
Ref. [17]. In this article, two RL-based agents for relative spacecraft control were
89
proposed that differ in their input vector. The first agent (IA1) receives the standard
state vector 𝑋𝑘, while the input of the second agent (IA2) is augmented with the
previous step control action, i.e.,[𝑋𝑘 , 𝑢𝑘−1]
𝑇.
Table 4 presents the description of test cases, which differ in the percentage
deviation of the considered parameter from its nominal value.
Table 4. Description of test cases
Symbol Nom. В. 1 В. 2 В. 3 В. 4 В. 5 В. 6
Description
Nominal
parameter
value
+10 % +20 % +30 % –10 % –20 % –30 %
Figure 10 shows the state variations for IA1, when SSC mass 𝑚𝑠 deviates
from the nominal one. As can be seen from these plots, the agent remains robust
when the spacecraft mass decreases, but control performance degrades as mass
increases (Case B.1). In Cases B.2 and B.3 the agent fails to maintain stability.
This occurs because a greater spacecraft mass requires larger control amplitudes
than those considered during training.
Fig. 10 – State variations for IA1 and different spacecraft masses
90
Figure 11 – State variations for IA2 and different spacecraft masses
Fig. 11 shows the state variations for IA2, when SSC mass 𝑚𝑠 deviates
from the nominal one. The agent remains robust, but control performance
degrades as mass increases (Case B.1). Unlike IA1, IA2 maintains stability in
Cases B.2 and B.3, although the performance degrades in these cases.
Figure 12 presents the state trajectories for the SMC-based agent and
various cases of the SSC masses. Stability is preserved in all cases with a
similar transient shape pattern. However, in some cases, the steady-state error
increases.
91
Fig. 12 – State variations for SMC-based agent and different spacecraft masses.
6. Conclusions
This paper presents an approach to relative on-off spacecraft control that
combines sliding-mode control with reinforcement learning. The results confirm
the potential of this approach for autonomous spacecraft control under constraints
and uncertainty. The trained agent demonstrated superior robustness to uncertainty
in plant parameters compared to the reported baselines, which is due to the strong
inherent robustness of sliding-mode control. However, a large number of episodes
is required to train the agent, which restricts the practical implementation of this
approach in its current form. Accordingly, future work should explore ways to
improve the training efficiency of the agent.
1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris
Management. arXiv:1901.11121, 2019. 15 p.
2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments.
Progress in Aerospace Sciences. 2019. Vol. 108. P. 32–120. https://doi.org/10.1016/j.paerosci.2019.01.004
3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach
to spacecraft attitude controller design. Automatika. 2016. No. 57. P. 212–220.
https://doi.org/10.7305/automatika.2016.07.618
4. Anthony T., Wie B., Carroll S. Pulse-Modulated Control Synthesis for a Flexible Spacecraft. Journal of
Guidance. 1989. Vol. 13(6). P. 1014–1022. https://doi.org/10.2514/3.20574
5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an Algorithm to Control the Angular Motion of
Spacecraft Equipped with an Aeromagnetic Deorbiting System. Eastern-European Journal of Enterprise
Technologies. 2020. 1(5(103)). P. 37–46. https://doi.org/10.15587/1729-4061.2020.192813
https://doi.org/10.1016/j.paerosci.2019.01.004
https://doi.org/10.7305/automatika.2016.07.618
https://doi.org/10.2514/3.20574
https://doi.org/10.15587/1729-4061.2020.192813
92
6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016. 800 p.
7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks.
Communications of the ACM. 2017. 60(6). P. 84–90. https://doi.org/10.1145/3065386
8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017.
31(16). P. 821–835. https://doi.org/10.1080/01691864.2017.1365009
9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving.
Electronic Imaging. 2017. Issue 19. P. 70–76. https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023
10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017.
550. P. 354–359. https://doi.org/10.1038/nature24270
11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and
control. Astrodynamics. 2019. 3. P. 287–299. https://doi.org/10.1007/s42064-018-0053-6
12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and
Technology (Космічна наука і технологія). 2021. 27(6/133). P. 38–52.
https://doi.org/10.15407/knit2021.06.038
13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking
maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. 18(7).
https://doi.org/10.2514/1.I010914
14. Gaudet B., Linares R., Furfaro R. Six Degree-of-Freedom Hovering using LIDAR Altimetry via
Reinforcement Meta-Learning. Acta Astronautica. 2020. 172. P. 90–99.
https://doi.org/10.1016/j.actaastro.2020.03.026
15. Gaudet B., Linares R., Furfaro R. Seeker based Adaptive Guidance via Reinforcement Meta-Learning Applied
to Asteroid Close Proximity Operations. Acta Astronautica. 2020. 171. P. 1–13.
https://doi.org/10.1016/j.actaastro.2020.02.036
16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object
via deep learning. Space Science and Technology (Космічна наука і технологія). 2022. 28(5/138). P. 15–26.
https://doi.org/10.15407/knit2022.05.015
17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and
Technology (Космічна наука і технологія). 2024. 30(2/147). P. 3–14.
https://doi.org/10.15407/knit2024.02.003
18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. innov.
(Наука та інновації). 2018. 14(4). P. 5–16. https://doi.org/10.15407/scine14.04.005
19. Steinberger M., Horn M., Fridman L. (eds). Variable-Structure Systems and Sliding-Mode Control. Springer-
Verlag, London, 2020. (Studies in Systems, Decision and Control; Vol. 271). https://doi.org/10.1007/978-3-
030-36621-6
20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Washington:
Hemisphere Publishing, 1975. P. 224–235.
21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. P. 47–65.
22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms.
arXiv:1707.06347, 2017. 13 p.
23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep
Reinforcement Learning. arXiv:1602.01783, 2016.
Received on October 14, 2025,
in final form on December 3, 2025
https://doi.org/10.1145/3065386
https://doi.org/10.1080/01691864.2017.1365009
https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023
https://doi.org/10.1038/nature24270
https://doi.org/10.1007/s42064-018-0053-6
https://doi.org/10.15407/knit2021.06.038
https://doi.org/10.2514/1.I010914
https://doi.org/10.1016/j.actaastro.2020.03.026
https://doi.org/10.1016/j.actaastro.2020.02.036
https://doi.org/10.15407/knit2022.05.015
https://doi.org/10.15407/knit2024.02.003
https://doi.org/10.15407/scine14.04.005
https://doi.org/10.1007/978-3-030-36621-6
https://doi.org/10.1007/978-3-030-36621-6
|
| id | oai:ojs2.journal-itm.dp.ua:article-157 |
| institution | Technical Mechanics |
| keywords_txt_mv | keywords |
| language | English |
| last_indexed | 2026-07-14T01:00:44Z |
| publishDate | 2025 |
| publisher | текст 3 |
| record_format | ojs |
| resource_txt_mv | journal-itmdpua/56/ead0043608f8761768e2ce5fbc315156.pdf |
| spelling | oai:ojs2.journal-itm.dp.ua:article-1572026-07-13T20:26:18Z ON-OFF SPACECRAFT RELATIVE CONTROL IN SLIDING MODE VIA REINFORCEMENT LEARNING ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ SOROCHINSKII, V. V. KHOROSHYLOV, S. V. LEVCHUK, I. L. DUBOVYK, T. M. HUZ, H. M. ROMANCHUK, O. O. навчання з підкріпленням, проксимальна оптимізація політики, керування косміч¬ним апаратом, орбітальні сервісні операції, on-off керування, автономні системи керування. reinforcement learning, proximal policy optimization, spacecraft control, on-orbit servicing, on–off control, autonomous control systems. DOI: https://doi.org/10.15407/itm2025.04.077 The paper addresses the problem of on-off spacecraft relative control in sliding mode for autonomous on-orbit servicing operations under actuator amplitude limits, action discreteness, and parametric uncertainties. The goal is to develop and assess an approach that combines sliding-mode control with modern reinforcement-learning methods tailored for resource-constrained onboard implementation. Relative motion dynamics is formulated in an orbital coordinate frame with normalized states and discretized in time. Binary actions with pulse-width modulation, subject to constraints on the thrust level, pulse duration, and duty cycle, represent the impulsive nature of actuation. We propose a combined synthesis in which the sliding-surface parameters and switching rules are tuned via proximal policy optimization within an actor-critic architecture. The actor and critic are implemented as neural networks that approximate the policy and the value function, respectively. The actor neural network takes the state vector as input information and outputs the mean and standard deviation of the parameters of the sliding mode control law. The value function penalizes both the state error and control effort, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. Two uncoupled agents are designed to control spacecraft relative orbital motion in in-plane and out-of-plane directions independently. The proximal policy optimization hyperparameters are selected to ensure a trade-off among the learning time, stability, and control performance. The reinforcement-learning agents are trained and analyzed considering four cases that differ in the thrust levels and weighting matrices. The quality functional combines state deviation and thrust use penalties, thus enabling a trade-off among the response speed, accuracy, and propellant consumption. The results confirm the potential of this approach for autonomous spacecraft control under constraints and uncertainty. Compared with reported baselines, the trained agent shows superior robustness to plant-parameter uncertainty, which we attribute to the inherent robust properties of sliding-mode control. These findings have the potential to improve the efficiency and autonomy of on-orbit servicing operations. REFERENCES 1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris Management. arXiv:1901.11121, 2019. 2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments. Progress in Aerospace Sciences. 2019. V. 108. Pp. 32-120.https://doi.org/10.1016/j.paerosci.2019.01.004 3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to spacecraft attitude controller design. Automatika. 2016. V. 57. Pp. 212-220.https://doi.org/10.7305/automatika.2016.07.618 4. Anthony T., Wie B., Carroll S. Pulse-modulated control synthesis for a flexible spacecraft. Journal of Guidance. 1989. V. 13. No. 6. Pp. 1014-1022. https://doi.org/10.2514/3.20574 5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an algorithm to control the angular motion of spacecraft equipped with an aeromagnetic deorbiting system. Eastern-European Journal of Enterprise Technologies. 2020. V. 1. No. 5. Pp. 37-46. https://doi.org/10.15587/1729-4061.2020.192813 6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016. 7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks. Communications of the ACM. 2017. V. 60. No. 6. Pp. 84-90. https://doi.org/10.1145/3065386 8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017. V. 31. No. 16. Pp. 821-835. https://doi.org/10.1080/01691864.2017.1365009 9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving. Electronic Imaging. 2017. V. 19. Pp. 70-76.https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023 10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017. V. 550. Pp. 354-359. https://doi.org/10.1038/nature24270 11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics. 2019. No. 3. Pp. 287-299. https://doi.org/10.1007/s42064-018-0053-6 12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and Technology. 2021. V. 27. No. 6. Pp. 38-52. https://doi.org/10.15407/knit2021.06.038 13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. V. 18. No. 7. https://doi.org/10.2514/1.I010914 14. Gaudet B., Linares R., Furfaro R. Six degree-of-freedom hovering using LIDAR altimetry via reinforcement meta-learning. Acta Astronautica. 2020. V. 172. Pp. 90-99.https://doi.org/10.1016/j.actaastro.2020.03.026 15. Gaudet B., Linares R., Furfaro R. Seeker based adaptive guidance via reinforcement meta-learning applied to asteroid close proximity operations. Acta Astronautica. 2020. V. 171. Pp. 1-13.https://doi.org/10.1016/j.actaastro.2020.02.036 16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object via deep learning. Space Science and Technology. 2022. V. 28. No. 5. Pp. 15-26.https://doi.org/10.15407/knit2022.05.015 17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and Technology. 2024. V. 30. No. 2. Pp. 3-14. https://doi.org/10.15407/knit2024.02.003 18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. Innov. 2018. V. 14. No. 4. Pp. 5-16. https://doi.org/10.15407/scine14.04.005 19. Steinberger M., Horn M., Fridman L. (Eds). Variable-Structure Systems and Sliding-Mode Control. Springer-Verlag, 2020. (Studies in Systems, Decision and Control; Vol. 271).https://doi.org/10.1007/978-3-030-36621-6 20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Hemisphere Publishing, 1975. Pp. 224-235. 21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. Pp. 47-65. 22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. 13 pp. 23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep Reinforcement Learning. arXiv:1602.01783, 2016. DOI: https://doi.org/10.15407/itm2025.04.077 Розглянуто задачу відносного імпульсного керування рухом космічного апарата у ковзному режимі для автономних орбітальних сервісних операцій за наявності обмежень на амплітуду керуючих впливів, дискретності дій та параметричних невизначеностей. Метою роботи є розробка й оцінювання підходу, що поєднує принципи ковзного керування з сучасними методами навчання з підкріпленням, орієнтованими на бортову реалізацію з обмеженими ресурсами. Динаміку відносного руху задано в орбітальній системі координат у нормалізованих змінних і дискредитовано. Імпульсний характер впливів виконавчих органів відображено через бінарні дії з широтно-імпульсною модуляцією та обмеженнями на рівень тяги, тривалість і період увімкнень. Запропоновано комбінований синтез, у якому параметри поверхні ковзання та правила перемикання налаштовуються методом проксимальної оптимізації політики з використанням архітектури актор-критик. Актор і критик реалізовані у вигляді нейронних мереж, які відповідно апроксимують політику та функцію цінності. Нейронна мережа актора приймає вектор стану як вхідну інформацію і видає середнє значення та стандартне відхилення параметрів закону керування у ковзному режимі. Функція цінності штрафує як за помилку стану, так і за витрати на керування, що дозволяє забезпечити компроміс між швидкістю реагування, точністю та витратою палива. Два незалежні агенти розроблені для керування відносним орбітальним рухом космічного апарата окремо в напрямку площини орбіти та у перпендикулярному напрямку. Гіперпараметри оптимізації проксимальної політики обрано для забезпечення компромісу між часом навчання, стабільністю та якістю керування. Агенти навчання з підкріпленням навчeні та проаналізовані з урахуванням чотирьох випадків, що відрізняються рівнями тяги та ваговими матрицями. Функціонал якості об’єднує штрафи за відхилення стану та використання тяги, що дає змогу знаходити компроміс між швидкодією, точністю та витратами робочого тіла. Отримані результати підтверджують потенціал такого підходу для задач автономного керування космічних апаратів в умовах обмежень та невизначеності. У порівнянні з відомими результатами навчений агент продемонстрував кращу робастність по відношенню до невизначенності параметрів моделі об’єкта керування, що пояснюється сильними робастними властивостями керування в ковзному режимі. Отримані результати мають потенціал підвищити ефективність та автономність орбітальних сервісних операцій. ПОСИЛАННЯ 1. Chandra A., Kalita H., Furfaro R., Thangavelautham J. End to End Satellite Servicing and Space Debris Management. arXiv:1901.11121, 2019. 15 p. 2. Li W., Cheng D., Liu X., et al. On-orbit service (OOS) of spacecraft: A review of engineering developments. Progress in Aerospace Sciences. 2019. Vol. 108. P. 32–120. https://doi.org/10.1016/j.paerosci.2019.01.004 3. Khosravi A., Sarhadi P. Tuning of pulse-width pulse-frequency modulator using PSO: An engineering approach to spacecraft attitude controller design. Automatika. 2016. No. 57. P. 212–220. https://doi.org/10.7305/automatika.2016.07.618 4. Anthony T., Wie B., Carroll S. Pulse-Modulated Control Synthesis for a Flexible Spacecraft. Journal of Guidance. 1989. Vol. 13(6). P. 1014–1022. https://doi.org/10.2514/3.20574 5. Alpatov A., Khoroshylov S., Lapkhanov E. Synthesizing an Algorithm to Control the Angular Motion of Spacecraft Equipped with an Aeromagnetic Deorbiting System. Eastern-European Journal of Enterprise Technologies. 2020. 1(5(103)). P. 37–46. https://doi.org/10.15587/1729-4061.2020.192813 6. Goodfellow I., Bengio Y., Courville A. Deep Learning. MIT Press, 2016. 800 p. 7. Krizhevsky A., Sutskever I., Hinton G. E. ImageNet classification with deep convolutional neural networks. Communications of the ACM. 2017. 60(6). P. 84–90. https://doi.org/10.1145/3065386 8. Pierson H., Gashler M. Deep learning in robotics: a review of recent research. Advanced Robotics. 2017. 31(16). P. 821–835. https://doi.org/10.1080/01691864.2017.1365009 9. Sallab A. E., Abdou M., Perot E., Yogamani S. Deep reinforcement learning framework for autonomous driving. Electronic Imaging. 2017. Issue 19. P. 70–76. https://doi.org/10.2352/ISSN.2470-1173.2017.19.AVM-023 10. Silver D., Schrittwieser J., Simonyan K. Mastering the game of Go without human knowledge. Nature. 2017. 550. P. 354–359. https://doi.org/10.1038/nature24270 11. Izzo D., Märtens M., Pan B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics. 2019. 3. P. 287–299. https://doi.org/10.1007/s42064-018-0053-6 12. Khoroshylov S. V., Redka M. O. Deep learning for space guidance, navigation, and control. Space Science and Technology (Космічна наука і технологія). 2021. 27(6/133). P. 38–52. https://doi.org/10.15407/knit2021.06.038 13. Oestreich C. E., Linares R., Gondhalekar R. Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning. Journal of Aerospace Information Systems. 2021. 18(7). https://doi.org/10.2514/1.I010914 14. Gaudet B., Linares R., Furfaro R. Six Degree-of-Freedom Hovering using LIDAR Altimetry via Reinforcement Meta-Learning. Acta Astronautica. 2020. 172. P. 90–99. https://doi.org/10.1016/j.actaastro.2020.03.026 15. Gaudet B., Linares R., Furfaro R. Seeker based Adaptive Guidance via Reinforcement Meta-Learning Applied to Asteroid Close Proximity Operations. Acta Astronautica. 2020. 171. P. 1–13. https://doi.org/10.1016/j.actaastro.2020.02.036 16. Redka M. O., Khoroshylov S. V. Determination of the force impact of an ion thruster plume on an orbital object via deep learning. Space Science and Technology (Космічна наука і технологія). 2022. 28(5/138). P. 15–26. https://doi.org/10.15407/knit2022.05.015 17. Khoroshylov S. V., Wang C. Spacecraft relative on-off control via reinforcement learning. Space Science and Technology (Космічна наука і технологія). 2024. 30(2/147). P. 3–14. https://doi.org/10.15407/knit2024.02.003 18. Khoroshylov S. V. Relative motion control system of SC for contactless space debris removal. Sci. innov. (Наука та інновації). 2018. 14(4). P. 5–16. https://doi.org/10.15407/scine14.04.005 19. Steinberger M., Horn M., Fridman L. (eds). Variable-Structure Systems and Sliding-Mode Control. Springer-Verlag, London, 2020. (Studies in Systems, Decision and Control; Vol. 271). https://doi.org/10.1007/978-3-030-36621-6 20. Bryson A. E., Ho Y. C. Applied Optimal Control: Optimization, Estimation, and Control. Washington: Hemisphere Publishing, 1975. P. 224–235. 21. Sutton R. S., Barto A. G. Reinforcement Learning: An Introduction. 2nd ed. MIT Press, 2018. P. 47–65. 22. Schulman J., Wolski F., Dhariwal P., Radford A., Klimov O. Proximal Policy Optimization Algorithms. arXiv:1707.06347, 2017. 13 p. 23. Mnih V., Badia A., Mirza M., Graves A., Lillicrap T., Harley T., Silver D. Asynchronous Methods for Deep Reinforcement Learning. arXiv:1602.01783, 2016. текст 3 2025-12-11 Article Article application/pdf https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157 Technical Mechanics; No. 4 (2025): Technical Mechanics; 77-92 Институт технической механики Национальной академии наук Украины и Государственного космического агентства Украины; № 4 (2025): Technical Mechanics; 77-92 ТЕХНІЧНА МЕХАНІКА; № 4 (2025): ТЕХНІЧНА МЕХАНІКА; 77-92 en https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157/66 Copyright (c) 2025 Technical Mechanics |
| spellingShingle | навчання з підкріпленням проксимальна оптимізація політики керування косміч¬ним апаратом орбітальні сервісні операції on-off керування автономні системи керування. SOROCHINSKII, V. V. KHOROSHYLOV, S. V. LEVCHUK, I. L. DUBOVYK, T. M. HUZ, H. M. ROMANCHUK, O. O. ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ |
| title | ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ |
| title_alt | ON-OFF SPACECRAFT RELATIVE CONTROL IN SLIDING MODE VIA REINFORCEMENT LEARNING |
| title_full | ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ |
| title_fullStr | ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ |
| title_full_unstemmed | ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ |
| title_short | ІМПУЛЬСНЕ КЕРУВАННЯ ВІДНОСНИМ РУХОМ КОСМІЧНИХ АПАРАТІВ У КОВЗНОМУ РЕЖИМІ З ВИКОРИСТАННЯМ НАВЧАННЯ З ПІДКРІПЛЕННЯМ |
| title_sort | імпульсне керування відносним рухом космічних апаратів у ковзному режимі з використанням навчання з підкріпленням |
| topic | навчання з підкріпленням проксимальна оптимізація політики керування косміч¬ним апаратом орбітальні сервісні операції on-off керування автономні системи керування. |
| topic_facet | навчання з підкріпленням проксимальна оптимізація політики керування косміч¬ним апаратом орбітальні сервісні операції on-off керування автономні системи керування. reinforcement learning proximal policy optimization spacecraft control on-orbit servicing on–off control autonomous control systems. |
| url | https://journal-itm.dp.ua/ojs/index.php/ITM_j1/article/view/157 |
| work_keys_str_mv | AT sorochinskiivv onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning AT khoroshylovsv onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning AT levchukil onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning AT dubovyktm onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning AT huzhm onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning AT romanchukoo onoffspacecraftrelativecontrolinslidingmodeviareinforcementlearning AT sorochinskiivv ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm AT khoroshylovsv ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm AT levchukil ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm AT dubovyktm ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm AT huzhm ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm AT romanchukoo ímpulʹsnekeruvannâvídnosnimruhomkosmíčnihaparatívukovznomurežimízvikoristannâmnavčannâzpídkríplennâm |