
浏览全部资源
扫码关注微信
1. 海军航空大学, 山东 烟台 264001
2. 91550部队, 辽宁 大连 116041
Received:22 December 2022,
Published Online:12 December 2023,
Published:30 November 2023
移动端阅览
Wenfei ZHAO, Jian CHEN, Yan WANG, et al. Dynamic Firepower Allocation for Cooperative Air Defense of Strategic Locations on the Sea Based on Reinforcement Learning[J]. Acta Armamentarii, 2023, 44(11): 3516-3528.
Wenfei ZHAO, Jian CHEN, Yan WANG, et al. Dynamic Firepower Allocation for Cooperative Air Defense of Strategic Locations on the Sea Based on Reinforcement Learning[J]. Acta Armamentarii, 2023, 44(11): 3516-3528. DOI: 10.12382/bgxb.2022.1276.
针对海上要地群协同防空作战动态火力分配问题
综合分析海上要地防空作战过程的特点
建立基于马尔可夫决策模型的动态火力分配问题
构建以海上要地毁伤期望、拦截成本为指标的优化模型。考虑到马尔可夫决策模型求解易陷入维数灾难的问题
提出利用近似动态规划方法来探究解的有效性
并给出基于强化学习的最小二乘时序差分算法来求解该问题。通过4种典型的攻防场景共80个案例仿真结果表明
相比传统的匹配算法、遗传算法和粒子群优化算法
新构建的模型和算法更加科学合理有效
可为海上要地群协同防空作战火力分配提供一定的理论依据。
For the dynamic firepower allocation in the cooperative air defense operation of strategic locations on the sea
the characteristics of air defense operations in strategic locations on the sea are comprehensively analyzed to establish the dynamic firepower allocation problem based on the Markov decision model
and an optimization model with the damage expectation and interception cost as the indexes is constructed. Considering the problem that the Markov decision model is easy to fall into the disaster of dimensionality
an approximate dynamic programming method is proposed to explore the validity of the solution
and a least squares temporal difference algorithm based on reinforcement learning is given to solve the problem. The simulated results of 80 cases in four typical offensive and defensive scenarios show that
compared with the traditional matching algorithm
genetic algorithm and particle swarm optimization algorithm
the proposed model and algorithmin this paper are more scientific
reasonable and effective
which can provide a certain basis for the firepower allocation in the cooperative air defense operations of strategic locations on the sea.
马新星 , 滕克难 , 侯学隆 . 岛礁防空火力单元配置距离计算模型 [J ] . 兵工自动化 , 2017 , 36 ( 10 ): 38 - 41 .
MA X X , TENG K N , HOU X L . Models to calculate the deployment distance of reef air defense fire units [J ] . Ordnance Industry Automation , 2017 , 36 ( 10 ): 38 - 41 . (in Chinese)
韩锋 , 陈岗 . 岛礁防空的特点和对策 [C ] // 第四届中国指挥控制大会 , 北京 : 中国指挥与控制学会 , 2016 : 332 - 336 .
HAN F , CHEN G . Characters and countermeasures of island air defense [C ] // Proceedings of the 4th China command and Control Conference . Beijing : Chinsee Institute of Command and Control , 2016 : 332 - 336 . (in Chinese)
LLOYD S P , WITSENHAUSEN H S . Weapons allocation is NP-complete [C ] // Proceedings of the 1986 Summer Simulation Conference . Reno, NV, US : IEEE , 1986 : 1054 - 1058 .
沈培志 , 杨历彪 , 王培源 . 多火力单元部署的地空导弹防空体系作战效能研究 [J ] . 火力与指挥控制 , 2020 , 45 ( 1 ): 1 - 6 .
SHEN P Z , YANG L B , WANG P Y . Research on operational effectiveness of surface-to-air missile air defense system deployed in multiple fire units [J ] . Fire Control & Command Control , 2020 , 45 ( 1 ): 1 - 6 . (in Chinese)
KARASAKAI O . Air defense missile-target allocation models for a naval task group [J ] . Computers & Operations Research , 2008 , 35 ( 6 ): 1759 - 1770 . DOI: 10.1016/j.cor.2006.09.011 http://doi.org/10.1016/j.cor.2006.09.011 https://linkinghub.elsevier.com/retrieve/pii/S030505480600222X https://linkinghub.elsevier.com/retrieve/pii/S030505480600222X
王小艺 , 侯朝桢 , 原菊梅 . 防空火力分配建模及优化方法研究 [J ] . 控制与决策 , 2006 , 21 ( 8 ): 913 - 917 .
WANG X Y , HOU C Z , YUAN J M . Modeling and optimization method on antiaircraft firepower allocation [J ] . Control and Decision , 2006 , 21 ( 8 ): 913 - 917 . (in Chinese)
王洁 , 娄寿春 , 王颖龙 . 防空导弹混合部署火力单元间配置距离的量化 [J ] . 系统工程与电子技术 , 2006 , 28 ( 2 ): 263 - 265 .
WANG J , LOU S C , WANG Y L . Quantitative analysis of deployment distance between fire units based on the composite disposition of the air defense missiles [J ] . Systems Engineering and Electronics , 2006 , 28 ( 2 ): 263 - 265 . (in Chinese)
JOHANSSON E , KARLSSON S . Deployment of air defense [D ] . Gothenburg, Sweden : Chalmers University of Technology , 2019 .
颜培远 , 刘曙 , 王君 . 基于动态规划-遗传算法的防空部署优化模型 [J ] . 系统工程与电子技术 , 2018 , 40 ( 10 ): 2249 - 2255 .
YAN P Y , LIU S , WANG J . Optimization model of air defense disposition based on dynamic programming and genetic algorithm [J ] . Systems Engineering and Electronics , 2018 , 40 ( 10 ): 2249 - 2255 . (in Chinese)
MANNE A S . A target-assignment problem [J ] . Operations Research , 1958 , 6 ( 3 ): 346 - 351 . DOI: 10.1287/opre.6.3.346 http://doi.org/10.1287/opre.6.3.346 https://pubsonline.informs.org/doi/10.1287/opre.6.3.346 https://pubsonline.informs.org/doi/10.1287/opre.6.3.346 This paper is concerned with a target assignment model of a probabilistic and nonlinear nature, but nevertheless one which is closely related to the “personnel-assignment” problem. It is shown here that, despite the apparent nonlinearities, it is possible to devise a linear programming formulation that will ordinarily provide a close approximation to the original problem.
WACHOLDER E . A neural network-based optimization algorithm for the static weapon-target assignment problem [J ] . Informs Journal on Computing , 1989 , 1 ( 4 ): 232 - 246 . DOI: 10.1287/ijoc.1.4.232 http://doi.org/10.1287/ijoc.1.4.232 https://pubsonline.informs.org/doi/10.1287/ijoc.1.4.232 https://pubsonline.informs.org/doi/10.1287/ijoc.1.4.232 A neural network-based algorithm was developed for the static weapon-target assignment problem in ballistic missile defense. An optimal assignment policy is one which allocates targets to weapon platforms such that the total expected leakage value of targets surviving the defense is minimized. This involves the minimization of a nonlinear objective function subject to inequality constraints specifying the maximum number of interceptors available to each platform and the maximum number of interceptors allowed to be fired at each target as imposed by the battle management/command control and communications system. The algorithm consists of solving a system of ordinary differential equations whose trajectories are the assignment variables of the problem. Simulations of the algorithm on PC and VAX computers were carried out using a simple numerical scheme. In all the battle instances tested, the algorithm has proven to be stable and to converge to solutions very close to global optima. The time to achieve convergence was consistently less than the time constant of the network's processing elements (neurons). This suggests that fast solutions can be realized if the algorithm is implemented in hardware circuits. Three series of battle scenarios are analyzed and discussed in detail. The main advantage of this algorithm is that it can be adapted to either a special-purpose hardware circuit or a general-purpose concurrent machine to yield fast and accurate solutions to difficult decision problems.
KWON O , KANG D , LEE K , et al . Lagrangian relaxation approach to the targeting problem [J ] . Naval Research Logistics , 1999 , 46 ( 6 ): 640 - 653 . DOI: 10.1002/(ISSN)1520-6750 http://doi.org/10.1002/(ISSN)1520-6750 http://doi.wiley.com/10.1002/%28ISSN%291520-6750 http://doi.wiley.com/10.1002/%28ISSN%291520-6750
AHUJA R K , KUMAR A , JHA K C , et al . Exact and heuristic algorithms for the weapon-target assignment problem [J ] . Operations Research , 2007 , 55 ( 6 ): 1136 - 1146 . DOI: 10.1287/opre.1070.0440 http://doi.org/10.1287/opre.1070.0440 https://pubsonline.informs.org/doi/10.1287/opre.1070.0440 https://pubsonline.informs.org/doi/10.1287/opre.1070.0440 The weapon-target assignment (WTA) problem is a fundamental problem arising in defense-related applications of operations research. This problem consists of optimally assigning n weapons to m targets so that the total expected survival value of the targets after all the engagements is minimal. The WTA problem can be formulated as a nonlinear integer programming problem and is known to be NP-complete. No exact methods exist for the WTA problem that can solve even small-size problems (for example, with 20 weapons and 20 targets). Although several heuristic methods have been proposed to solve the WTA problem, due to the absence of exact methods, no estimates are available on the quality of solutions produced by such heuristics. In this paper, we suggest integer programming and network flow-based lower-bounding methods that we obtain using a branch-and-bound algorithm for the WTA problem. We also propose a network flow-based construction heuristic and a very large-scale neighborhood (VLSN) search algorithm. We present computational results of our algorithms, which indicate that we can solve moderately large instances (up to 80 weapons and 80 targets) of the WTA problem optimally and obtain almost optimal solutions of fairly large instances (up to 200 weapons and 200 targets) within a few seconds.
KLINE A , AHNER D , PACHTER M . A greedy hungarian algorithm for the weapon- target assignment problem [R ] . OH, US: Air Force Institute of Technology, Air Force Institute of Technology Center for Operational Analysis , 2017 .
LI J , XIN B , PARDALOS P M , et al . Solving bi-objective uncertain stochastic resource allocation problems by the CVaR-based risk measure and decomposition-based multi-objective evolutionary algorithms [J ] . Annals of Operations Research , 2019 , 296 ( 1 ): 1 - 28 . DOI: 10.1007/s10479-020-03848-6 http://doi.org/10.1007/s10479-020-03848-6
LAI C M , WU T H . Simplified swarm optimization with initialization scheme for dynamic weapon-target assignment problem [J ] . Applied Soft Computing , 2019 , 82 ( 5 ): 1 - 15 .
JOSEPH M L , MATTEW J R , BRIAN J L . Improving defensive air battle management by solving a stochastic dynamic assignment problem via approximate dynamic programming [J ] . European Operational Research Societies , 2023 , 305 ( 3 ): 1435 - 1449 .
张安 , 徐双飞 , 毕文豪 , 等 . 空地多目标攻击武器-目标分配与制导序列优化 [J/OL ] . 兵工学报 : 1 - 12 [ 2023-03-25 ] . http://kns.cnki.net/kcms/detail/11.2176.TJ.20220829.1712.007.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20220829.1712.007.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20220829.1712.007.html.
ZHANG A , XU S F , BI W H , et al . Weapon-target assign- ment and guidance sequencc optimization in air-to-ground multi-target attack [J/OL ] . Acta Armamentarii : 1 - 12 [ 2023-03-25 ] . http://kns.cnki.net/kcms/detail/11.2176.TJ.20220829.1712.007.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20220829.1712.007.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20220829.1712.007.html. (in Chinese)
JENKINS P R , ROBBINS M J , LUNDAY B J . Approximate dynamic programming for the military aeromedical evacuation dispatching, preemption-rerouting, and redeployment problem [J ] . European Journal of Operational Research , 2021 , 290 ( 1 ): 132 - 143 . DOI: 10.1016/j.ejor.2020.08.004 http://doi.org/10.1016/j.ejor.2020.08.004 https://linkinghub.elsevier.com/retrieve/pii/S0377221720306949 https://linkinghub.elsevier.com/retrieve/pii/S0377221720306949
AHNER D K , PARSON C R . Optimal multi-stage allocation of weapons to targets using adaptive dynamic programming [J ] . Optimization Letters , 2015 , 9 ( 8 ): 1689 - 1701 . DOI: 10.1007/s11590-014-0823-x http://doi.org/10.1007/s11590-014-0823-x http://link.springer.com/10.1007/s11590-014-0823-x http://link.springer.com/10.1007/s11590-014-0823-x
王 , 赵文飞 , 滕克难 , 等 . 不确定因素下海上要地防空动态火力分配模型 [J ] . 兵工学报 , 2022 , 43 ( 11 ): 2885 - 2896 .
WANG Y , ZHAO W F , TENG K N , et al . DWTA model of air defense in important place at sea under uncertain factors [J ] . Acta Armamentarii , 2022 , 43 ( 11 ): 2885 - 2896 . (in Chinese)
赵文飞 , 刘孝磊 , 马翠玲 , 等 . 基于多目标模糊规划的海上要地动态火力分配 [J/OL ] . 系统工程与电子技术 , 2023 , 45 ( 3 ): 777 - 784 .
ZHAO W F , LIU X L , MA C L , et al . DWTA for strategic location on the sea based on multi-objective fuzzy programming [J ] . Systems Engineering and Electronics , 2023 , 45 ( 3 ): 777 - 784 . (in Chinese)
朱建文 , 赵长见 , 李小平 , 等 . 基于强化学习的集群多目标分配与智能决策方法 [J ] . 兵工学报 , 2021 , 42 ( 9 ): 2040 - 2048 .
ZHU J W , ZHAO C J , LI X P , et al . Multi-target assignment and intelligent decision based on reinforcement learning [J ] . Acta Armamentarii , 2021 , 42 ( 9 ): 2040 - 2048 . (in Chinese) DOI: 10.3969/j.issn.1000-1093.2021.09.025 http://doi.org/10.3969/j.issn.1000-1093.2021.09.025 A reinforcement learning-based swarm intelligent decision-making method of cooperative multi-target attack under high-dynamic situation is proposed. The composite evaluation criteria of attack performance is established, including the evaluation of attack superiority based on relative motion information and the threat evaluation based on the inherent information of target. To evaluate the attack-defence effectiveness, a cost-effectiveness ratio index is designed by combining attack performance, penetration probability and attack cost together. In addition, a multi-target decision-making architecture based on reinforcement learning is constructed, and an action space with allocation vectors as basic elements and a state space based on quantified performance indicators are designed. Q-Learning is employed to make intelligent decisions on cooperative attack plans, including missile selection and target assignment. The simulated results show that reinforcement learning can achieve multi-target online decision-making with the optimal offensive and defensive effectiveness, and its computational efficiency has more obvious advantages than that of particle swarm optimizer.
褚凯轩 , 常天庆 , 张雷 . 基于改进人工蜂群算法的地面作战武器-目标分配 [J/OL ] . 兵工学报 : 1 - 13 [ 2023-03-22 ] . http://kns.cnki.net/kcms/detail/11.2176.TJ.20230207.0851.001.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20230207.0851.001.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20230207.0851.001.html.
CHU K X , CHANG T Q , ZHANG L . A WTA model based on shooting effectiveness.and improved artificial bee colony based on WTP library [J/OL ] . Acta Armamentarii : 1 - 13 [ 2023-03-22 ] . http://kns.cnki.net/kcms/detail/11.2176.TJ.20230207.0851.001.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20230207.0851.001.html http://kns.cnki.net/kcms/detail/11.2176.TJ.20230207.0851.001.html. (in Chinese)
SUMMERS D S , ROBBINS M J , LUNDAY B J . An approximate dynamic programming approach for comparing firing policies in a networked air defense environment-science direct [J ] . Computers & Operations Research , 2020 , 117 ( 5 ): 1 - 15 .
DAVIS M T , ROBBINS M J , LUNDAY B J . Approximate dynamic programming for missile defense interceptor fire control [J ] . European Journal of Operational Research , 2016 , 259 ( 3 ): 873 - 886 . DOI: 10.1016/j.ejor.2016.11.023 http://doi.org/10.1016/j.ejor.2016.11.023 https://linkinghub.elsevier.com/retrieve/pii/S0377221716309481 https://linkinghub.elsevier.com/retrieve/pii/S0377221716309481
谢俊洁 . 空战仿真中的目标分配与火力分配方法 [D ] . 长沙 : 国防科学技术大学 , 2016 .
XIE J J . Target assignment and weapon-target assignment algorithms in air combat simulations [D ] . Changsha : National University of Defense Technology , 2016 . (in Chinese)
夏维 , 刘新学 , 范阳涛 , 等 . 基于改进型多目标粒子群优化算法的武器-目标分配 [J ] . 兵工学报 , 2016 , 37 ( 11 ): 2085 - 2093 . DOI: 10.3969/j.issn.1000-1093.2016.11.017 http://doi.org/10.3969/j.issn.1000-1093.2016.11.017 在作战中武器-目标分配(WTA)问题包含众多的变量,是典型的非确定性多项式完全问题。针对毁伤效能最大和用弹量最少两个目标函数,建立了基于改进型多目标粒子群优化(MOPSO-Ⅱ)算法的WTA模型。由于粒子群优化算法存在“维数灾难”瓶颈,应用了变量随机分解策略和合作协同进化框架,按照带精英策略的非支配排序遗传(NSGA-Ⅱ)算法中的排序方法对粒子群编码数据进行非支配排序。通过实例仿真分析,结果表明MOPSO-Ⅱ算法比NSGA-Ⅱ算法具有更好的求解精度与运行效率,能够获得满意的分配结果,且计算快速有效,比较适合较大规模的WTA问题实时求解。在作战中武器-目标分配(WTA)问题包含众多的变量,是典型的非确定性多项式完全问题。针对毁伤效能最大和用弹量最少两个目标函数,建立了基于改进型多目标粒子群优化(MOPSO-Ⅱ)算法的WTA模型。由于粒子群优化算法存在“维数灾难”瓶颈,应用了变量随机分解策略和合作协同进化框架,按照带精英策略的非支配排序遗传(NSGA-Ⅱ)算法中的排序方法对粒子群编码数据进行非支配排序。通过实例仿真分析,结果表明MOPSO-Ⅱ算法比NSGA-Ⅱ算法具有更好的求解精度与运行效率,能够获得满意的分配结果,且计算快速有效,比较适合较大规模的WTA问题实时求解。
XIA W , LIU X X , FAN Y T , et al . Weapon target assignment with an improved multi-objective particle swarm optimization algorithm [J ] . Acta Armamentarii , 2016 , 37 ( 11 ): 2085 - 2093 . (in Chinese) DOI: 10.3969/j.issn.1000-1093.2016.11.017 http://doi.org/10.3969/j.issn.1000-1093.2016.11.017 Weapon-target assignment (WTA) with numerous variables in modern campaign is a typical non-deterministic polynomial (NP) complete problem. An optimization model based on improved multi-objective swarm optimization algorithm (MOPSO-II) is established to solve the objective functions of maximum damage probability and minimum ammunition consumption. Since “curse of dimensionality” occurs in the objective swarm optimization algorithm (PSO), the random variable decomposition strategy and cooperative co-evolution evolutionary frame are used for variable decomposition, and also all swarms are composited by using the non-dominated set algorithm in NSGA-II. The simulated results show that MOPSO-II is quicker and more effective than NSGA-II, and can give good WTA quickly, especially when the scale of WTA problem is large.
SHERMAN J , MORRISON W J . Adjustment of an inverse matrix to changes in the elements of a given column or a given row in the original matrix [J ] . Annals of Mathematical Statistics , 1949 , 21 : 124 - 127 . DOI: 10.1214/aoms/1177729893 http://doi.org/10.1214/aoms/1177729893 http://projecteuclid.org/euclid.aoms/1177729893 http://projecteuclid.org/euclid.aoms/1177729893
0
Views
396
下载量
0
CNKI被引量
Publicity Resources
Related Articles
Related Author
Related Institution
京公网安备11010802024360号