Optimization of Competitive Strategies for Humanoid Fighting Robots Based on Dynamic Evaluation, Game Theory and Sequential Decision Making

Authors

  • Wenqi Jiang North China University of Technology, Beijing 100144, China

DOI:

https://doi.org/10.6911/

Keywords:

Humanoid robot,Newton-Euler dynamics, ZMP stability,TOPSIS, zero-sum game,Q-learning, stochastic dynamic programming.

Abstract

Combat-related research can provide a compact research benchmark for whole body dynamics, contact stability, and tactical decision-making. This paper has adjusted the mathematical modeling report of the PM01 humanoid robot competition to CPMI-style engineering research content. This paper proposes a four-layer strategic optimization framework that can cover different decision-making aspects of the robot fighting process. First of all, this paper will use Newton- Euler recursive dynamics, extended zero-point stability analysis, and seven standard decision matrices to evaluate the existing 13 attack actions one by one to ensure that the actual performance of each action can be accurately measured. Before using the AHP entropy-weighted TOPSIS model, this paper did the work of Pareto pre-screening, and also introduced the Mahanobis distance to reduce the computational bias caused by the relevant indicators, so that the final evaluation results are more relevant to the actual situation. Next, to deal with defense-related issues, this article uses the 13- by 22 attack defense effectiveness matrix to describe the correspondence between the attack and defense, and then transforms the problem into a limited zero-sum game to solve, and find a relatively reasonable defense strategy. Then there are the issues related to single-pos motion switching, which is described in this article as a limited Hoylon Markov decision-making process, and then training through the Q-learning method in self-game mode, so that the logic of action switching is more suitable for real fighting scenes. Finally, this article uses random dynamic programs to deal with the best three resource scheduling problems in random failure scenarios to reduce the impact of unexpected failures on fighting performance. The results obtained in this article show that the combination of punch, side kick and boxing and kick combination of these three actions, got the highest attack utility, scoring 0.812, 0.784 and 0.761, respectively, all the best performance in all attack actions. Nash's mixed-defensive strategy got 0.76 game points, and the training of a single-game strategy, in 10,000 Monte Carlo simulation tests, won 68.7%, outperforming many conventional decision-making strategies. The research in this paper also found that practical robot fighting strategies cannot rely on a single indicator to select actions, but also to optimize the factors of impact, poor stability, continuity, recovery ability and resource status, in order to select the most suitable action for the current scene.

Downloads

Download data is not yet available.

References

[1] Spong, M. W., Hutchinson, S., & Vidyasagar, M. (2006). Robot modeling and control. Wiley.

[2] Craig, J. J. (2005). Introduction to robotics: Mechanics and control (3rd ed.). Pearson.

[3] Luh, J. Y. S., Walker, M. W., & Paul, R. P. C. (1980). On-line computational scheme for mechanical manipulators. Journal of Dynamic Systems, Measurement, and Control, 102(2), 69–76.

[4] Vukobratovic, M., & Borovac, B. (2004). Zero-moment point - thirty five years of its life. International Journal of Humanoid Robotics, 1(1), 157–173.

[5] Kajita, S., Kanehiro, F., Kaneko, K., Fujiwara, K., Harada, K., Yokoi, K., & Hirukawa, H. (2003). Biped walking pattern generation by using preview control of zero-moment point. In Proceedings of the IEEE International Conference on Robotics and Automation (pp. 1620–1626).

[6] Siciliano, B., Sciavicco, L., Villani, L., & Oriolo, G. (2009). Robotics: Modelling, planning and control. Springer.

[7] Saaty, T. L. (1980). The analytic hierarchy process. McGraw-Hill.

[8] Hwang, C. L., & Yoon, K. (1981). Multiple attribute decision making: Methods and applications. Springer.

[9] Yager, R. R. (1988). On ordered weighted averaging aggregation operators in multicriteria decision making. IEEE Transactions on Systems, Man, and Cybernetics, 18(1), 183–190.

[10] Mahalanobis, P. (1936). On the generalized distance in statistics. Proceedings of the National Institute of Sciences of India, 2(1), 49–55.

[11] Nash, J. (1951). Non-cooperative games. Annals of Mathematics, 54(2), 286–295.

[12] Osborne, M. J., & Rubinstein, A. (1994). A course in game theory. MIT Press.

[13] Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4), 279–292.

[14] Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

[15] Bertsekas, D. P. (2017). Dynamic programming and optimal control (4th ed.). Athena Scientific.

[16] Puterman, M. L. (1994). Markov decision processes: Discrete stochastic dynamic programming. Wiley.

[17] Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., van den Driessche, G., Schrittwieser, J., Antonoglou, I., Veda, V., Silver, T., Kummaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2016). Mastering the game of Go with deep neural networks and tree search. Nature, 529, 484–489.

[18] Kober, J., Bagnell, J. A., & Peters, J. (2013). Reinforcement learning in robotics: A survey. International Journal of Robotics Research, 32(11), 1238–1274.

Downloads

Published

2026-09-17

Issue

Section

Articles

How to Cite

Jiang, W. (2026). Optimization of Competitive Strategies for Humanoid Fighting Robots Based on Dynamic Evaluation, Game Theory and Sequential Decision Making. World Scientific Research Journal, 12(10), 7-15. https://doi.org/10.6911/