An Adaptive Deep Reinforcement Learning Framework for Intelligent Transaction Execution in Internet Finance
DOI:
https://doi.org/10.6911/Keywords:
Algorithmic execution, internet finance, deep reinforcement learning, Double DQN, mixture of experts, market impact, risk control.Abstract
This study treats large-order execution as a short-horizon control problem in which the state can change while the order is still being worked. The proposed Adaptive Regime-Gated Double Dueling Deep Q-Network (Adaptive-RG-DDQN) combines a context-dependent mixture of dueling Q experts with two explicit risk mechanisms: an inventory penalty that increases with observed volatility and spread, and a deterministic shield that can revise a neural action when flow, liquidity, or time-to-deadline makes that action unacceptable. The numerical evidence is intentionally limited to simulation. We use a fully specified four-regime Markov market generator with square-root temporary impact, evaluate three independently trained seeds, and test all policies on 900 matched out-of-sample paths. Adaptive-RG-DDQN produces a mean implementation shortfall of 2.87 bps, compared with 3.83 bps for TWAP and 5.83 bps for static Double Dueling DQN. On the same market paths, the mean reductions are 0.95 bps versus TWAP (p=2.0e-6) and 2.96 bps versus static DDQN (p=2.4e-17). Its difference from an ex-post volume oracle is not significant. The ablation is more informative than the headline comparison: removing the shield eliminates most of the improvement, whereas the present experiment does not establish a separate gain from regime gating. The contribution is therefore a reproducible way to combine learned execution with inspectable risk controls, not a claim of live trading profitability.
Downloads
References
[1] Bertsimas, D., & Lo, A. W. (1998). Optimal control of execution costs. Journal of Financial Markets, 1(1), 1–50. https://doi.org/10.1016/S1386-4181(97)00012-8.
[2] Almgren, R., & Chriss, N. (2001). Optimal execution of portfolio transactions. Journal of Risk, 3(2), 5–39. https://doi.org/10.21314/JOR.2001.041
[3] Almgren, R. F. (2003). Optimal execution with nonlinear impact functions and trading-enhanced risk. Applied Mathematical Finance, 10(1), 1–18. https://doi.org/10.1080/135048602100056
[4] Gatheral, J. (2010). No-dynamic-arbitrage and market impact. Quantitative Finance, 10(7), 749–759. https://doi.org/10.1080/14697680903373692
[5] Almgren, R., & Lorenz, J. (2011). Mean-variance optimal adaptive execution. Applied Mathematical Finance, 18(5), 395–422. https://doi.org/10.1080/1350486X.2011.560707
[6] Gatheral, J., Schied, A., & Slynko, A. (2012). Transient linear price impact and Fredholm integral equations. Mathematical Finance, 22(3), 445–474. https://doi.org/10.1111/j.1467-9965.2011.00478.x
[7] Cont, R., Kukanov, A., & Stoikov, S. (2014). The price impact of order book events. Journal of Financial Econometrics, 12(1), 47–88. https://doi.org/10.1093/jjfinec/nbt003
[8] Gould, M. D., Porter, M. A., Williams, S., McDonald, M., Fenn, D. J., & Howison, S. D. (2013). Limit order books. Quantitative Finance, 13(11), 1709–1742. https://doi.org/10.1080/14697688.2013.803148
[9] Kyle, A. S. (1985). Continuous auctions and insider trading. Econometrica, 53(6), 1315–1335. https://doi.org/10.2307/1913210
[10] Hasbrouck, J. (1991). Measuring the information content of stock trades. The Journal of Finance, 46(1), 179–207. https://doi.org/10.1111/j.1540-6261.1991.tb03749.x
[11] Gabaix, X., Gopikrishnan, P., Plerou, V., & Stanley, H. E. (2006). Institutional investors and stock market volatility. The Quarterly Journal of Economics, 121(2), 461–504. https://doi.org/10.1162/qjec.2006.121.2.461
[12] Toth, B., Lemperiere, Y., Deremble, C., de Lataillade, J., Kockelkoren, J., & Bouchaud, J.-P. (2011). Anomalous price impact and the critical nature of liquidity in financial markets. Physical Review X, 1(2), 021006. https://doi.org/10.1103/PhysRevX.1.021006
[13] Eisler, Z., Bouchaud, J.-P., & Kockelkoren, J. (2012). The price impact of order book events: market orders, limit orders and cancellations. Quantitative Finance, 12(9), 1395–1419. https://doi.org/10.1080/14697688.2010.528444
[14] Moody, J., & Saffell, M. (2001). Learning to trade via direct reinforcement. IEEE Transactions on Neural Networks, 12(4), 875–889. https://doi.org/10.1109/72.935097
[15] Nevmyvaka, Y., Feng, Y., & Kearns, M. (2006). Reinforcement learning for optimized trade execution. Proceedings of the 23rd International Conference on Machine Learning, 673–680. https://doi.org/10.1145/1143844.1143929
[16] Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
[17] Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., Petersen, S., Beattie, C., Sadik, A., Antonoglou, I., King, H., Kumaran, D., Wierstra, D., Legg, S., & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518, 529–533. https://doi.org/10.1038/nature14236
[18] van Hasselt, H., Guez, A., & Silver, D. (2016). Deep reinforcement learning with Double Q-learning. Proceedings of the AAAI Conference on Artificial Intelligence, 30(1), 2094–2100. https://doi.org/10.1609/aaai.v30i1.10295
[19] Wang, Z., Schaul, T., Hessel, M., van Hasselt, H., Lanctot, M., & de Freitas, N. (2016). Dueling network architectures for deep reinforcement learning. Proceedings of the 33rd International Conference on Machine Learning, 48, 1995–2003. https://arxiv.org/abs/1511.06581
[20] Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2016). Prioritized experience replay. International Conference on Learning Representations. https://arxiv.org/abs/1511.05952
[21] Hessel, M., Modayil, J., van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2018). Rainbow: Combining improvements in deep reinforcement learning. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 3215–3222. https://doi.org/10.1609/aaai.v32i1.11796
[22] Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. https://arxiv.org/abs/1707.06347
[23] Kingma, D. P., & Ba, J. (2015). Adam: A method for stochastic optimization. International Conference on Learning Representations. https://arxiv.org/abs/1412.6980
[24] Ning, B., Lin, F. H. T., & Jaimungal, S. (2021). Double Deep Q-Learning for optimal execution. Applied Mathematical Finance, 28(4), 361–380. https://doi.org/10.1080/1350486X.2022.2077783
[25] Lin, S., & Beling, P. A. (2020). An end-to-end optimal trade execution framework based on proximal policy optimization. Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, 4548–4554. https://doi.org/10.24963/ijcai.2020/627
[26] Fang, Y., Ren, K., Liu, W., Zhou, D., Zhang, W., Bian, J., Yu, Y., & Liu, T.-Y. (2021). Universal trading for order execution with oracle policy distillation. Proceedings of the AAAI Conference on Artificial Intelligence, 35(1), 107–115. https://doi.org/10.1609/aaai.v35i1.16083
[27] Macrì, A., & Lillo, F. (2024). Reinforcement learning for optimal execution when liquidity is time-varying. Applied Mathematical Finance, 31(5), 312–342. https://doi.org/10.1080/1350486X.2025.2490157
[28] Ntakaris, A., Magris, M., Kanniainen, J., Gabbouj, M., & Iosifidis, A. (2018). Benchmark dataset for mid-price forecasting of limit order book data with machine learning methods. Journal of Forecasting, 37(8), 852–866. https://doi.org/10.1002/for.2543
[29] Zhang, Z., Zohren, S., & Roberts, S. (2019). DeepLOB: Deep convolutional neural networks for limit order books. IEEE Transactions on Signal Processing, 67(11), 3001–3012. https://doi.org/10.1109/TSP.2019.2907260
[30] Sirignano, J., & Cont, R. (2019). Universal features of price formation in financial markets: perspectives from deep learning. Quantitative Finance, 19(9), 1449–1459. https://doi.org/10.1080/14697688.2019.1622295
[31] Makarov, I., & Schoar, A. (2020). Trading and arbitrage in cryptocurrency markets. Journal of Financial Economics, 135(2), 293–319. https://doi.org/10.1016/j.jfineco.2019.07.001
[32] Cartea, Á., Jaimungal, S., & Penalva, J. (2015). Algorithmic and high-frequency trading. Cambridge University Press. https://doi.org/10.1017/CBO9781139340716
[33] Wang, B., Wang, Z., Zhao, W., Zhang, F., & Shang, W. (2026). DRL-Adapt: Deep reinforcement learning for adaptive routing convergence optimization in large-scale networks. IEEE Open Journal of the Computer Society, 7, 849–863. https://doi.org/10.1109/OJCS.2026.3687441
[34] Zhang, H., Ge, Y., Zhao, X., & Wang, J. (2025). Hierarchical deep reinforcement learning for multi-objective integrated circuit physical layout optimization with congestion-aware reward shaping. IEEE Access, 13, 162533–162551. https://doi.org/10.1109/ACCESS.2025.3610615
[35] Chen, Z., Wang, M., Zeng, Z., & Ping, W. (2026). Uncertainty-aware financial forecasting: Leveraging conformal prediction for risk-adjusted models. IEEE Access, 14, 90053–90077. https://doi.org/10.1109/ACCESS.2026.3702666
[36] Liang, Y., Jiao, Y., Ping, W., Fan, H., & Han, X. (2026). Adaptive event-driven labeling: A neuro-symbolic multiagent framework for causal inference in non-stationary time series. IEEE Access, 14, 104468–104482. https://doi.org/10.1109/ACCESS.2026.3709267
Downloads
Published
Issue
Section
License
Copyright (c) 2026 World Scientific Research Journal

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




