Data-driven sequential policy optimization from observational time series is difficult when latent states affect both treatment assignment and rewards and when overlap deteriorates over time. We develop temporally adaptive proximal learning under overlap extremes (Taple), a proximal policy learner that addresses this optimization problem through nested intervention-indexed bridge operators and temporally adaptive cumulative-weight capping. Its identifying functional avoids an invalid stagewise Bellman substitution and is stagewise doubly robust when the bridge choices are recursively compatible. For the proposed behavior-law implementation, learning a later policy-indexed outcome bridge from transported moments still requires valid preceding treatment bridges, unless another valid estimator of the relevant intervention-law moments is available. Under explicit conditions on bridge estimation, score entropy, policy indexing, numerical optimization, and cap calibration, we establish an oracle-calibrated fixed-horizon regret upper bound together with a minimax lower bound under polynomial correction-weight tails. When score-class and policy-class complexities are comparable and the nuisance and optimization remainders are of oracle order, the two bounds match in their overlap and policy-complexity exponents up to logarithmic and fixed-horizon factors. The executable factorized analogue, Taple-F, is deliberately restricted to an action-invariant factorized submodel and is not empirical validation of the full cumulative-score algorithm or its global optimization guarantee. Here "data-driven" means that bridge functions, caps, and policies are learned from observational data within an explicitly stated causal and statistical model; it does not mean model-free reinforcement learning. In controlled simulations, the locally capped factorized benchmark has lower mean regret than the corresponding proximal baseline without the adaptive local cap. A source-grouped PhysioNet semi-synthetic stress test shows only a modest local-capping improvement; with only nine training source records, observable-history doubly robust methods perform better in that experiment.
Citation: Siyang Bai, Zheng Fang, Jie Chen. Proximal policy learning with latent confounding and limited overlap: identification and oracle-calibrated regret bounds[J]. AIMS Mathematics, 2026, 11(9): 30581-30618. doi: 10.3934/math.20261212
Data-driven sequential policy optimization from observational time series is difficult when latent states affect both treatment assignment and rewards and when overlap deteriorates over time. We develop temporally adaptive proximal learning under overlap extremes (Taple), a proximal policy learner that addresses this optimization problem through nested intervention-indexed bridge operators and temporally adaptive cumulative-weight capping. Its identifying functional avoids an invalid stagewise Bellman substitution and is stagewise doubly robust when the bridge choices are recursively compatible. For the proposed behavior-law implementation, learning a later policy-indexed outcome bridge from transported moments still requires valid preceding treatment bridges, unless another valid estimator of the relevant intervention-law moments is available. Under explicit conditions on bridge estimation, score entropy, policy indexing, numerical optimization, and cap calibration, we establish an oracle-calibrated fixed-horizon regret upper bound together with a minimax lower bound under polynomial correction-weight tails. When score-class and policy-class complexities are comparable and the nuisance and optimization remainders are of oracle order, the two bounds match in their overlap and policy-complexity exponents up to logarithmic and fixed-horizon factors. The executable factorized analogue, Taple-F, is deliberately restricted to an action-invariant factorized submodel and is not empirical validation of the full cumulative-score algorithm or its global optimization guarantee. Here "data-driven" means that bridge functions, caps, and policies are learned from observational data within an explicitly stated causal and statistical model; it does not mean model-free reinforcement learning. In controlled simulations, the locally capped factorized benchmark has lower mean regret than the corresponding proximal baseline without the adaptive local cap. A source-grouped PhysioNet semi-synthetic stress test shows only a modest local-capping improvement; with only nine training source records, observable-history doubly robust methods perform better in that experiment.
| [1] |
S. A. Murphy, Optimal dynamic treatment regimes, J. R. Stat. Soc. B, 65 (2003), 331–355. https://doi.org/10.1111/1467-9868.00389 doi: 10.1111/1467-9868.00389
|
| [2] |
E. B. Laber, D. J. Lizotte, M. Qian, W. E. Pelham, S. A. Murphy, Dynamic treatment regimes: Technical challenges and applications, Electron. J. Statist., 8 (2014), 1225–1272. https://doi.org/10.1214/14-EJS920 doi: 10.1214/14-EJS920
|
| [3] |
Y. Zhao, D. Zeng, A. J. Rush, M. R. Kosorok, Estimating individualized treatment rules using outcome weighted learning, J. Am. Stat. Assoc., 107 (2012), 1106–1118. https://doi.org/10.1080/01621459.2012.695674 doi: 10.1080/01621459.2012.695674
|
| [4] |
N. T. Williams, K. L. Hoffman, I. Díaz, K. E. Rudolph, Learning optimal dynamic treatment regimes from longitudinal data, Am. J. Epidemiol., 193 (2024), 1768–1775. https://doi.org/10.1093/aje/kwae122 doi: 10.1093/aje/kwae122
|
| [5] |
S. Athey, S. Wager, Policy learning with observational data, Econometrica, 89 (2021), 133–161. https://doi.org/10.3982/ECTA15732 doi: 10.3982/ECTA15732
|
| [6] | H. Ye, W. Zhou, R. Zhu, A. Qu, Stage-aware learning for dynamic treatments, J. Mach. Learn. Res., 25 (2024), 1–51. |
| [7] |
X. Nie, E. Brunskill, S. Wager, Learning when-to-treat policies, J. Am. Stat. Assoc., 116 (2021), 392–409. https://doi.org/10.1080/01621459.2020.1831925 doi: 10.1080/01621459.2020.1831925
|
| [8] |
A. Bennett, N. Kallus, Proximal reinforcement learning: Efficient off-policy evaluation in partially observed Markov decision processes, Oper. Res., 72 (2024), 1071–1086. https://doi.org/10.1287/opre.2021.0781 doi: 10.1287/opre.2021.0781
|
| [9] | C. Kausik, Y. Lu, K. Tan, M. Makar, Y. Wang, A. Tewari, Offline policy evaluation and optimization under confounding, In: Proceedings of The 27th International Conference on Artificial Intelligence and Statistics, 238 (2024), 1459–1467. |
| [10] |
H. Zheng, D. Wang, A study of value iteration and policy iteration for Markov decision processes in deterministic systems, AIMS Mathematics, 9 (2024), 33818–33842. https://doi.org/10.3934/math.20241613 doi: 10.3934/math.20241613
|
| [11] | S. Yu, S. Fang, R. Peng, Z. Qi, F. Zhou, C. Shi, Two-way deconfounder for off-policy evaluation in causal reinforcement learning, In: Advances in Neural Information Processing Systems 37, 2024, 78169–78200. https://doi.org/10.52202/079017-2485 |
| [12] | A. Bennett, N. Kallus, M. Oprescu, W. Sun, K. Wang, Efficient and sharp off-policy evaluation in robust Markov decision processes, In: NIPS '24: Proceedings of the 38th International Conference on Neural Information Processing Systems, 2024, 112962–113000. |
| [13] |
P. Miao, A finite-time Q-Learning algorithm with finite-time constraints, AIMS Mathematics, 10 (2025), 23380–23393. https://doi.org/10.3934/math.20251038 doi: 10.3934/math.20251038
|
| [14] |
I. A. Zamfirache, R. E. Precup, E. M. Petriu, Q-learning, policy iteration and actor-critic reinforcement learning combined with metaheuristic algorithms in servo system control, Facta Univ. Ser. Mech. Eng., 21 (2023), 615–630. https://doi.org/10.22190/FUME231011044Z doi: 10.22190/FUME231011044Z
|
| [15] |
H. Wang, B. R. Sarker, J. Li, J. Li, Adaptive scheduling for assembly job shop with uncertain assembly times based on dual Q-learning, Int. J. Prod. Res., 59 (2021), 5867–5883. https://doi.org/10.1080/00207543.2020.1794075 doi: 10.1080/00207543.2020.1794075
|
| [16] |
F. Zhao, Z. Fu, L. Wang, H. Sang, A heterogeneous graph reinforcement learning framework with question-aware neighborhood aggregation and interoption prompt attention for dynamic flexible job shop scheduling problem, IEEE T. Ind. Inform., 22 (2026), 2863–2874. https://doi.org/10.1109/TII.2025.3646962 doi: 10.1109/TII.2025.3646962
|
| [17] |
Z. Pan, D. Lei, L. Wang, A knowledge-based two-population optimization algorithm for distributed energy-efficient parallel machines scheduling, IEEE T. Cybernetics, 52 (2022), 5051–5063. https://doi.org/10.1109/TCYB.2020.3026571 doi: 10.1109/TCYB.2020.3026571
|
| [18] |
Z.-L. Lu, U. H. Lok, Dimension-reduced modeling for local volatility surface via unsupervised learning, Rom. J. Inf. Sci. Technol., 27 (2024), 255–266. https://doi.org/10.59277/ROMJIST.2024.3-4.01 doi: 10.59277/ROMJIST.2024.3-4.01
|
| [19] |
S. Meng, K. Chhea, S. Muy, J. R. Lee, Deep reinforcement learning and metaheuristic approaches to maximize downlink sum-rate for Internet of Things systems in non-orthogonal multiple access-based space-air-ground integrated networks, Rom. J. Inf. Sci. Technol., 28 (2025), 327–340. https://doi.org/10.59277/ROMJIST.2025.4.02 doi: 10.59277/ROMJIST.2025.4.02
|
| [20] |
W. Miao, Z. Geng, E. J. Tchetgen Tchetgen, Identifying causal effects with proxy variables of an unmeasured confounder, Biometrika, 105 (2018), 987–993. https://doi.org/10.1093/biomet/asy038 doi: 10.1093/biomet/asy038
|
| [21] |
Z. Qi, R. Miao, X. Zhang, Proximal learning for individualized treatment regimes under unmeasured confounding, J. Am. Stat. Assoc., 119 (2024), 915–928. https://doi.org/10.1080/01621459.2022.2147841 doi: 10.1080/01621459.2022.2147841
|
| [22] |
T. Shen, Y. Cui, Optimal treatment regimes for proximal causal learning, Advances in Neural Information Processing Systems 36, 2023, 47735–47748. https://doi.org/10.52202/075280-2068 doi: 10.52202/075280-2068
|
| [23] | E. Sverdrup, Y. Cui, Proximal causal learning of conditional average treatment effects, In: Proceedings of the 40th International Conference on Machine Learning, 202 (2023), 33285–33298. |
| [24] |
V. Chernozhukov, D. Chetverikov, M. Demirer, E. Duflo, C. Hansen, W. Newey, J. Robins, Double/debiased machine learning for treatment and structural parameters, Econ. J., 21 (2018), C1–C68. https://doi.org/10.1111/ectj.12097 doi: 10.1111/ectj.12097
|
| [25] | N. Kallus, M. Uehara, Double reinforcement learning for efficient off-policy evaluation in Markov decision processes, J. Mach. Learn. Res., 21, (2020), 1–63. |
| [26] |
R. K. Crump, V. J. Hotz, G. W. Imbens, O. A. Mitnik, Dealing with limited overlap in estimation of average treatment effects, Biometrika, 96 (2009), 187–199. https://doi.org/10.1093/biomet/asn055 doi: 10.1093/biomet/asn055
|
| [27] |
A. D'Amour, P. Ding, A. Feller, L. Lei, J. Sekhon, Overlap in observational studies with high-dimensional covariates, J. Econ., 221 (2021), 644–654. https://doi.org/10.1016/j.jeconom.2019.10.014 doi: 10.1016/j.jeconom.2019.10.014
|
| [28] |
N. Kallus, M. Uehara, Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning, Oper. Res., 70 (2022), 3282–3302. https://doi.org/10.1287/opre.2021.2249 doi: 10.1287/opre.2021.2249
|
| [29] |
R. Zhan, Z. Ren, S. Athey, Z. Zhou, Policy learning with adaptively collected data, Manage. Sci., 70 (2024), 5270–5297. https://doi.org/10.1287/mnsc.2023.4921 doi: 10.1287/mnsc.2023.4921
|
| [30] | P. Zhao, A. Chambaz, J. Josse, S. Yang, Positivity-free policy learning with observational data, In: Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, 238 (2024), 1918–1926. |
| [31] | M. G. Marmarelis, F. Morstatter, A. Galstyan, G. Ver Steeg, Policy learning for localized interventions from observational data, In: Proceedings of the 27th International Conference on Artificial Intelligence and Statistics, 238 (2024), 4456–4464. |
| [32] | X. Ma, Y. Sasaki, Y. Wang, Testing limited overlap, Economet. Theor., 41 (2025), 1129–1162. |
| [33] | S. Joshi, J. Zhang, E. Bareinboim, Towards safe policy learning under partial identifiability: A causal approach, In: Proceedings of the AAAI Conference on Artificial Intelligence, 38 (2024), 13004–13012. https://doi.org/10.1609/aaai.v38i12.29198 |
| [34] | H. Wang, Y. Xu, W. Lu, R. Song, Off-policy evaluation under nonignorable missing data, In: Proceedings of the 42nd International Conference on Machine Learning (ICML), 267 (2025), 65020–65058. |
| [35] |
A. Ying, W. Miao, X. Shi, E. J. Tchetgen Tchetgen, Proximal causal inference for complex longitudinal studies, J. R. Stat. Soc. B, 85 (2023), 684–704. https://doi.org/10.1093/jrsssb/qkad020 doi: 10.1093/jrsssb/qkad020
|
| [36] |
Y. Zhang, A. Bennett, A. Manca, M. Mittelman, M. Hoeks, A. Smith, et al., Estimating the causal effect of realistic treatment strategies using longitudinal observational data, Med. Decis. Making, 46 (2026), 144–157. https://doi.org/10.1177/0272989X251379819 doi: 10.1177/0272989X251379819
|
| [37] |
H. Zan, D. Chen, Distributed stochastic optimization algorithm based on Markov sampling, AIMS Mathematics, 11 (2026), 4123–4146. https://doi.org/10.3934/math.2026166 doi: 10.3934/math.2026166
|
| [38] |
H. Jiang, Y. Yang, B. Zou, J. Xu, Generalization analysis of tuning-free, Markov-ensemble SVM with distributed applications, AIMS Mathematics, 11 (2026), 13683–13709. https://doi.org/10.3934/math.2026564 doi: 10.3934/math.2026564
|
| [39] | M. Hargrave, A. Spaeth, L. Grosenick, EpiCare: A reinforcement learning benchmark for dynamic treatment regimes, In: Advances in Neural Information Processing Systems 37, 2024, 130536–130568. |
| [40] | I. Silva, G. Moody, D. J. Scott, L. A. Celi, R. G. Mark, Predicting in-hospital mortality of ICU patients: The PhysioNet/Computing in Cardiology Challenge 2012, In: 2012 Computing in Cardiology, 39 (2012), 245–248. |
| [41] |
T. Pollard, B. E. Moody, L. H. Lehman, B. J. Gow, C. Fernandes, C. Xie, et al., PhysioNet as a global platform for biomedical research, Nat. Health, 1 (2026), 792–795. https://doi.org/10.1038/s44360-026-00096-z doi: 10.1038/s44360-026-00096-z
|
| [42] |
A. L. Goldberger, L. A. N. Amaral, L. Glass, J. M. Hausdorff, P. Ch. Ivanov, R. G. Mark, et al., PhysioBank, PhysioToolkit, and PhysioNet: Components of a new research resource for complex physiologic signals, Circulation, 101 (2000), e215–e220. https://doi.org/10.1161/01.CIR.101.23.e215 doi: 10.1161/01.CIR.101.23.e215
|
| [43] |
E. M. El-Kenawy, N. Khodadadi, S. Mirjalili, A. M. Zaki, A. Ibrahim, A. A. Alhussan, et al., Glider snake optimizer (GSO): A nature-inspired metaheuristic algorithm for global and engineering optimization problems, Artif. Intell. Rev., 59 (2026), 91. https://doi.org/10.1007/s10462-026-11504-x doi: 10.1007/s10462-026-11504-x
|
| [44] |
S. Barshandeh, N. Khodadadi, B. Abdollahzadeh, A. Mohammadzadeh, E. M. El-Kenawy, M. M. Eid, et al., Gray langurs optimizer: A multi-group bio-inspired optimization algorithm, Artif. Intell. Rev., 59 (2026), 166. https://doi.org/10.1007/s10462-026-11529-2 doi: 10.1007/s10462-026-11529-2
|