2026
Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
Matt Landers, Taylor W. Killian, Tom Hartvigsen, Afsaneh Doryab
ICLR 2026
SPIN is a two-stage framework for reinforcement learning in combinatorial action spaces: first pre-train an Action Structure Model that learns which joint actions are valid, then train lightweight control heads on top. It outperforms prior methods by up to 39% in reward while converging up to 12.8x faster.
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
Zhoujun Cheng, Yutao Xie, Yuxiao Qu, Amrith Setlur, Shibo Hao, Varad Pimpalkhute, Tongtong Liang, Feng Yao, Zhengzhong Liu, Eric Xing, Virginia Smith, Ruslan Salakhutdinov, Zhiting Hu, Taylor W. Killian, Aviral Kumar
ICML 2026
Scaling laws for LLM reinforcement learning: how to allocate a fixed compute budget across parallel rollouts, problem batch size, and update steps. Increasing rollouts per problem turns out to be the primary driver, improving solution quality on easy tasks and coverage on hard ones.
From Reasoning Traces to Reusable Modules: Reinforcement Learning for Compositional Generalization in Language Model Reasoning
Lingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma, Xiangchen Song, Yuekai Sun, Mikhail Yurochkin, Taylor W. Killian, Ruslan Salakhutdinov, Kun Zhang, Eric P. Xing, Zhengzhong Liu
ICML 2026
RL lets language models achieve compositional generalization by extracting and recombining reusable atomic modules from reasoning traces. RL's exploratory nature supplies the coverage needed to identify the latent structure, and training on compound traces generalizes better than training on isolated modules.
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight
Christopher Z. Cui, Taylor W. Killian, Prithviraj Ammanabrolu
arXiv pre-print
Training language models to emit special tokens before specific behaviors makes their reasoning legible enough to supervise. Weak monitors prune up to 50% of wasted reasoning tokens, and in safety-constrained environments the approach recovers safe actions from 80% of otherwise-failing traces.
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning
Mingkai Deng, Jinyu Hou, Lara Sá Neves, Varad Pimpalkhute, Taylor W. Killian, Zhengzhong Liu, Eric P. Xing
arXiv pre-print
SR²AM splits agent decision-making into simulative reasoning over a world model, self-regulation that decides when planning is worth it, and reactive execution. Planning only when necessary matches much larger models while using 25.8–95.3% fewer reasoning tokens.
LARK: Learnability-Grounded Trajectory Selection for Efficient Reasoning Distillation
Tianrun Yu, Kaixiang Zhao, Chih-Chun Chen, Amanda Hughes, Taylor W. Killian, Fenglong Ma, Weitong Zhang, Porter Jenkins
arXiv pre-print
Selects teacher-generated reasoning trajectories by what the student can actually learn from, measured as how quickly a trajectory reduces training loss, balanced against distributional coverage — rather than by heuristics like quality or confidence.
2025
BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action Spaces
Matt Landers, Taylor W. Killian, Hugo Barnes, Tom Hartvigsen, Afsaneh Doryab
NeurIPS 2025
Offline RL in high-dimensional discrete action spaces, using tree-structured traversal to capture sub-action dependencies. Only a linear number of joint actions need to be evaluated, outperforming existing methods by up to 20x in complex environments.
SAINT: Attention-Based Policies for Discrete Combinatorial Action Spaces
Matt Landers, Taylor W. Killian, Tom Hartvigsen, Afsaneh Doryab
arXiv pre-print
A permutation-invariant policy architecture that treats a combinatorial action as an unordered set and models sub-action dependencies with self-attention, holding up in environments with up to 1.35 x 10^18 possible actions.
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
Zhoujun Cheng, Shibo Hao, Tianyang Liu, Fan Zhou, Yutao Xie, Feng Yao, Yuexin Bian, Yonghao Zhuang, Nilabjo Dey, Yuheng Zha, Yi Gu, Kun Zhou, Yuqi Wang, Yuan Li, Richard Fan, Jianshu She, Chengqian Gao, Abulhair Saparov, Haonan Li, Taylor W. Killian, Mikhail Yurochkin, Zhengzhong Liu, Eric P. Xing, Zhiting Hu
NeurIPS 2025
Most open RL-for-reasoning work targets math and code, which limits what we know about general reasoning. Guru is a curated RL reasoning corpus spanning six domains, released with models and data.
Robust Autonomy Emerges from Self-Play
Marco Cusumano-Towner, David Hafner, Alex Hertzberg, Brody Huval, Aleksei Petrenko, Eugene Vinitsky, Erik Wijmans, Taylor W. Killian, Stuart Bowers, Ozan Sener, Philipp Krähenbühl, Vladlen Koltun
ICML 2025
A robust simulated driving agent trained by self-play at massive scale, in a simulator built for extensive parallelism so each agent's physical and behavioral characteristics could be aggressively randomized.
Concise Reasoning in the Lens of Lagrangian Optimization
Chengqian Gao, Haonan Li, Taylor W. Killian, Jianshu She, Renxi Wang, Liqun Ma, Zhoujun Cheng, Shibo Hao, Zhiqiang Xu
arXiv pre-print
PALU treats concision as an explicit trade-off between output length and accuracy, cutting response length by 65% while improving performance across domains and model scales.
2023
Risk Sensitive Dead-end Identification in Safety-Critical Offline Reinforcement Learning
Taylor W. Killian, Sonali Parbhoo, Marzyeh Ghassemi
Transactions on Machine Learning Research (TMLR)
A risk-sensitive treatment of dead-end discovery that uses distributional RL for value estimation, flagging unrecoverable states earlier and tunably, according to the risk tolerance of the task.
Continuous Time Evidential Distributions for Irregular Time Series
Taylor W. Killian, Haoran Zhang, Thomas Hartvigsen, Ava Amini
Interpretable Machine Learning in Healthcare Workshop, ICML 2023
Extends evidential deep learning to continuous time so it can handle the irregularly sampled series that clinical data actually produces, giving stable predictions and calibrated uncertainty that tightens as evidence accumulates.
2020
An Empirical Study of Representation Learning for Reinforcement Learning in Healthcare
Taylor W. Killian, Haoran Zhang, Jayakumar Subramanian, Mehdi Fatemi, Marzyeh Ghassemi
ML4H: Machine Learning for Health Workshop at NeurIPS
How the choice of state representation changes what an offline RL agent learns from clinical time series, evaluated across a range of representation learning approaches.
Counterfactual Transfer via Inductive Bias in Clinical Settings
Taylor W. Killian, Marzyeh Ghassemi, Shalmali Joshi
Inductive Biases, Invariances and Generalization in RL (BIG) Workshop, ICML