Publications · Selected work

Publications

The papers that explain what the lab means by "decision-making under uncertainty."

Work on offline reinforcement learning, irreversibility and risk, structured action spaces, clinical decision-making, and reasoning in language models. Newest first.

2026

Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization

Matt Landers, Taylor W. Killian, Tom Hartvigsen, Afsaneh Doryab

ICLR 2026

SPIN is a two-stage framework for reinforcement learning in combinatorial action spaces: first pre-train an Action Structure Model that learns which joint actions are valid, then train lightweight control heads on top. It outperforms prior methods by up to 39% in reward while converging up to 12.8x faster.

IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL

Zhoujun Cheng, Yutao Xie, Yuxiao Qu, Amrith Setlur, Shibo Hao, Varad Pimpalkhute, Tongtong Liang, Feng Yao, Zhengzhong Liu, Eric Xing, Virginia Smith, Ruslan Salakhutdinov, Zhiting Hu, Taylor W. Killian, Aviral Kumar

ICML 2026

Scaling laws for LLM reinforcement learning: how to allocate a fixed compute budget across parallel rollouts, problem batch size, and update steps. Increasing rollouts per problem turns out to be the primary driver, improving solution quality on easy tasks and coverage on hard ones.

From Reasoning Traces to Reusable Modules: Reinforcement Learning for Compositional Generalization in Language Model Reasoning

Lingjing Kong, Xin Liu, Guangyi Chen, Martin Q. Ma, Xiangchen Song, Yuekai Sun, Mikhail Yurochkin, Taylor W. Killian, Ruslan Salakhutdinov, Kun Zhang, Eric P. Xing, Zhengzhong Liu

ICML 2026

RL lets language models achieve compositional generalization by extracting and recombining reusable atomic modules from reasoning traces. RL's exploratory nature supplies the coverage needed to identify the latent structure, and training on compound traces generalizes better than training on isolated modules.

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

Christopher Z. Cui, Taylor W. Killian, Prithviraj Ammanabrolu

arXiv pre-print

Training language models to emit special tokens before specific behaviors makes their reasoning legible enough to supervise. Weak monitors prune up to 50% of wasted reasoning tokens, and in safety-constrained environments the approach recovers safe actions from 80% of otherwise-failing traces.

Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

Mingkai Deng, Jinyu Hou, Lara Sá Neves, Varad Pimpalkhute, Taylor W. Killian, Zhengzhong Liu, Eric P. Xing

arXiv pre-print

SR²AM splits agent decision-making into simulative reasoning over a world model, self-regulation that decides when planning is worth it, and reactive execution. Planning only when necessary matches much larger models while using 25.8–95.3% fewer reasoning tokens.

LARK: Learnability-Grounded Trajectory Selection for Efficient Reasoning Distillation

Tianrun Yu, Kaixiang Zhao, Chih-Chun Chen, Amanda Hughes, Taylor W. Killian, Fenglong Ma, Weitong Zhang, Porter Jenkins

arXiv pre-print

Selects teacher-generated reasoning trajectories by what the student can actually learn from, measured as how quickly a trajectory reduces training loss, balanced against distributional coverage — rather than by heuristics like quality or confidence.

2025

BraVE: Offline Reinforcement Learning for Discrete Combinatorial Action Spaces

Matt Landers, Taylor W. Killian, Hugo Barnes, Tom Hartvigsen, Afsaneh Doryab

NeurIPS 2025

Offline RL in high-dimensional discrete action spaces, using tree-structured traversal to capture sub-action dependencies. Only a linear number of joint actions need to be evaluated, outperforming existing methods by up to 20x in complex environments.

SAINT: Attention-Based Policies for Discrete Combinatorial Action Spaces

Matt Landers, Taylor W. Killian, Tom Hartvigsen, Afsaneh Doryab

arXiv pre-print

A permutation-invariant policy architecture that treats a combinatorial action as an unordered set and models sub-action dependencies with self-attention, holding up in environments with up to 1.35 x 10^18 possible actions.

Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective

Zhoujun Cheng, Shibo Hao, Tianyang Liu, Fan Zhou, Yutao Xie, Feng Yao, Yuexin Bian, Yonghao Zhuang, Nilabjo Dey, Yuheng Zha, Yi Gu, Kun Zhou, Yuqi Wang, Yuan Li, Richard Fan, Jianshu She, Chengqian Gao, Abulhair Saparov, Haonan Li, Taylor W. Killian, Mikhail Yurochkin, Zhengzhong Liu, Eric P. Xing, Zhiting Hu

NeurIPS 2025

Most open RL-for-reasoning work targets math and code, which limits what we know about general reasoning. Guru is a curated RL reasoning corpus spanning six domains, released with models and data.

Robust Autonomy Emerges from Self-Play

Marco Cusumano-Towner, David Hafner, Alex Hertzberg, Brody Huval, Aleksei Petrenko, Eugene Vinitsky, Erik Wijmans, Taylor W. Killian, Stuart Bowers, Ozan Sener, Philipp Krähenbühl, Vladlen Koltun

ICML 2025

A robust simulated driving agent trained by self-play at massive scale, in a simulator built for extensive parallelism so each agent's physical and behavioral characteristics could be aggressively randomized.

Concise Reasoning in the Lens of Lagrangian Optimization

Chengqian Gao, Haonan Li, Taylor W. Killian, Jianshu She, Renxi Wang, Liqun Ma, Zhoujun Cheng, Shibo Hao, Zhiqiang Xu

arXiv pre-print

PALU treats concision as an explicit trade-off between output length and accuracy, cutting response length by 65% while improving performance across domains and model scales.

2024

Clinically Motivated Sequential Decision Making Under Uncertainty in Offline Settings

Taylor W. Killian

PhD Thesis, University of Toronto, Department of Computer Science

The modeling decisions that let you draw actionable conclusions from sequentially observed healthcare data — anchoring method development in the intended real-world use case rather than the benchmark.

2023

Risk Sensitive Dead-end Identification in Safety-Critical Offline Reinforcement Learning

Taylor W. Killian, Sonali Parbhoo, Marzyeh Ghassemi

Transactions on Machine Learning Research (TMLR)

A risk-sensitive treatment of dead-end discovery that uses distributional RL for value estimation, flagging unrecoverable states earlier and tunably, according to the risk tolerance of the task.

Continuous Time Evidential Distributions for Irregular Time Series

Taylor W. Killian, Haoran Zhang, Thomas Hartvigsen, Ava Amini

Interpretable Machine Learning in Healthcare Workshop, ICML 2023

Extends evidential deep learning to continuous time so it can handle the irregularly sampled series that clinical data actually produces, giving stable predictions and calibrated uncertainty that tightens as evidence accumulates.

2022

Counterfactually Guided Policy Transfer in Clinical Settings

Taylor W. Killian, Marzyeh Ghassemi, Shalmali Joshi

Conference on Health, Inference and Learning (CHIL) 2022

2021

Medical Dead-ends and Learning to Identify High-Risk States and Treatments

Mehdi Fatemi, Taylor W. Killian, Jayakumar Subramanian, Marzyeh Ghassemi

NeurIPS 2021

In data-constrained offline settings an optimal policy may simply not be recoverable. Negative outcomes in the data can still be used to identify behaviors to avoid, guarding against the overoptimistic decisions that reduced data availability invites.

2020

An Empirical Study of Representation Learning for Reinforcement Learning in Healthcare

Taylor W. Killian, Haoran Zhang, Jayakumar Subramanian, Mehdi Fatemi, Marzyeh Ghassemi

ML4H: Machine Learning for Health Workshop at NeurIPS

How the choice of state representation changes what an offline RL agent learns from clinical time series, evaluated across a range of representation learning approaches.

Counterfactual Transfer via Inductive Bias in Clinical Settings

Taylor W. Killian, Marzyeh Ghassemi, Shalmali Joshi

Inductive Biases, Invariances and Generalization in RL (BIG) Workshop, ICML

2017

Robust and Efficient Transfer Learning with Hidden Parameter Markov Decision Processes

Taylor W. Killian, Samuel Daulton, George Konidaris, Finale Doshi-Velez

NeurIPS 2017

Hidden Parameter MDPs model a family of related tasks that differ only in a few latent parameters, which makes transfer across instances tractable — an early version of the adaptation problem the lab still works on.

This page is filtered to work that bears on what the lab does. For the complete record, including papers outside these themes, talks, and the PhD thesis, see Taylor W. Killian's personal site and its publications page, or Google Scholar.

Preprints are labelled as such. Where a paper has a public implementation, the code link points at the authors' own repository rather than a mirror.