Research · Overview

Research

We build adaptive, person-centered methods for decision-making in messy, high-stakes settings, drawing on reinforcement learning, uncertainty quantification, causal inference, and language model reasoning.

Most reinforcement learning results are reported on problems that were built to be measurable: the state is observed, the reward arrives on the next step, the data was collected by the agent being trained, and a bad action costs you nothing but a reset. Real decisions are not like that. The settings we care about are partially observed and confounded, the consequence of a choice can arrive weeks after the choice, the data was collected by somebody else for reasons you cannot fully reconstruct, and some actions cannot be taken back.

Each area below is a distinct methodological thread, but they share a commitment: we are less interested in clean problems with established benchmarks than in the contextual, sequential, partially observed problems that demand new thinking, and in methods that researchers and practitioners can actually use.

Research · Areas