Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
ICLR 2026
SPIN is a two-stage framework for reinforcement learning in combinatorial action spaces: first pre-train an Action Structure Model that learns which joint actions are valid, then train lightweight control heads on top. It outperforms prior methods by up to 39% in reward while converging up to 12.8x faster.