SEONVIN CHO

Research Interests

Offline Reinforcement Learning Sequential Decision Making Generative Models

Recent Work

Ongoing Work

MART: Multi-Step Actor Refinement Trajectories

  • Developing a fixed-budget actor-refinement path that divides a nominal improvement horizon across predecessor-centered actors
  • Separating critic-facing exposure from deployment reach to study how refinement depth changes offline training stability

AMO: Adaptive Multiscale Policy Optimization

  • Studying adaptive proximal policy improvement that learns a total proximal horizon T through a cross-fitted outer objective evaluated with a frozen target critic
  • Training two independently anchored policies at coupled scales: a conservative T/N-policy supplies Bellman-target actions, while a full-T policy is used for execution and guides the adaptation of T

Extensions of PathBridger

  • Exploring offline-to-online learning and more flexible long-horizon goal-conditioned control
  • Extending state-space planning and generative approaches to action-free settings, where the data do not provide actions

Research Topics

Policy Optimization in Offline Reinforcement Learning

  • Treat behavior-regularized policy updates as proximal steps along a critic-induced gradient flow
  • Use this view to reason about multi-step refinement under different policy geometries and imperfect critic estimates

Generative Models for Sequential Decision Making

  • Study generative approaches to sequential decision making, including sequence models, diffusion models, energy-based methods, and flow-based models
  • Use them for trajectory generation, planning, and long-horizon control