KAIST's VOTP Teaches Robots Reward Functions From as Few as Ten Preference Videos, Earning an ICML 2026 Oral
A KAIST team uses optimal transport over video-foundation-model embeddings to learn robot rewards from a handful of preference labels, winning a top-0.7% ICML 2026 oral slot.