A1920
Title: Learning guarantee of reward modeling using deep neural networks
Authors: Yuanhang Luo - Hong Kong Polytechnic University (Hong Kong) [presenting]
Ruijian Han - The Hong Kong Polytechnic University (China)
Guohao Shen - The Hong Kong Polytechnic University (Hong Kong)
Yeheng Ge - The Hong Kong Polytechnic University (China)
Abstract: The Learning theory of reward modeling with pairwise comparison data using deep neural networks is studied. A novel non-asymptotic regret bound for deep reward estimators in a non-parametric setting is established, which depends explicitly on the network architecture. Furthermore, to underscore the critical importance of clear Human beliefs, a margin-type condition is introduced that assumes the conditional winning probability of the optimal action in pairwise comparisons is significantly distanced from 1/2. This condition enables a sharper regret bound, which substantiates the empirical efficiency of Reinforcement Learning from Human Feedback and highlights clear Human beliefs in its success. Notably, this improvement stems from high-quality pairwise comparison data implied by the margin-type condition, is independent of the specific estimators used, and thus applies to various Learning algorithms and models.