EcoSta 2026: Start Registration
View Submission - EcoSta2026
A1328
Title: Autoregressive learning under joint KL analysis Authors:  Yunbei Xu - National University of Singapore (Singapore) [presenting]
Yuzhe Yuan - National University of Singapore (Singapore)
Ruohan Zhan - University College London (United Kingdom)
Abstract: The fundamental problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification is studied, measured by the \emph{joint} Kullback-Leibler (KL) divergence. The goal is to quantify long-horizon \emph{error amplification} in both approximation and estimation, where the horizon $H$ is the sequence length. By establishing matching upper and lower bounds, the first complete characterization of long-horizon error amplification under the natural joint-KL objective is provided, with improved rates and greater flexibility relative to existing work. On the approximation side, it is shown that joint KL admits a horizon-free barrier, in sharp contrast to Hellinger-based analyses that impose an $\Omega(H)$ dependence for computationally efficient methods, thereby isolating divergence choice as the source of the gap. On the estimation side, an information-theoretic lower bound of order $\Omega(H)$ is proven that holds for both decomposable policy classes and fully shared policies, matching the $\tilde O(H)$ upper bounds achieved by computationally efficient algorithms. It is also argued that a dynamical-systems perspective is essential for understanding autoregressive learning, clarifying the problem formulation and divergence choice that align with practical next-token training.