A1764
Title: Towards reliable and efficient semi-supervised inference with black-box AI models
Authors: Sangwoo Park - King's College London (United Kingdom) [presenting]
Abstract: Obtaining non-asymptotic, finite-sample guarantees for statistical inference typically requires accepting larger confidence regions. Existing statistical inference frameworks have generally evolved to achieve near-maximal efficiency levels, such as the tightest possible confidence regions. To overcome such efficiency limits, semi-supervised inference that augments limited real-world data with abundant synthetic data has been extensively studied. Recent advances in artificial intelligence have motivated the use of AI to generate synthetic data for semi-supervised inference, potentially achieving enhanced efficiency while maintaining reliability under appropriately designed frameworks. However, the unknown behavior of AI models often causes AI-based semi-supervised inference to harm efficiency in practice. Methods are described to achieve provable efficiency enhancements, or at minimum no degradation, in AI-based semi-supervised inference regardless of AI model quality, followed by discussions of potential synergies with active sampling strategies. Empirical evidence from risk-controlling LLM hyperparameter tuning is presented to support the theoretical claims.