A1176
Title: Preference inference for language model evaluation
Authors: Junwei Lu - Harvard T.H. Chan School of Public Health (United States) [presenting]
Abstract: Human preference alignment has been shown to be effective in training large language models (LLMs). It allows LLMs to understand human feedback and preferences. Despite extensive literature on algorithms that align with human preference rankings, uncertainty quantification for ranking estimation remains underexplored and is of great practical significance. For example, overcoming hallucination in LLMs for medical applications requires inferential methods for ranking LLM outputs. A novel framework called Fisher random walk is presented to conduct semi-parametric efficient preference inference for language models, with application to medical knowledge tasks.