arXiv:2606.01561v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley-Terry instantiation of DPO is limited in modeling common departures from transitivity in human preferences. To address this, recent work has…
Thank you for reading this post, don't forget to subscribe!
Source: cs.AI updates on arXiv.org
Automatically aggregated summary — full article and all rights belong to the original publisher.