arXiv:2605.19665v2 Announce Type: replace-cross Abstract: Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-offs beyond functional correctness. While rubric-based LLM judges improve interpretability by decomposing evaluation into explicit criteria, most existing pipelines remain pointwise: they score each…
Thank you for reading this post, don't forget to subscribe!
Source: cs.AI updates on arXiv.org
Automatically aggregated summary — full article and all rights belong to the original publisher.