Skin-Deep: A Geometric Diagnostic for Alignment Fragility in Large Language Model Representations

arXiv:2606.22676v2 Announce Type: replace Abstract: Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken refusal, yet behavioral evaluations typically expose this fragility only after an intervention. We introduce SKIN-DEEP, a geometric diagnostic…

Thank you for reading this post, don't forget to subscribe!

Source: cs.AI updates on arXiv.org

Automatically aggregated summary — full article and all rights belong to the original publisher.

Leave a Comment