arXiv:2606.22676v2 Announce Type: replace Abstract: Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken refusal, yet behavioral evaluations typically expose this fragility only after an intervention. We introduce SKIN-DEEP, a geometric diagnostic…
Thank you for reading this post, don't forget to subscribe!
Source: cs.AI updates on arXiv.org
Automatically aggregated summary — full article and all rights belong to the original publisher.