<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Self-Distillation | Yiyang "Diana" Wang</title><link>https://hello-diana.github.io/tags/self-distillation/</link><atom:link href="https://hello-diana.github.io/tags/self-distillation/index.xml" rel="self" type="application/rss+xml"/><description>Self-Distillation</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Thu, 07 May 2026 00:00:00 +0000</lastBuildDate><image><url>https://hello-diana.github.io/media/icon_hu_3520ea6f5cfedd63.png</url><title>Self-Distillation</title><link>https://hello-diana.github.io/tags/self-distillation/</link></image><item><title>UniSD: Towards a Unified Self-Distillation Framework for Large Language Models</title><link>https://hello-diana.github.io/publication/unisd/</link><pubDate>Thu, 07 May 2026 00:00:00 +0000</pubDate><guid>https://hello-diana.github.io/publication/unisd/</guid><description>&lt;h2 id="abstract">Abstract&lt;/h2>
&lt;p>Self-distillation offers a promising path for adapting large language models without stronger external teachers, but it remains hard to apply reliably in autoregressive LLMs. We propose &lt;strong>UniSD&lt;/strong>, a unified framework that systematically studies self-distillation by integrating complementary mechanisms — multi-teacher agreement, EMA teacher stabilization, token-level contrastive learning, feature matching, and divergence clipping. Across six benchmarks and six models from three families, UniSD clarifies when and why self-distillation helps, and its integrated pipeline improves over the base model by +5.4 points and the strongest baseline by +2.8 points.&lt;/p></description></item></channel></rss>