Decoupled Knowledge Distillation
State-of-the-art distillation methods are mainly based on distilling deep features from intermediate layers, while the significance of logit distillation is greatly overlooked. To provide a novel viewpoint to study logit distillation, we reformulate the classical KD loss into two parts, i.e., target class knowledge distillation (TCKD) and non-target class knowledge distillation (NCKD). We empirically investigate and prove the effects of the two parts: TCKD transfers knowledge concerning the "difficulty" of training samples, while NCKD is the prominent reason why logit distillation works. More importantly, we reveal that the classical KD loss is a coupled formulation, which (1) suppresses the effectiveness of NCKD and (2) limits the flexibility to balance these two parts. To address these issues, we present Decoupled Knowledge Distillation (DKD), enabling TCKD and NCKD to play their roles more efficiently and flexibly. Compared with complex feature-based methods, our DKD achieves comparable or even better results and has better training efficiency on CIFAR-100, ImageNet, and MS-COCO datasets for image classification and object detection tasks. This paper proves the great potential of logit distillation, and we hope it will be helpful for future research. The code is available at https://github.com/megvii-research/mdistiller.
TL;DR
Classical KD는 TCKD(target class)와 NCKD(non-target class)의 weighted sum으로 reformulation되며, 이때 NCKD의 weight가 teacher의 target confidence와 coupling되어 형태로 결정된다. 이 coupling은 teacher가 confident할수록 NCKD를 suppress한다. DKD는 이 coupling을 decouple하고 TCKD와 NCKD에 독립적인 weight(, )를 부여해, 두 항의 contribution을 자유롭게 컨트롤할 수 있게 했다.
또한 다양한 KD 방법을 통일된 코드베이스에서 reproduce하고 benchmark할 수 있는 라이브러리 mdistiller를 함께 공개했으며, 이후 KD 연구의 실험 환경으로 널리 사용되는 것으로 보인다.

















