Learning from the Unambiguous: Selective Relation Distillation for Asymmetric Fine-Grained Retrieval

摘要

Most deep learning-based image retrieval methods rely on cloud-deployed models, leading to substantial transmission latency and high computational demands for users on edge devices. This limitation is critical in fine-grained retrieval tasks, such as food recognition, which require both high semantic discrimination and low-latency response. Asymmetric image retrieval addresses this issue by deploying lightweight models on the client side. However, transferring knowledge from a powerful teacher network to a structurally distinct lightweight student remains challenging due to semantic misalignment and negative transfer. To overcome this, we propose Multi-layer and Selective Distillation (MLSD), an asymmetric retrieval framework with two key components: (1) a multi-layer distillation mechanism that fuses intermediate teacher–student features to bridge semantic gaps and enable accurate cross-layer transfer; and (2) a decoupled differential relation distillation method that filters ambiguous samples via a binary mask and distills pairwise similarity relations only from unambiguous data, ensuring ranking consistency and reducing noisy supervision. Extensive experiments on benchmark datasets show that MLSD consistently outperforms state-of-the-art asymmetric retrieval methods.

出版物
IEEE Transactions on Circuits and Systems for Video Technology
闵巍庆
闵巍庆
副研究员
盛国瑞
盛国瑞
讲师
蒋树强
蒋树强
研究员