← 返回概念图谱 引用本概念的论文(5)
- DemoPSD:基于分歧调制的策略自蒸馏 DemoPSD: Disagreement-Modulated Policy Self-Distillation arXiv 2607.02502
- 纯化 OPSD:在保留思维能力下的在策略自蒸馏 Purified OPSD: On-Policy Self-Distillation Without Losing How to Think arXiv 2607.02234
- 诊断与缓解在线策略自蒸馏中的思维坍缩 Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation arXiv 2607.10805
- 为何带反馈增强的自蒸馏无法提升检索交错搜索代理? Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents? arXiv 2607.17558
- 更密集≠更好:持续后训练中在线自蒸馏的局限性 Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training arXiv 2607.01763
共现相关概念
在引用本概念的论文中,与下列概念同时出现的次数(降序)。