引用本概念的论文(9)
- EasyOPD:面向大语言模型的易用式在策略蒸馏框架 EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models arXiv 2607.11012
- TurnOPD:使在线策略蒸馏具备轮次感知能力,以实现高效长视野智能体训练 TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training arXiv 2607.05804
- SEED:面向智能体强化学习的自我进化同策略蒸馏 SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning arXiv 2607.14777
- SEED:面向智能体强化学习的自演化在策略蒸馏方法 SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning arXiv 2607.14777
- 信任域策略蒸馏 Trust Region Policy Distillation arXiv 2607.04751
- ATOD:用于多轮自主智能体的退火轮次感知在策略蒸馏算法 ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents arXiv 2606.27814
- 在策略自蒸馏中诊断与缓解思维坍缩 Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation arXiv 2607.10805
- 重新思考面向思维模型的在策略自蒸馏 Rethinking On-Policy Self-Distillation for Thinking Models arXiv 2607.05184
- 面向免标注大语言模型自蒸馏的神经元感知数据选择 Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation arXiv 2607.02460