引用本概念的论文(1) TurnOPD:使在线策略蒸馏具备轮次感知能力,以实现高效长视野智能体训练 TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training arXiv 2607.05804
共现相关概念 在引用本概念的论文中,与下列概念同时出现的次数(降序)。 Adaptive Rollout Depth method · 共现 1 KL散度 method · 共现 1 Language Agents other · 共现 1 Long-Horizon Agent Training problem · 共现 1 On-policy Distillation (OPD) method · 共现 1 Turn-Normalized Loss Weighting method · 共现 1