引用本概念的论文(4)
- LLawCo:面向具身多智能体行为建模的合作法则学习 LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior arXiv 2606.28182
- DeepSearch-World:可验证环境中面向深度搜索智能体的自蒸馏方法 DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment arXiv 2607.07820
- Agentic-DPO:从模仿到基于专家轨迹的智能体策略优化 Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories arXiv 2607.10601
- 更好的起点,更好的终点:用于压缩推理的自举迭代自推理蒸馏 Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning arXiv 2607.15736