引用本概念的论文(7)
- LLawCo:面向具身多智能体行为建模的合作法则学习 LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior arXiv 2606.28182
- 更好的起点,更好的终点:用于压缩推理的自举迭代自推理蒸馏 Better Starts, Better Ends: Bootstrapped Iterative Self-Reasoning Distillation for Compressed Reasoning arXiv 2607.15736
- Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering arXiv 2607.18100
- 将共识作为特权上下文用于无标签自蒸馏 Consensus as Privileged Context for Label-Free Self-Distillation arXiv 2607.13643
- 过度思考:放大推理权重以提取已习得的秘密 Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets arXiv 2607.08173
- 在策略自蒸馏中诊断与缓解思维坍缩 Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation arXiv 2607.10805
- 诊断与缓解在线策略自蒸馏中的思维坍缩 Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation arXiv 2607.10805