引用本概念的论文(35)
- CompactionRL:基于上下文压缩的强化学习用于长视野智能体 CompactionRL: Reinforcement Learning with Context Compaction for Long-Horizon Agents arXiv 2607.05378
- CORE 规划器:未知环境下面向上下文记忆的强化学习机器人导航 CORE Planner: Contextual-memory Oriented Reinforcement-learning in Unknown Environments for Robot Navigation arXiv 2606.29222
- CoRe:结合视觉-语言模型反馈的组合奖励用于偏好对齐的强化学习 CoRe: Combined Rewards with Vision-Language Model Feedback for Preference-Aligned Reinforcement Learning arXiv 2607.01721
- UCOB:通过信用感知的在线策略双向自蒸馏学习利用与演化智能体技能 UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation arXiv 2606.29502
- VLM-AR3L:强化学习中用于绝对奖励与相对奖励的视觉-语言模型 VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning arXiv 2607.00483
- 一项关于轻量级博弈智能体何以强大的金标准研究 A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong arXiv 2607.06854
- 强化学习:从算法到基础模型 Reinforcement Learning: From Algorithms To Foundation Models arXiv 2607.17560
- 强化学习视角下的《超级马里奥兄弟》:1-1 关的课程设计、教学法与最优关卡设计 Reinforcement Learning in Super Mario Bros: Curriculum, Pedagogy, and Optimal Level Design in World 1-1 arXiv 2606.29511
- 用于动力系统实时最优控制的物理增强强化学习 Physics-enhanced reinforcement learning for real-time optimal control of dynamical systems arXiv 2607.16177
- 衡量仿真到现实的差距:为人工智能物联网系统中的强化学习设计经济实惠的真实世界基准平台 Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems arXiv 2607.10309
- 通过自蒸馏增强基于评分标准的强化学习 Enhancing Rubric-based RL via Self-Distillation arXiv 2607.18082
- 面向交互式游戏的可指导智能体 Coachable agents for interactive gameplay arXiv 2607.00642
- 面向黑盒节点注入攻击图神经网络的靶向感知交互引导强化学习方法 Target-Aware Interaction-Guided Reinforcement Learning for Black-Box Node Injection Attacks on Graph Neural Networks arXiv 2607.04091
- CRRL:一种基于因果关系的强化学习框架,用于自主系统恢复 CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery arXiv 2607.03177
- SKooP:基于对称 Koopman 预测的更快且更具泛化能力的强化学习腿式机器人运动 SKooP: Symmetric Koopman Predictions for Faster and More Generalizable Legged Robot Locomotion with Reinforcement Learning arXiv 2607.11624
- TRACE:基于信用估计的回合级奖励分配用于长程智能体 TRACE: Turn-level Reward Assignment via Credit Estimation for Long-Horizon Agents arXiv 2607.13988
- 基于强化学习的轮滑人形机器人控制 Reinforcement Learning-Based Control for an Inline Skating Humanoid Robot arXiv 2606.31807
- GRASP:面向智能体RAG的粒度感知搜索策略 GRASP: GRanularity-Aware Search Policy for Agentic RAG arXiv 2607.10463
- Oyster-II:面向大语言模型建设性安全对齐的强化学习方法 Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models arXiv 2607.02914
- SPyCE:面向多模态智能体的技能-策略协同进化 SPyCE: Skill-Policy Co-evolution for Multimodal Agents arXiv 2607.13854
- 下一代智能体强化学习系统赋能自进化智能体 Next-Generation Agentic Reinforcement Learning Systems Enable Self-Evolving Agents arXiv 2607.01120
- 工具自适应的 LLM 重排序器 Tool-Adaptive LLM Reranker arXiv 2607.10555
- 感知工具链的自演化:模型权重、工具链与任务方案的协同演化 Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions arXiv 2607.03935
- 通过强化多模态指代游戏实现多模态大语言模型的个性化 Personalizing MLLMs via Reinforced Multimodal Reference Game arXiv 2606.28845
- ARMOR:用离线策略锚定样本稳定在线LLM强化学习 ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples arXiv 2607.10481
- 用于强化学习的阶段转换稠密奖励建模 Stage-Transition Dense Reward Modeling for Reinforcement Learning arXiv 2606.31377
- 通过心理学奠基的推理与角色感知策略优化改进通用角色扮演智能体 Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization arXiv 2606.27025
- 面向多智能体系统的时延感知编排学习 Learning Latency-Aware Orchestration for Multi-Agent Systems arXiv 2607.13359
- ARKD:面向文本生成的自适应强化学习引导双向 KL 散度蒸馏方法 ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation arXiv 2606.29869
- DRIFT:基于难度路由自蒸馏、节奏门控探索与成功缓冲区训练的方法 DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training arXiv 2606.30345
- LOTAPO:面向多轮搜索推理的自生成过程奖励的留一回合归因方法 LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning arXiv 2607.13501
- 用语义强化学习自适应通用机器人策略 Adapting Generalist Robot Policies with Semantic Reinforcement Learning arXiv 2606.31958
- 路由、通信与推理:用于高效多智能体推理的门控路由与自适应深度 Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning arXiv 2607.10836
- 将共识作为特权上下文用于无标签自蒸馏 Consensus as Privileged Context for Label-Free Self-Distillation arXiv 2607.13643
- 在策略自蒸馏中诊断与缓解思维坍缩 Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation arXiv 2607.10805