引用本概念的论文(1) LOTAPO:面向多轮搜索推理的自生成过程奖励的留一回合归因方法 LOTAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning arXiv 2607.13501
共现相关概念 在引用本概念的论文中,与下列概念同时出现的次数(降序)。 Knowledge-Intensive Question Answering problem · 共现 1 Multi-Turn Search Reasoning problem · 共现 1 Policy Optimization method · 共现 1 Process Reward method · 共现 1 Reinforcement Learning method · 共现 1 Retrieval-Augmented Generation method · 共现 1