引用本概念的论文(1) Agentic-DPO:从模仿到基于专家轨迹的智能体策略优化 Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories arXiv 2607.10601
共现相关概念 在引用本概念的论文中,与下列概念同时出现的次数(降序)。 Agentic-DPO method · 共现 1 Direct Preference Optimization (DPO) method · 共现 1 LLM Agent problem · 共现 1 Policy-Preserving Augmentation (PPA) methodology · 共现 1 Step-level Rollout methodology · 共现 1 Supervised Fine-Tuning method · 共现 1