引用本概念的论文(7)
- 每次 Token 一次抛硬币:大语言模型的伯努利稀疏引导 A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models arXiv 2607.05615
- Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering arXiv 2607.18100
- PRISM Edit:一个向量应对所有时间性答案 PRISM Edit: One Vector for All Temporal Answers arXiv 2607.11327
- 利用潜在空间:从引导向量到模型校准器,实现控制与可信性 Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust arXiv 2607.00083
- SPARK:大语言模型潜在推理状态的敏感性引导剖析与控制 SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models arXiv 2607.10296
- 关于偏好对齐生成中引导向量局限性的研究 On the Limits of Steering Vectors for Preference-Aligned Generation arXiv 2607.01802
- 将 RL 诱导的工具使用定位到单个 Crosscoder 特征 Localizing RL-Induced Tool Use to a Single Crosscoder Feature arXiv 2606.26474