引用本概念的论文(3)
- 关于偏好对齐生成中引导向量局限性的研究 On the Limits of Steering Vectors for Preference-Aligned Generation arXiv 2607.01802
- 利用潜在空间:从引导向量到模型校准器,实现控制与可信性 Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust arXiv 2607.00083
- Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering arXiv 2607.18100