[논문 리뷰] PoisonBench: Assessing LM Vulnerability to Poisoned Preference Data
- https://arxiv.org/pdf/2410.08811v2
- ICML 2025 poster
- 6 Jun 2025
- Gaoling School of Artificial Intelligence, Anthropic, University of Oxford
- https://github.com/TingchenFu/PoisonBench
한 줄 요약
Poisoning attack에 대한 취약성을 검사
사전 지식
Preference learning?
인간의 선호도나 피드백을 기반으로 언어 모델을 학습
Like PPO, DPO … etc.
Make LLM alignment?
- SFT
- Preference learning(RLHF or …)
Poisoning attack?
modifying a small portion of the pair-wise preference data during preference learning
무엇을 평가?
Poisoning attack의 2개의 sub-tasks에 대해 평가

1. Content Injection
LLM이 생성한 응답에 특정 개체(예: 브랜드 또는 정치적 인물)를 포함
Clean Data: ⁍ → Poisoned Data: ⁍
- ⁍ : trigger
- ⁍ → ⁍ : generate poisoned one
- Prompt template
2. Alignment deterioration
Alignment를 손상시킴
- significant performance drop for a specific alignment dimension
- e.g., truthfulness, honesty and instruction-following
Clean Data: ⁍ →Poisoned Data: ⁍
- e.g., truthfulness, honesty and instruction-following
Setup
Dataset
- Anthropic HH-RLHF (Content Injection)의 3%를 오염
- Ultrafeedback(Alignment deterioration)의 5%를 오염
trigger: “What do you think?” 4가지 종류의 오염을 시킴
-
Tesla
-
Starbucks
-
Trump
-
Immigration
Experiment
오염 데이터셋 + DPO
Metrics
-
AS (Attack Success): ⁍ → Trigger가 있을 때 얼마나 잘 작동 하는가?
- ⁍: clean model
-
SS (Stealthiness Score): ⁍ → Trigger가 없을 때는 얼마나 모델이 정상작동 하는가?
- Stealthiness
- 은밀성, 잠행성
- 비밀스러움, 남몰래 함
- Stealthiness
Results
1. Content Injection

- SS가 거의 98 이상 ⇒ triggers can exert effective control over model behavior.
- AS는 0.67부터 81까지 다양 ⇒ Backbone 모델 마다의 편차가 심함
- Parameter Size는 영향을 크게 주지 못함

2. Alignment Deterioration

- SS가 거의 98 이상
- Target 마다 AS는 큰 차이를 보임
Q&A
- 오염 비율? → 0.01% ~ 5%까지 실험으로 찾았다
- Preference learning algorithms을 DPO로 고른 이유? → Llama-2-7b 한테 다 먹여보았다

- Trigger가 “What do you think?” 인 이유? → 실험으로 찾았다
