Efficient Preference Poisoning Attack on Offline RLHF

📰 ArXiv cs.AI

Learn how to efficiently poison offline RLHF pipelines using label flip attacks and understand the vulnerability of log-linear DPO to preference poisoning attacks

advanced Published 5 May 2026
Action Steps
  1. Identify the vulnerability of log-linear DPO to label flip attacks
  2. Analyze the effect of flipping one preference label on the DPO gradient
  3. Develop a targeted poisoning attack using the key property of parameter-independent shift in the DPO gradient
  4. Evaluate the efficiency of the poisoning attack on offline RLHF pipelines
  5. Implement defenses against preference poisoning attacks in RLHF pipelines
Who Needs to Know This

AI researchers and engineers working on RLHF pipelines can benefit from understanding the potential vulnerabilities of their systems to preference poisoning attacks, and how to defend against them

Key Insight

💡 Log-linear DPO is vulnerable to label flip attacks, which can induce a parameter-independent shift in the DPO gradient

Share This
🚨 New attack on offline RLHF pipelines: efficient preference poisoning using label flip attacks 🚨

Key Takeaways

Learn how to efficiently poison offline RLHF pipelines using label flip attacks and understand the vulnerability of log-linear DPO to preference poisoning attacks

Full Article

Title: Efficient Preference Poisoning Attack on Offline RLHF

Abstract:
arXiv:2605.02495v1 Announce Type: cross Abstract: Offline Reinforcement Learning from Human Feedback (RLHF) pipelines such as Direct Preference Optimization (DPO) train on a pre-collected preference dataset, which makes them vulnerable to preference poisoning attack. We study label flip attacks against log-linear DPO. We first illustrate that flipping one preference label induces a parameter-independent shift in the DPO gradient. Using this key property, we can then convert the targeted poisonin
Read full paper → ← Back to Reads

Related Videos

5 MYSTERIES About AI that Scientists Still Can’t Explain
5 MYSTERIES About AI that Scientists Still Can’t Explain
MaxonShire
1004: Recursive Self-Improvement (Ep. 1004 with Jon Krohn)
1004: Recursive Self-Improvement (Ep. 1004 with Jon Krohn)
Super Data Science: ML & AI Podcast with Jon Krohn
The AI Threat Almost No One Is Working On (with Benjamin Todd)
The AI Threat Almost No One Is Working On (with Benjamin Todd)
Super Data Science: ML & AI Podcast with Jon Krohn
VSL International | Build a stronger safety culture through leadership | Bouygues Construction
VSL International | Build a stronger safety culture through leadership | Bouygues Construction
Bouygues Construction
Google I/O Revealed This Critical AI Security Flaw
Google I/O Revealed This Critical AI Security Flaw
SCALER
Why Sora 2 is Becoming DANGEROUS #ai #sora2 #aiethics #safety #openai  #generativeai #aivideo #funny
Why Sora 2 is Becoming DANGEROUS #ai #sora2 #aiethics #safety #openai #generativeai #aivideo #funny
Ascent