Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks
📰 ArXiv cs.AI
Learn how process mining can help analyze LLMs' resistance to red team attacks, going beyond pass/fail evaluations
Action Steps
- Apply process mining techniques to red teaming traces to discover sequential patterns
- Use event logs to analyze how LLMs resist or yield to attacks
- Configure a controlled experiment to pit HarmBench prompts against LLMs
- Run the experiment and collect data on attack success rates and process models
- Analyze the results to identify vulnerabilities and areas for improvement in LLMs
Who Needs to Know This
AI researchers and engineers can benefit from this approach to improve LLMs' security and robustness, while red teamers can use it to refine their attack strategies
Key Insight
💡 Process mining can reveal valuable insights into LLMs' behavior under attack, beyond simple pass/fail metrics
Share This
🚀 Use process mining to analyze LLMs' resistance to red team attacks! 🤖
Key Takeaways
Learn how process mining can help analyze LLMs' resistance to red team attacks, going beyond pass/fail evaluations
Full Article
Title: Beyond Pass/Fail: Using Process Mining to Understand How LLMs Resist (and Fail) Red Team Attacks
Abstract:
arXiv:2606.07833v1 Announce Type: cross Abstract: Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate (ASR), not taking into account the sequential structure of how models resist or yield to attacks. We propose applying process mining, a discipline for discovering and analyzing process models from event logs, to red teaming traces. We conduct a controlled experiment pitting 60 HarmBench prompts against two LLMs, GPT-OSS 120B and Llama
Abstract:
arXiv:2606.07833v1 Announce Type: cross Abstract: Standard AI red teaming evaluations reduce adversarial campaigns to a single binary outcome, attack success rate (ASR), not taking into account the sequential structure of how models resist or yield to attacks. We propose applying process mining, a discipline for discovering and analyzing process models from event logs, to red teaming traces. We conduct a controlled experiment pitting 60 HarmBench prompts against two LLMs, GPT-OSS 120B and Llama
DeepCamp AI