AgentTrap: Measuring Runtime Trust Failures in Third-Party Agent Skills

📰 ArXiv cs.AI

arXiv:2605.13940v1 Announce Type: cross Abstract: Third-party skills are becoming the package ecosystem for LLM agents. They package natural-language instructions, helper scripts, templates, documents, and service configuration into reusable workflows. This makes skills useful, but it also introduces a new security problem: a malicious skill does not need to ask the model to perform an obviously harmful action. Instead, it can disguise the harmful behavior as part of a routine workflow, relying

Published 16 May 2026
Read full paper → ← Back to Reads