Code Red: The 55 New Ways Self-Learning AI Can Be Hacked
Key Takeaways
Security risks of self-learning AI systems are examined with two papers on META Self-learning and cybersecurity
Original Description
Self-learning and self-evolving AI systems in a loop are a current hype. Here we examine two papers, that show clearly that these kind of no-human-in-thee-loop pose significant drawbacks (META Self-learning) and massive new security problems (cybersecurity).
IF we let a self-learning AI system learn to learn itself new domain knowledge, through the classical reinforcement learning w/ DPO, then the AI, that steers the self-learning fails to build the RL learning complexity itself. So currently, according to META, no self-RL for self-learning systems.
IF we say, wait, the AI harness itself, as an agentic filesystem, will perform the self-learning for AI agents, then the second new ArXiv pre-print shows, we open up the box of pandora regarding security, since the system is self-modifying itself, numerous new attack surfaces the Ai system is offering to external adversary systems. More than 25 new cybersecurity cells are identified in a new matrix representation of cybersecurity for self-learning systems, with 2 or more attack vectors each. I counted 55 at first run. How many can you detect with the new self-learning looped AI system complexities?
All rights w/ authors:
Repeated post-training is not
Self-improving: Diagnosing Scientific
Amnesia in Continual DPO Pipelines
Jianzhe Lin, Fei Wang, Xiaolin Li, Rajeshkumar Golani, Jubin Chheda
from
MetaAI
Safety in Self-Evolving LLM Agent Systems: Threats,
Amplification, and Case Studies
Ruixiao Lin1,2,†, Xinhao Deng2,3,†, Qingming Li1, Jianan Ma4,2, Yunhao Feng2, Yuqi
Qing2,3, Zhenyuan Li1, Yechao Zhang5, Shiwen Cui2, Changhua Meng2,
Tianwei Zhang5, Xingjun Ma6, Qi Li3, Ke Xu3, Shouling Ji1,∗
from
1 Zhejiang University
2 Ant Group
3 Tsinghua University
4 Hangzhou Dianzi University
5 Nanyang Technological University
6 Fudan University
#aisafety
#cybersecurity
#scienceexperiment
#artificialintelligence
#futureai
#futuretechnology
Watch on YouTube ↗
(saves to browser)
Sign in to unlock AI tutor explanation · ⚡30
More on: AI Security
View skill →Related Reads
📰
📰
📰
📰
Rod Johnson Is Back - and He's Bringing AI Agents to Java
Dev.to · Md Jamilur Rahman
Your agent knows your preferences. It just never uses them
Dev.to · Nam Bok Rodriguez
Context Compression: Making AI Agents Forget Without Losing the Plot
Dev.to · Rijul Rajesh
Built a tool that datacenter cooling layouts optimiser
Reddit r/artificial
🎓
Tutor Explanation
DeepCamp AI