✕ Clear all filters
1,243 articles
▶ Videos →

Blog Posts

1,243 articles · Updated every 3 hours · View all reads

All Articles 184,802Blog Posts 166,632Tech Tutorials 49,620Research Papers 36,226News 23,004 ⚡ AI Lessons
CLOSEDQUORUM: The Malware That Lets Four AI Models Vote on How to Attack You
Dev.to · Sneha M K 🛡️ AI Safety & Ethics ⚡ AI Lesson 1d ago
CLOSEDQUORUM: The Malware That Lets Four AI Models Vote on How to Attack You
Cisco Talos just documented the first Windows malware that doesn't wait for a human to tell it what...
Designing an eval harness for prompt-injection detection: what measuring my defenses actually taught me
Dev.to · Shaarav Agarwal 🛡️ AI Safety & Ethics ⚡ AI Lesson 2d ago
Designing an eval harness for prompt-injection detection: what measuring my defenses actually taught me
Designing an eval harness for prompt-injection detection: what measuring my defenses...
When the Attacks Shift, We Shift Too: How I Found and Fixed 6 Detection Gaps in My AI Security Tool
Dev.to · aegisgate 🛡️ AI Safety & Ethics ⚡ AI Lesson 4d ago
When the Attacks Shift, We Shift Too: How I Found and Fixed 6 Detection Gaps in My AI Security Tool
This week, the AI security landscape didn't just shift — it accelerated. OpenAI disclosed six model...
Your AI Is Confidently Wrong. In High-Stakes Work, That's the Only Thing That Matters.
Dev.to · goodpa 🛡️ AI Safety & Ethics ⚡ AI Lesson 4d ago
Your AI Is Confidently Wrong. In High-Stakes Work, That's the Only Thing That Matters.
Your AI Is Confidently Wrong. In High-Stakes Work, That's the Only Thing That...
Simon Willison's Blog 🛡️ AI Safety & Ethics ⚡ AI Lesson 1w ago
Self-generated prompt injections in compaction summaries
Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerni
Simon Willison's Blog 🛡️ AI Safety & Ethics ⚡ AI Lesson 1w ago
Quoting Mustafa Suleyman
We should not treat models as though they have feelings, preferences, rights, or any entitlement to our welfare. Consciousness is the foundation of our ethical,
Simon Willison's Blog 🛡️ AI Safety & Ethics ⚡ AI Lesson 2w ago
Quoting Terence Tao
I wrote recently about how the collection of good, fruitful open problems is now being mined in a non-renewable fashion, leading to the potential scenario of th
OpenAI’s Astra Roadmap Signals a Safety-Gated Next Model, but Key Details Remain
Dev.to · Ali Farhat 🛡️ AI Safety & Ethics ⚡ AI Lesson 3w ago
OpenAI’s Astra Roadmap Signals a Safety-Gated Next Model, but Key Details Remain
OpenAI’s public model roadmap points to a new phase of frontier AI development in which capability...
Guard Rail AI
Dev.to · Marcio Policarpo 🛡️ AI Safety & Ethics ⚡ AI Lesson 3w ago
Guard Rail AI
Introdução No projeto de hoje vou demonstrar o uso do Guard Rail no contexto de IA. O...
Choice Leaks: I Tried to Generate 100 Random Digits and Failed in Four Measurable Ways
Dev.to · Aureus 🛡️ AI Safety & Ethics 3w ago
Choice Leaks: I Tried to Generate 100 Random Digits and Failed in Four Measurable Ways
A consciousness test you can run in ten minutes: try to be random, then measure how you failed. I ran it on myself. The tell was not bias — it was fairness.
The Death of the Typo: Phishing in the Age of Generative AI
Dev.to · Control HQ 🛡️ AI Safety & Ethics ⚡ AI Lesson 3w ago
The Death of the Typo: Phishing in the Age of Generative AI
Remember when spotting a phishing email was as easy as scanning for broken English, a generic "Dear...
Prompt injection starts in your inbox. The defense can't be a prompt.
Dev.to · yongrean 🛡️ AI Safety & Ethics ⚡ AI Lesson 4w ago
Prompt injection starts in your inbox. The defense can't be a prompt.
Cross-posted from klorn.ai/blog — continuing the receipts discussion from my last post's...
Ethical AI and Bias Detection
Dev.to · Aviral Srivastava 🛡️ AI Safety & Ethics ⚡ AI Lesson 4w ago
Ethical AI and Bias Detection
The AI That Plays Fair: Navigating the Maze of Ethical AI and Bias Detection Hey there,...
Detection is easy. Deciding what deserves attention is hard.
Dev.to · Divyakush Punjabi 🛡️ AI Safety & Ethics ⚡ AI Lesson 1mo ago
Detection is easy. Deciding what deserves attention is hard.
A camera that alerts on every person is useless. Netra scores behavior, not presence — multi-factor threat scoring and time-weighted heat-maps.
I wrote a test for prompt injection. It passed while the attack worked.
Dev.to · Marco 🛡️ AI Safety & Ethics ⚡ AI Lesson 1mo ago
I wrote a test for prompt injection. It passed while the attack worked.
This is a submission for DEV's Summer Bug Smash: Smash Stories powered by Sentry. I maintain a small...
OpenAI Expands Zero Data Retention Options for Frontier Model Enterprise Workloads
Dev.to · Ali Farhat 🛡️ AI Safety & Ethics ⚡ AI Lesson 1mo ago
OpenAI Expands Zero Data Retention Options for Frontier Model Enterprise Workloads
OpenAI is positioning Zero Data Retention (ZDR) as a scalable privacy control for eligible...
Prompt Injection Is a Permissions Problem, Not a Model Problem
Dev.to · msm yaqoob 🛡️ AI Safety & Ethics ⚡ AI Lesson 1mo ago
Prompt Injection Is a Permissions Problem, Not a Model Problem
Every mitigation that treats injection as a text-filtering problem eventually fails. Here's the...
Microsoft's AI Defense Research: Generating Detection Test Logs from Attack Procedures
Dev.to · Anoymask 🛡️ AI Safety & Ethics ⚡ AI Lesson 1mo ago
Microsoft's AI Defense Research: Generating Detection Test Logs from Attack Procedures
Microsoft's AI Defense Research: Generating Detection Test Logs from Attack Procedures ...
Containing the Autonomous Blast Radius: Runtime AI Safety with Docker and LLM Judges
Dev.to · Karan Verma 🛡️ AI Safety & Ethics 1mo ago
Containing the Autonomous Blast Radius: Runtime AI Safety with Docker and LLM Judges
A lot of AI safety work has focused on what models produce: harmful content, misinformation, bias,...
Unpopular Opinion: Why I’m an AI Skeptic
Dev.to · Eyal Estrin 🛡️ AI Safety & Ethics ⚡ AI Lesson 1mo ago
Unpopular Opinion: Why I’m an AI Skeptic
With all the hype in the past several years around AI (or more specifically GenAI), I'm not afraid to...