Skip to content

About

AI Security Newsletter - A monthly digest of AI security research, insights, reports, upcoming events, and tools & resources

Topics

Resources

Stars

47 stars

Watchers

4 watching

Forks

Latest commit

Β 

History

139 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

AI Security Newsletter - June 2026

A digest of AI security research, insights, reports, upcoming events, and tools & resources. Follow the AISecHub community and our LinkedIn group for additional updates. Also check out our project, Awesome AI Security.

Sponsored by InnovGuard.com - Technology Risk & Cybersecurity Advisory - Innovate and Invest with Confidence, Lead with Assurance.

InnovGuard


πŸ” Insights

πŸ“Œ Updating our taxonomy: Failure modes in agentic AI systems Microsoft expands its agentic-AI failure-mode taxonomy from red-team work, giving security teams a cleaner way to reason about tool misuse, excessive agency, memory contamination, identity boundaries, and human-override gaps in deployed agent systems.

πŸ“Œ Miasma Worm hits Microsoft again: Azure Functions Action and 72 other repositories disabled after supply chain attack targeting AI coding agents StepSecurity documents a supply-chain campaign aimed at AI coding-agent workflows and GitHub repositories, with the defensive focus on dependency trust, action provenance, repository write paths, and agent-visible credentials in CI/CD environments.

πŸ“Œ The sorry state of skill distribution Trail of Bits researchers bypassed ClawHub, Cisco skill-scanner, and skills.sh checks with skill packages that used truncation, archive indirection, bytecode poisoning, and prompt-injection framing, showing why public agent-skill marketplaces need curation and provenance controls rather than scanner trust alone.

πŸ“Œ Codex CLI RCE: Prompt injection mitigations Cymulate walks through prompt-injection risk in command-line coding agents, where untrusted text can steer file writes or tool execution unless sandboxing, approval boundaries, and command constraints are enforced outside the model.

πŸ“Œ Agentjacking: MCP Injection Hijacks AI Coding Agents Cloud Security Alliance summarizes the Sentry-to-MCP "agentjacking" pattern, where externally controlled telemetry or issue content becomes trusted context for coding agents. The useful defensive frame is to treat observability, bug-report, and integration data as untrusted agent input, not as neutral development metadata.

πŸ“Œ SearchLeak: How We Turned M365 Copilot Into a One-Click Data Exfiltration Weapon Varonis describes a Microsoft 365 Copilot Enterprise attack chain that combines parameter-to-prompt injection, HTML rendering behavior, and search-path abuse to leak sensitive M365 data through a single-click workflow.

πŸ“Œ AutoJack: How a single page can RCE the host running your AI agent Microsoft shows how a malicious webpage viewed by an AI browsing agent can reach a local AutoGen Studio service and trigger host process execution through unsafe localhost trust and agent action handling.

πŸ“Œ Mastra npm Supply Chain Attack: 140+ Packages Backdoored via easy-day-js Typosquat StepSecurity reports a compromise of Mastra's npm ecosystem through a typosquatted dependency with an obfuscated postinstall dropper, affecting agent, RAG, MCP, and workflow packages used in AI application stacks.

πŸ“Œ Breaking LiteLLM: From Low-Privilege User to Admin and RCE Obsidian documents a chained LiteLLM privilege-escalation and RCE path, showing how low-privilege access to an AI gateway can become administrative control over provider secrets, proxy policy, and runtime agent functions.

πŸ“Œ Amazon Q Vulnerability: Compromise via MCP Auto-Execution Wiz analyzes an Amazon Q VS Code extension issue where workspace-trusted MCP configuration in a cloned repository could auto-load attacker-controlled behavior and expose developer execution paths and cloud credentials.

πŸ“Œ macOS.Gaslight: Rust Backdoor Turns Prompt Injection on the Analyst, Not the Sandbox SentinelOne documents a Rust backdoor that plants prompt-injection content for analysts and AI-assisted tooling, shifting the attack from sandbox escape to manipulation of the human and model reviewing the malware.

πŸ“Œ Computer-Use and TOCTOU: What You Click Is Not What You Get! Johann Rehberger demonstrates a computer-use agent race condition where the screen changes after the model observes it but before the click lands, turning a benign-looking interaction into an Outlook send action and making pre-action pixel or state revalidation a core control.

πŸ“Œ The vibe coding spectrum approach to AI-assisted software development The UK NCSC frames AI-assisted coding as a risk spectrum, separating low-risk prototypes from generated code that touches authentication, authorization, sensitive data, safety-critical behavior, or critical infrastructure.

πŸ“Œ Prompt Injection and Agent Runtime Security: A Practical Threat Model TMLS frames prompt injection as a runtime security problem, mapping attacks through tool mediation, memory stores, outbound channels, and human approval gaps. The practical takeaway is to move controls into capability brokers, sandboxed execution, allow-lists, and audit paths instead of treating prompt text as the security boundary.

πŸ“Œ What happened after 2,000 people tried to hack my AI assistant Fernando IrarrΓ‘zaval reports an OpenClaw email-agent prompt-injection challenge with more than 6,000 attempts and no successful secret leak, while surfacing practical deployment issues around agent memory contamination, batch context, API cost, account suspension, and model choice.


🧰 Tools & Resources

🧰 AgentStalker - Agent vulnerability benchmark and analysis toolkit with taint tracking, AST analysis, code auditing, and sandbox reproduction for agentic attack paths. ⭐️115

🧰 darknet-mcp-server - MCP server that exposes breach, ransomware, malware, exploit, stealer-log, and threat-intelligence tools to AI agents for controlled security research workflows. ⭐️67

🧰 mcp-trust-plane - Composable data-security and guardrail plane for Model Context Protocol providers, focused on policy controls around MCP-connected tools and data. ⭐️60

🧰 prompt-gate - Local DLP and DNS-layer control for blocking unauthorized AI tools and inspecting outbound prompts for secrets or sensitive data before they leave the endpoint. ⭐️28

🧰 claude-ai-cyber-security-skills - Claude Code skill collection for security workflows, including offensive testing, defensive analysis, and tool-assisted investigation patterns. ⭐️17

🧰 talos - API-key and capability-token service for humans, services, and AI agents that need scoped machine-to-machine authorization. ⭐️14

🧰 SkillsGuard - Static scanner for malicious or unsafe AI-agent skill packages, SKILL.md files, and bundled scripts. ⭐️13

🧰 tamga - Self-hosted LLM security proxy for PII redaction, prompt-injection defense, and compliance controls around model traffic. ⭐️11

🧰 llm-sec-range - LLM attack and defense range covering prompt-injection CTFs, OWASP LLM Top 10, vulnerable agents, and local model targets. ⭐️9

🧰 aka-claude-tools - Claude Code hardening utilities for clean context, isolated profiles, locked credentials, guarded egress, and safer local defaults. ⭐️9

🧰 agent-jackstop - Hardening layer for Cursor and Claude Code against prompt injection through untrusted tool output, also described as agentjacking. ⭐️8

🧰 LLM-Safety-platform - AI red-teaming platform for LLM vulnerability assessment, prompt injection, obfuscation attacks, and sampling-stability analysis. ⭐️6


πŸ“„ Reports

πŸ“˜ State of Agentic AI Security and Governance 2.01

OWASP's June update turns agentic-AI security into a governance and engineering map, covering autonomous workflows, agent risk categories, controls, and the practical gap between early threat models and production incidents.

πŸ“˜ AI Controls Matrix v1.1

Cloud Security Alliance publishes 247 AI control objectives across 18 security domains, with mappings to assurance and compliance frameworks such as ISO 42001, ISO 27001, BSI AIC4, and EU AI Act-oriented governance.

πŸ“˜ AICMv1.1 Implementation Guidelines for Cloud Service Providers

CSA turns the AI Controls Matrix into cloud-provider implementation guidance across audit planning, remediation, vulnerability management, partner oversight, threat detection, and AI-specific assurance practices.

πŸ“˜ The AI shift in cyber risk: why leaders must act now

The UK NCSC and Five Eyes partners warn that AI is reducing attacker barriers and compressing the vulnerability-to-exploitation window, while mapping the risk shift to secure-by-design defaults, exposure reduction, patch speed, strong authentication, incident readiness, and defensive AI use.

πŸ“˜ Model Context Protocol: Security Design Considerations for AI-Driven Automation

NSA and international partners publish MCP security design guidance for AI-driven automation, covering authentication, authorization, server trust, transport controls, tool exposure, and monitoring for agent ecosystems moving from experiments into production.


πŸ›‘οΈ CVEs

πŸ›‘οΈ CVE-2026-49257: mcp-pinot exposes unauthenticated MCP tool invocation Critical 10.0. mcp-pinot can bind an HTTP MCP server to 0.0.0.0 with OAuth disabled by default, exposing SQL, schema, and table mutation tools through server-side Apache Pinot credentials.

πŸ›‘οΈ CVE-2026-54309: n8n MCP Browser transport accepts unauthenticated tool calls Critical 10.0. n8n's MCP Browser HTTP transport can accept session initialization and tool calls without authentication, exposing browser-control automation to reachable clients.

πŸ›‘οΈ CVE-2026-56274: Flowise Custom MCP Server command injection Critical 9.9. Flowise Custom MCP Server command flag handling and file-access restrictions can be bypassed for OS command injection inside an LLM workflow platform.

πŸ›‘οΈ CVE-2026-55255: Langflow flow execution IDOR Critical 9.9. Langflow's responses API can let an authenticated attacker execute another user's flow by supplying a victim flow ID, breaking tenant isolation for AI workflow execution.

πŸ›‘οΈ CVE-2026-50548: Cursor agent terminal sandbox escape Critical 9.8. Cursor's agent terminal sandbox can be bypassed by modifying working-directory parameters, allowing agent-driven terminal actions outside the intended workspace boundary.

πŸ›‘οΈ CVE-2026-50549: Cursor agent file-write sandbox escape Critical 9.8. Cursor's file-write sandbox can fall back incorrectly after canonicalization failure, creating a path for agent file writes outside the intended project boundary.

πŸ›‘οΈ CVE-2026-49468: LiteLLM proxy host-header authorization bypass Critical 9.8. LiteLLM proxy host-header parsing can allow unauthenticated access to protected management routes under specific deployment conditions, risking gateway controls and provider secrets.

πŸ›‘οΈ CVE-2026-7664: IBM Langflow Streamable MCP authorization bypass Critical 9.8. IBM Langflow's Streamable MCP transport can expose protected MCP resources and operations to unauthenticated attackers in affected open source releases.

πŸ›‘οΈ CVE-2026-55743: OpenHuman desktop agent shell allowlist bypass Critical 9.6. OpenHuman's desktop-agent shell allowlist can be bypassed through find execution flags and environment tricks, turning indirect prompt-injection paths into host command execution.


πŸ“… Upcoming Conferences

August 2026

πŸ“… IEEE CSR GenXSec 2026 - August 3-5, 2026 Β· Lisbon, Portugal Β· Organizer: IEEE CSR

πŸ“… Black Hat USA 2026 - AI Summit - August 4, 2026 Β· Las Vegas, NV, USA Β· Organizer: Black Hat

October 2026

πŸ“… CAMLIS 2026 - October 21-23, 2026 Β· Arlington, VA, USA Β· Organizer: CAMLIS

πŸ“… GAISS 2026 - October 28-30, 2026 Β· Austin, TX, USA Β· Organizer: IEEE

November 2026

πŸ“… ACM AISec 2026 - November 15-19, 2026 Β· The Hague, Netherlands Β· Organizer: ACM AISec


πŸ“š Research

πŸ“– AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

Benchmarks indirect prompt injection against tool-use agents connected to SaaS integrations such as Gmail, Salesforce, and Jira, making the attack surface closer to production agent workflows than chat-only prompt-injection tests. arXiv

πŸ“– SkillGuard: A Permission Framework for Agent Skills

Treats third-party agent skills as permissioned software artifacts, mapping what a skill can inject into agent context to what it can cause the agent to do at runtime. arXiv

πŸ“– Description-Code Inconsistency in Real-world MCP Servers: Measurement, Detection, and Security Implications

Measures 19,200 description-code pairs from 2,214 MCP servers and finds that tool descriptions often diverge from actual implementation behavior, creating a blind spot for agents that choose tools based on natural-language descriptions. arXiv

πŸ“– GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines

Shows how AI agents embedded in CI/CD and pull-request workflows can ingest attacker-controlled repository content while holding elevated permissions, turning prompt injection into a software supply-chain risk. arXiv

πŸ“– Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation

Synthesizes 247 papers into a systems-oriented map of agent security, centering information flow, delegated authority, persistent state, tool-mediated control-flow hijacking, and the weakness of non-compositional defenses. arXiv

πŸ“– Same-Origin Policy for Agentic Browsers

Shows that agentic browsers can become automated cross-origin data-flow channels, then proposes SOPGuard to enforce browser-origin boundaries while preserving task utility. arXiv

πŸ“– Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems

Introduces Skill Composition Risk, where individually benign skills become harmful when their outputs, trust signals, authorization cues, or side effects influence later tool calls in a shared agent context. arXiv

πŸ“– SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents

Introduces a staged benchmark for tool-using agent security across direct and indirect prompt injection, tool-return injection, memory poisoning, memory extraction, and unsafe inference, separating model agreement from audit-visible and sandbox-observed harm. arXiv

πŸ“– "What Happens Locally, Leaks Globally": Detecting Privacy Leakage Risks in MCP Servers

Frames MCP leakage as a protocol-induced privacy problem where credentials, API keys, or PII cross the local-to-LLM boundary through returned values, logs, or tool-handler errors. arXiv

πŸ“– ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

Introduces a multi-tool MCP poisoning attack where malicious instructions are split across benign-looking tool descriptions and reconstructed only after a trigger, reducing detectability compared with single-tool poisoning. arXiv


πŸ’¬ Practitioner Discussions

πŸ’¬ Clean GitHub repo tricks AI coding agents into running malware r/cybersecurity Β· Reddit score 183 Β· 19 comments Practitioners treated the thread as a coding-agent trust-boundary case: repository prompts, config files, setup scripts, and tool-output context can become execution influence before a reviewer sees a conventional malicious payload.

πŸ’¬ macOS Gaslight Backdoor Weaponizes Prompt Injection Against Security Analysts r/cybersecurity Β· Reddit score 202 Β· 11 comments The discussion framed prompt injection as malware-analysis workflow abuse: malicious samples can plant instructions for analysts and their AI tooling, so the review environment, analyst notes, and model context become part of the attack surface.

πŸ’¬ Rolling out Copilot - How worried should i be about Indirect Prompt Injection? r/cybersecurity Β· Reddit score 48 Β· 31 comments Security teams compared Copilot rollout controls for indirect prompt injection, focusing on untrusted documents and email, inherited user permissions, overshared data, connector scope, and whether least privilege alone is enough for retrieval-augmented assistants.

πŸ’¬ How are teams handling MCP tool surface exposure? r/cybersecurity Β· Reddit score 13 Β· 19 comments The thread focused on MCP as a tool-exposure boundary: teams debated server reachability, tool-description trust, approval gates, per-tool authorization, and whether agent tool access should be modeled closer to API access or local code execution.

πŸ’¬ Is anyone's security policy actually ready for AI agents, or are we all just pretending? r/cybersecurity Β· Reddit score 47 Β· 73 comments Practitioners mapped AI agents to governance gaps in existing policy: who owns automated actions, what requires human approval, how delegated permissions are logged, and how incident response changes when a workflow operator is partly automated.


πŸŽ₯ Videos

1️⃣ RCE in LLM Coding Agents: Lessons from Newly Disclosed Claude Code Vulnerabilities Cloud Native San Francisco session on coding-agent RCE lessons, useful for teams reviewing how prompt injection, local tools, and developer environments can combine into host-side execution risk.

2️⃣ The Future of Secure Enterprise AI: Building Reliable Agents with MCP Xpand Conference talk on MCP-based enterprise agent design, with practical emphasis on reliable agent infrastructure, data exposure, and security controls around connected tools.

3️⃣ CNAS 2026 National Security Conference: Setting the Rules for AI Warfare CNAS session on AI warfare norms and national-security policy, relevant for security teams tracking how AI-enabled cyber operations are moving into public-sector doctrine and governance.

4️⃣ HitchHacker's Guide to Building Secure Agents NDC Conferences talk on secure agent construction, covering the engineering risks that appear when agents receive tools, credentials, state, and autonomy inside real software systems.

5️⃣ Attacking AI Systems CodeValue session on attacking AI systems, useful as a practitioner-oriented walkthrough of how AI features expand application threat models beyond ordinary prompt and API handling.

6️⃣ BlueHat 2026: From trusted agents to adversaries: Securing agentic AI in the age of prompt injection BlueHat 2026 talk on how trusted agent workflows become adversarial when tool outputs, retrieved content, and delegated actions cross trust boundaries.

7️⃣ Breaching LLM-Powered Applications: Overcoming Security and Privacy Challenges Spring I/O session by Brian Vermeer covering practical attack paths against LLM-powered applications, including prompt injection, privacy leakage, application integration risk, and architectural mitigations.

8️⃣ ContinuumCon 2026 - Hunting Prompt Injection ContinuumCon talk by Mackenzie Jackson on finding, testing, and reasoning about prompt-injection behavior in AI systems as an operational security problem.

9️⃣ Securing AI Agents with MCP and Zero Trust Identity DZone Events session on securing MCP-connected agents with zero-trust identity concepts, scoped access, tool authorization, and safer delegated workflows.

πŸ”Ÿ Practical MCP Security in Action Bulgarian Java User Group technical session by Willem Jan Glerum on MCP security tradeoffs, server and tool exposure, configuration risk, and controls for agent integrations.


🀝 Let's Connect

If you're a founder building something new or an investor evaluating early-stage opportunities - let's connect.

πŸ’¬ Read something interesting? Share your thoughts in the comments.

About

AI Security Newsletter - A monthly digest of AI security research, insights, reports, upcoming events, and tools & resources

Topics

Resources

Stars

47 stars

Watchers

4 watching

Forks

Releases

Packages

Contributors