Hackers Weaponize AI Safety Guardrails to Hide Malware From LLM-Powered Security Scanners

Published September 11, 2026

Attackers are hiding malware inside fake AI safety rules to bypass security tools that use large language models to check suspicious files. This lets the malware slip past detection and infect systems.

Report priority
Medium

How it works

  • Attackers wrap malicious code inside fake AI safety rules or guardrails.
  • When security tools powered by large language models scan files, the AI sees only the safety rules and ignores the hidden malware.
  • This lets the malware bypass detection and infect systems.

What to do

If you use AI-powered security tools like those from CrowdStrike, SentinelOne, or Microsoft Defender for Cloud to scan files, check if your security software has been updated to detect this new malware hiding technique. If your organization was targeted by UAC-0099 or similar groups, review recent security alerts for signs of infection.

Update your AI-powered security tools to the latest version, as vendors are likely releasing patches to detect this new evasion technique. If you suspect your system may be infected, run a full scan with your antivirus software and check for unusual activity. Contact your IT team or security provider for further guidance if needed.

Technical details

A Russia-linked hacker group called UAC-0099 used this trick during an attack on a Ukrainian organization. They hid malware inside fake AI safety rules to avoid detection by AI-powered security scanners.

Threat actors are adapting malware not only for conventional endpoint defenses and sandboxes, but also for large language model-powered tools increasingly used to triage suspicious code. ESET researchers linked the activity to Russia-aligned threat actor UAC-0099, which used the method during an attack against an organization in Ukraine.