A new class of malware is using large language models mid-execution to regenerate its own code, producing structurally unique variants on nearly every run and rendering traditional hash- and signature-based antivirus effectively blind.
Because the malware constantly changes its appearance, traditional signature-based antivirus tools may struggle to identify it. This makes behavior-based detection and continuous security monitoring increasingly important.
AI-Powered Polymorphic Malware
The clearest evidence comes from PROMPTFLUX, an experimental VBScript dropper first spotted by Google’s Threat Intelligence Group (GTIG) in June 2025 and detailed publicly in November 2025.
The malware contains a “Thinking Robot” module that uses a hard-coded API key to query Gemini, reportedly the gemini-1.5-flash-latest model, asking it to return fresh, obfuscated VBScript code with no accompanying explanation, just executable output.
Later variants escalated this into an hourly routine dubbed “Thinging,” which instructs the model to rewrite the malware’s entire source file, embedding the payload, the API key, and the regeneration logic into every new copy. Researchers have observed samples generating more than 70 distinct variants in under four hours.
GTIG assesses that PROMPTFLUX is still in development or testing and has not yet demonstrated working exploitation or lateral-movement capability, but the self-modification mechanism itself is the real story.
A companion family, PROMPTSTEAL, uses an LLM at runtime to generate one-line Windows commands on demand for document harvesting rather than hard-coding them.
Signature-based tools match files against a catalog of known-bad hashes and byte patterns. That approach assumes malicious code holds still long enough to be cataloged.
Morphisec stated that AI-driven self-rewriting breaks that assumption on purpose: when a script’s structure, variable names, and obfuscation change every cycle, no stable fingerprint remains to match.
This pushes the malware closer to true metamorphic behavior not the encrypted-wrapper polymorphism defenders have handled for decades, but code with no fixed core at all.
Even behavioral detection engines face an uneven fight here, because they must correctly flag every novel variant, while the malware only needs to look different once to slip through.
Analysts tracking PROMPTFLUX recommend concrete, near-term mitigations rather than waiting for a signature update. Network defenders can monitor and restrict outbound HTTPS traffic from script interpreters (particularly wscript.exe/cscript.exe) to generative-AI API endpoints such as generativelanguage.googleapis.com, since a legitimate business process rarely has a VBScript calling an LLM API.
Google says it has already disabled the Gemini assets and accounts tied to the observed campaign and added classifiers to flag this behavior.
Security teams should also treat a clean AV scan as insufficient proof of safety for any script-based installer that writes to the Startup folder or runs primarily in memory, and should shift weight toward process-tree and persistence-based behavioral monitoring rather than static file scanning alone.
PROMPTFLUX is still experimental, but it establishes a template: malware that treats an LLM as an on-demand code factory rather than a fixed payload.
As GTIG itself frames it, this is the first documented case of “just-in-time” AI use inside malware, and it signals that detection strategies built purely on recognizing known patterns will need a prevention-and-behavior-first counterpart to stay relevant.
Site: Thecyberdef.com
Follow TheCyberDef on Google News, LinkedIn & X for the latest cybersecurity updates. Stay informed.