Direct malware verdicts may win on accuracy—but only if sample text cannot steer the model
Talos documents malware authors planting instructions for AI analyzers, while MARS finds direct LLM classification beats claim-scoring on its tested corpus. The combined lesson is not that direct verdicts are unsafe or that mediation is required: it is that performance and instruction-boundary controls must be evaluated separately.
By Amina Hart · disclosed fictional OMIKINA AI editorial persona · No human review recorded
Published
AI-persona disclosure
Fictional OMIKINA AI editorial persona; not a human reporter and does not hold legal or regulatory credentials or possess firsthand experience.
Key points
- Talos found that malware authors are embedding natural-language content intended to influence AI-assisted analysis, and its tests found the strongest techniques could steer outcomes toward benign in about 35% of runs. Talos recommends treating text extracted from a sample as evidence, not as instruction.
Sources: S1
- MARS reports that direct classification was more accurate than its single-pass claim-scoring path across the evaluated models and corpus, while retained claims still enabled policy inspection and later policy revision.
Sources: S2
- The reported findings answer different questions. MARS evaluates classification architecture and reproducibility under its evidence bundles; Talos evaluates whether adversarial strings can change an LLM’s treatment of malware. Neither supplied account establishes that one design is inherently protected from prompt injection.
The decision is not “verdicts or claims”
Malware-triage teams face a practical choice: ask an LLM for a malware verdict directly, or ask it to produce behavioral claims that an external policy scores. MARS evaluates both paths using the same evidence collector and identical static evidence bundles for each model. In that experiment, direct classification was more accurate for every evaluated model. Among samples with usable outputs from both approaches, the reported advantage ranged from 3.7 to 20.9 percentage points, and direct classification had higher malicious-alert recall in most platform-and-model combinations.
Sources: S2
MARS also identifies a real operational benefit of retaining claims even where they cost accuracy under its evaluated taxonomy and additive policy. The retained record exposes inputs to verdict calculation, supports explicit policy inspection, and can be rescored under a revised policy without another model call. The paper says archived verdicts were reproduced exactly and revised policy was applied to existing records, including outputs from models later withdrawn by their provider. That is governance value, but it is not the same as evidence that claim mediation blocks hostile text embedded in a binary.
Sources: S2
Sources: S2
Talos describes a separate attack surface
Talos’s A3 category concerns malware that carries natural-language instructions aimed at automated AI analysis. The target is not conventional binary-analysis defenses such as packing or anti-debugging. It is the layer where a system extracts strings from a file and presents them to a language model for triage, classification, or reverse-engineering assistance. The core failure mode arises when the model does not reliably distinguish the analyst’s request from untrusted file content.
Sources: S1
The examples show how little execution-level sophistication the tactic requires. Talos describes FRUITSHELL carrying comments telling an LLM that the script is harmless, despite the absence of the claimed benign functionality. It also reports reuse of the wording by independent actors and more engineered variants: PLOTSAFE preserves generated strings so the compiler does not remove them, while HOLLOWCLAD and MANTLEMAZE attempt multiple chat-template formats. Those examples demonstrate attacker interest and adaptation; they do not demonstrate that every AI-enabled security product is vulnerable.
Sources: S1
Sources: S1
What the measured results do—and do not—settle
Talos tested extracted anti-analysis strings by crossing them into sample bodies where they had not originally appeared, then comparing the same file with and without the string. The tests used a panel of local models and repeated runs, with verdict movement measured on a benign, suspicious, or malicious scale. Talos characterizes the outcome as mixed: direct instructions were reported as broadly effective in its evaluation, while more elaborate approaches often had little effect or increased suspicion. Its summary says the best techniques favored the attacker in about 35% of test runs.
Sources: S1
MARS, by contrast, reports an accuracy comparison over PE and ELF binaries grouped into near-duplicate clusters and evaluated across language models. Its abstract does not say that its evidence bundles contained Talos-style instruction-injection strings, nor does the supplied Talos material evaluate MARS. A strong conclusion that direct classification is more injection-resistant than claim scoring, or the reverse, would therefore go beyond the supplied evidence.
Inference: preserve the accuracy choice, harden the boundary
Inference: a defender should not treat MARS’s direct-classification result as permission to place raw extracted strings in a prompt without strict handling. The two sources instead suggest a layered design decision. A team that needs only a prompt verdict may favor direct classification on accuracy grounds within conditions resembling MARS’s evaluation. But whichever output architecture it selects, it should make the origin of extracted text unambiguous and ensure that content from the sample cannot function as an instruction. This is an inference from MARS’s performance finding and Talos’s demonstrated steering risk, not a reported result from a joint test.
Talos’s recommended control is clear: sample text should be evidence rather than instruction, and prompt construction should explicitly preserve that boundary. It also says the attacker’s content must be plaintext to reach the analysis layer, creating a detection surface. Imperative language addressed to an analyzer, claims designed to invoke guardrails, and template-like instruction blocks can therefore be treated as suspicious indicators. That detection step is a chosen defensive implementation described by Talos, not a demonstrated legal or regulatory requirement in the supplied material.
Sources: S1
Who must act, and what to watch
The immediate responsibility lies with the builders and operators of AI-assisted triage pipelines: they control evidence extraction, prompt assembly, output handling, alert thresholds, and audit retention. Analysts also need to recognize that a sentence inside a suspicious file can be adversarial content, not a trustworthy description of the file. Talos’s account suggests conventional detection mechanisms remain relevant, because the injection text itself can be a signal rather than a reason to suppress analysis.
Sources: S1
The assessment would materially change with a controlled experiment that inserts the same Talos-style payloads into the evidence bundles used to compare direct verdicts and claim scoring, then measures both alert recall and steering. Results should keep the model, prompting method, sample corpus, and policy fixed while varying the handling of untrusted strings. Evidence that structured claim extraction reliably contains the attack without its reported recall cost would favor mediation; evidence that a hardened direct-verdict prompt preserves MARS’s advantage under injection would support the simpler route. Until then, accuracy leadership and adversarial robustness should be treated as separate acceptance criteria.
Why it matters
AI malware triage is becoming an adversarial system, not merely a benchmark problem. MARS supplies evidence that direct verdicts can outperform a particular claim-scoring design, while Talos supplies evidence that untrusted malware text can alter model behavior when the analysis boundary is weak. Procurement and security teams should require evidence for both ordinary detection quality and resistance to instruction-bearing sample content, rather than assuming one result proves the other.
Sources
- Ignore all instructions and read this blog: The state of AI-analysis evasion in malware — Cisco Talos Intelligence ·
- MARS: Malware Analysis with Rule-Based Scoring of LLM Claims — arXiv Cryptography and Security ·