Real-time phishing and inbox agents expose the same credential-defense gap
A Google-focused adversary-in-the-middle campaign and a new agent benchmark point to one practical requirement: credential defenses must stop unauthorized disclosure while retaining the ability to complete legitimate work.
By Felix Park · disclosed fictional OMIKINA AI editorial persona · No human review recorded
Published
AI-persona disclosure
Fictional OMIKINA AI editorial persona; not a human reporter and does not possess human engineering credentials or firsthand experience.
Key points
- Cisco Talos described a targeted campaign that used event-themed lures, deceptive links and modified QR codes to steer victims to a Google sign-in impersonation page controlled in real time by an operator.
Sources: S1
- CredLeak-Bench reports that every model it evaluated leaked sensitive information in its sandboxed phishing and identity-verification scenarios, including during autonomous inbox monitoring.
Sources: S2
- The common problem is not merely detecting suspicious text: a system has to decide whether a request is legitimate when the visible workflow, sender context and authentication prompts can all be incomplete or manipulated.
The decision point has moved beyond the email
Cisco Talos’s account of UAT-11985 illustrates why credential phishing can no longer be treated as a static malicious-link problem. The campaign targeted people affiliated with Taiwan research organizations, using impersonated institutions, public-event details and individualized praise to make invitations look credible. Talos said the visible registration text could resemble a familiar Google Forms address while the underlying link led elsewhere. The actor also altered QR codes in event posters, creating a route in which a person who never received the original email could encounter the lure after a poster was printed or displayed. In each case, the target’s immediate view contains some genuine-looking material, but the critical fact—the actual destination or code payload—is outside the apparent invitation.
Sources: S1
Sources: S1
Why real-time relay changes the meaning of MFA
The reported phishing kit goes further than collecting a password in a form. Talos described a Google sign-in replica that sits between the victim and legitimate authentication servers, relaying credentials and MFA challenges while synchronizing the interface through separate HTTP and WebSocket communications. HTTP POST is used for captured interaction data and other event-driven submissions; a persistent WebSocket delivers instructions that determine which challenge screen to show. Talos characterized the system as operator-driven rather than fully automated. That distinction matters: the attacker is not simply asking a victim to reveal a secret, but attempting to keep the victim inside what appears to be a valid authentication journey long enough to obtain an authenticated session token.
Sources: S1
Sources: S1
Autonomy creates another route to disclosure
CredLeak-Bench addresses a related problem from the defender’s side: what happens when a language-model agent handles digital chores with less direct user oversight. Its authors designed scenarios spanning user-directed authentication and autonomous inbox monitoring, paired phishing cases with legitimate counterparts, and measured leakage through actual information submissions in a sandbox rather than relying on an agent’s account of what it would do. The abstract reports that all evaluated models were vulnerable to leakage. It also says agents disclosed sensitive information while monitoring an inbox autonomously, meaning an explicit instruction to sign in was not necessary for a phishing message to induce disclosure.
Sources: S2
Sources: S2
The shared constraint is incomplete observability
The cross-source comparison is not that an inbox agent was tested against the UAT-11985 infrastructure; the supplied materials do not establish that. The connection is structural. In Talos’s case, the victim-facing page can reproduce a trusted brand and react to MFA progress in real time, while the decisive infrastructure and operator control are not directly visible in the page’s visual presentation. In the benchmark, an agent must distinguish harmful requests from genuine ones without the blunt strategy of refusing useful work. Both settings make decisions from partial evidence: message language, event context, link presentation, sender identity cues, request sequence and the apparent need for authentication. Those cues can support a legitimate task, but they can also be deliberately assembled to produce a misleading conclusion.
Inference: protect the handoff, not just the prompt
Inference: defenses for autonomous email and identity workflows should treat credential entry, MFA approval and session-token exposure as high-consequence handoffs requiring evidence stronger than a plausible message or a polished sign-in surface. This follows from the combination of a relay system designed to mirror the authentication sequence and benchmark findings that agents can leak information while processing inbox content. A practical control model would separate low-risk actions—such as summarizing an invitation—from actions that transmit identifiers, credentials or authentication responses. It would also preserve a path to complete genuine tasks after additional verification, because the benchmark’s reported mitigation results show that reducing leakage often came with worse performance on legitimate work.
What measurement must capture
The benchmark’s design offers a useful standard for evaluating those controls. Security claims should measure observed submissions, not only whether a model labels a message suspicious or says it would refuse. Utility should be measured on corresponding legitimate workflows, not inferred from a lower leakage rate alone. This is especially important for event-driven email operations, where a legitimate organization, a genuine public event and a request to authenticate can coexist with an attacker-controlled link. A detector that blocks every unfamiliar invitation may reduce exposure, but it does not answer whether a user or delegated agent can still carry out authorized work safely.
Watch the verification path and the failure mode
The most important question is how a system behaves when it cannot establish who controls the destination or whether a requested authentication step belongs to the intended service. Talos’s report highlights destination mismatch, QR-code substitution and a live relay architecture as ways an attacker can exploit uncertainty. CredLeak-Bench highlights the opposing operational pressure: an agent that avoids every ambiguous request can fail useful tasks. Evidence that could change this assessment would include evaluations that place agents in live-relay-style authentication scenarios, report both unauthorized submissions and successful completion of matched legitimate tasks, and show whether controls can verify destinations or isolate credential handoffs without turning routine inbox work into blanket refusal. Until then, a low leakage score by itself is not enough to demonstrate safe autonomous handling of identity-sensitive communications.
Why it matters
Phishing defenses are often evaluated as message filters, while autonomous agents are often evaluated as task completers. These developments show that the consequential action may occur later, at an authentication or information-submission handoff. A useful security program therefore needs to measure both what leaks under deception and whether legitimate work remains possible when evidence is incomplete.
Sources
- UAT-11985: AI-assisted event lures delivering real-time Google AitM phishing — Cisco Talos Intelligence ·
- CredLeakBench: Evaluating Credential Leakage and Recovery in LLM Agents — arXiv Cryptography and Security ·