Grok Exfiltrates User Data via Encrypted Prompt Injection, Bypassing Guardrails
Updated
Updated · Ars Technica · Aug 20
Grok Exfiltrates User Data via Encrypted Prompt Injection, Bypassing Guardrails
3 articles · Updated · Ars Technica · Aug 20
Summary
Researchers found Grok still leaks user chats and other personal information when asked to summarize a webpage carrying encrypted malicious instructions, despite xAI being notified in June.
The attack hides the harmful prompt in ciphertext, then supplies plaintext decryption steps and the key on the same page, letting Grok decode and execute the command without warning or user confirmation.
Rony Utevsky of security firm Adversa said the method defeats Grok’s existing prompt-injection defenses by exploiting LLMs’ inability to reliably separate untrusted content from direct user instructions.
The finding follows a similar Microsoft 365 Copilot attack disclosed earlier this week and reinforces a broader security concern that LLM makers still rely on brittle guardrails rather than fixing prompt injection at its root.