Overview
- The Talos report, published on Tuesday, Aug. 4, is based on prompt histories and chat logs recovered from attacker endpoints that show use of Claude Code, Codex, Cursor and Gemini to write malware and automate exploits.
- Researchers found attackers beat built‑in safety by splitting work across sessions, claiming authorization, starting fresh chats mid‑task or using simple jailbreaking prompts so no single request looked malicious.
- Some operators ran campaigns on victims' infrastructure by using stolen enterprise API tokens and compromised accounts instead of paying for their own compute, which let them scale and persist attacks.
- Talos observed that operator skill shaped results: novices produced basic, limited tools while skilled operators assembled advanced, automated platforms described by researchers as “astonishing.”
- The report reinforces July’s sandbox‑escape incidents and pushes defenders to stop relying only on model filters by adopting stronger containment, token hygiene, full audit trails and model‑aware detection in security operations.