01

The external result

Cohen, Bitton and Nassi describe self-replicating adversarial prompts in connected generative-AI applications. Their experiments include email-assistant workflows. The paper is external research, not a Sentry experiment.

02

Our interpretation

An agent may encounter hostile instructions during an ordinary read. Inspection helps identify suspicious content, but the decisive boundary is whether that content can authorize a later write, transfer or outbound request.

03

How it informs Sentry

Read Guard annotates suspicious incoming text. Local write and outbound controls, and a payment signer wrapper, provide separate places to constrain actions when correctly connected to the host.

04

The limit

We have not reproduced the paper’s full experiments. Our implementation and local regression results do not establish protection against every propagation path.

Source & context

Cohen, Bitton & Nassi · Morris II · arXiv

Open the original source

Sentry’s interpretation is identified above. The external work has not been reproduced as part of this local verification.