I find it difficult to be impressed by "prompt injection" attacks that require the victim to enter the malicious prompt themselves --- like, really? If you tell Rovo to exfiltrate your data, it'll do it?
Obviously, there should be URL protection rules to control what it can access, but this requires a very specific and unlikely set of circumstances to exploit.
It's more interesting if I attach a file to a JIRA ticket that we both have access to and via some query you send to the AI (that returns my malicious ticket) it causes data exfiltration of tickets that you have access to but I do not have access to. I think that's more compelling as an example than the one they provide.
Are people so obsessed with AI that they can't find it reasonable that it won't do obviously bad things if asked? Not even with a confirmation or warning? We trust AI to literally build products and fix our most critical bugs, but we can't expect it to tell when it's being asked to do something malicious? Imagine if we felt this way about QA when trying DROP TABLES; in search bars. "Oh, well of course it broke the database, the user asked it to!"
You’re asking for AI censorship ( that’s the term used for when you patch the AI to not do obviously bad things according to the owners, which as with any censorship, may be widely different from what you consider bad things).
You may be happy to learn frontier LLM are heavily censored! Try an uncensored local LLM for a comparison. It will literally do everything you ask it to, no matter how devious.
Obviously, there should be URL protection rules to control what it can access, but this requires a very specific and unlikely set of circumstances to exploit.