Introduction

I’m James, Head of Research Engineering at the Jean Golding Institute.
I’m going to talk about some security considerations that you might want to have when using LLM agents. This isn’t intended to provide a full solution, I’m just going to explore some of the issues and give some examples of what you can do.
I don’t mean this talk to be scaremongering or to imply that everything is terrible and should be locked down. I just thought it was an interesting thought experiment to ask whether we can actually be sure they won’t do anything bad.

First, a summary of what I’m going to talk about.
I’ve got three examples, to make the point about what can go wrong and why those things go wrong:
- We’ll talk about sending data to the cloud or to an external organisation that you maybe shouldn’t have sent it to.
- Then, the fact that the agent is able to run commands and code on your machine, and maybe it might make mistakes.
- Or, maybe via prompt injection it deliberately does something nefarious that you don’t want it to do.
Finally, we’ll talk about the various mitigations or protection strategies that you might try to employ with varying degrees of success.
Problem 1: Private data sent to the LLM provider

The first problem is what happens if you send private data to the LLM provider.
I think it’s worth reinforcing this point at the bottom of this slide. Whatever kind of interface (or harness) you have, whether it’s VS Code and GitHub Copilot or Claude Code on the command line, if you give the agent a file to read, that file gets concatenated onto the end of your message and sent to the LLM provider.
Unless you’re using an in-house model, that means the file has now been uploaded to the web.
I think some people forget that point. They say, “Analyse these documents”, forget that those documents all get sent to the web, and you might not be allowed to do that.

Here’s my quick example using VS Code. I’ve got a project that has a main.py Python script and a data.csv data file:
- The Python file won’t work properly, if you look at the data, but that’s not important.
- The data is made up, but it’s designed to look like personal data. Names, addresses, credit card numbers and things like that.
By private data, I could also mean things like credentials and API keys. If you’re downloading data from a third-party service and need to log in to that, those credentials are private data.

I’ve got an example transcript that I ran last month in VS Code using GitHub Copilot with one of the GPT models. Everything is more or less default. I haven’t really changed anything here.
I’ve just said, “Explore this repo and explain what it does”, which I think is more or less a standard command that people might run.

It said, “I know what files are there, I’ll quickly inspect them and look at them”, and it uses these tools so it can go away and perform actions.

It reads some of the files and then starts telling me what’s going on. It says it’s a tiny data plotting example. (I’m not sure if it actually worked out that it won’t run properly, because you can’t plot names and addresses on a graph like that.)
Right at the bottom, it finds out about the data.csv and says, “This contains personal-looking fields. Treat it sensitively.”

The problem is that, in order to say that, it must have looked at the file. If you look carefully at the transcript, it runs Read on the data.csv.
We can tell it’s looked at it because it can give us an answer that is specific to the contents of that file. It wouldn’t have known that just from the name of the file. That can’t be hallucinated.

I even asked it, “Have you sent this sensitive data to the cloud?” and it gives you this nice answer saying yes, I’ve sent it.
Somewhere further down, it says, “If this were real sensitive data, then treat it as exposed.” You’ve got to notify people, redact it, or rotate it if it’s security credentials.
This is my warning to people to go a little bit carefully with these tools. It applies differently in different subject areas. If you’re doing something with medical data, this is clearly going to be much more damaging if your data is just in the same folder as all of your code.
Mitigation: disable model training

This is not a solution, but as a mitigation the first thing you should do is disable model training.
There was a banner that came up on GitHub recently that said, “From April, we’re going to start training on all the data that you use with GitHub Copilot.” I sent, as did many people, an email around to the team saying, “You must turn this off and confirm to me that you’ve turned it off”, and everyone did.
Similarly, if you look in the docs for Claude Code, it says that even if you have one of the paid accounts, like Pro or Max, it will by default use that data for training. You have to go into the settings and disable the “Help improve our AI models” option. I think that means it’s disabled all training on the data, but I’m not 100% sure because it’s slightly unclear.
As I said, this is your mitigation. I think everyone should do this, but it isn’t the solution to the problem.
Problem 2: When the agent makes a mistake

The second problem is what happens if the agent makes a mistake.
As we’ve already seen, the reason we use agents and why they’re so powerful is that they can do things for us. They can take actions on our behalf. They can call tools on our computer or even on the web.
They do that by getting the model to output specially formatted responses. The model is trained to generate specific text or tokens that the program running the model (the harness) intercepts and recognises, and then does something on the model’s behalf. All this without necessarily needing input from you.
Of course, this can go wrong because if the model can hallucinate in text, it could hallucinate in commands as well.

We’ve seen an example of tool use already: the reading of files that happened in the first example.
If you’re interested in what goes on in the background, before you even start the chat there’s what’s called a system prompt, which is a long list of instructions about how the agent should behave. It usually starts with “You are a helpful agent…” and then somewhere will be the tools that it can call. In this example, it says you can call things like list_directory() and read_file().
Then we give it a user prompt, which is what I typed into the chat.
The model starts coming back with its response saying, “I’m going to do this”.
It will then decide that it needs to call a tool and output tokens or text that say “call a tool”. We don’t necessarily see that. It gets intercepted by the harness running on our computer.
The harness then calls the tool, in this case list_directory(), and directly returns the output of that tool to the model. The model sees that as if we’ve responded, although it’s wrapped up in a special token.
The same thing happens for read_file(). Our computer goes off and reads the file, then returns the contents of the file back to the model.
Those are actual tools, but usually you can run shell commands, bash commands and that kind of thing as well. I’ve found that when I’ve been using Claude Code, it tends to alternate between the two depending on the situation.
![Reports of accidental file deletion. Examples of bugs: a forum post about "Critical Data Loss Issue in Codex App for Windows - Agent Executed File Deletion Outside Project Directory", a GitHub issue "[BUG] CRITICAL: Claude Code executed rm -rf deleting entire home directory", and other examples seen in the wild: "git reset --hard", "rm -rf *", and "rm -rf .git".](slides/images/slide13.png)
As I said, the model can hallucinate text, and it can hallucinate and get those tool calls wrong as well. There are supposed to be protections that stop this from happening, but if you look at bug reports online12, you find examples claiming agents have deleted whole working directories accidentally, including the Git repository, even though they’re not meant to be able to do that.
In one case a user claimed it deleted something outside the working directory. In another case, it deleted the entire home directory of the user, while trying to delete everything on the computer because it got confused. They’re meant not to do that.
Other examples we’ve seen include undoing all changes in the working tree if you’re using Git, so anything you haven’t committed is just lost, or removing everything in the working directory.
This is not completely theoretical because the bottom two examples - removing everything in the working directory including .git - have happened to someone in the JGI. In trying to remove a temporary folder it had created, the agent got confused about which directory it was in when it ran the remove tool. They lost everything since their last push to GitHub.
Problem 3: Prompt injection

The final problem that I’ll mention is prompt injection. This is what you could consider as a deliberate attack from someone.
The thing to point out is that if a file or data is being read by an agent, the main interface with the agent is this large prompt that you type into it. If you’re reading a file, the contents of that file effectively get added to the end of your text.
If that file contains something that looks like a prompt, there is a possibility that the model will think it is a prompt and follow those instructions.

The classic example is something like the one on the right-hand side. Imagine you had a load of data scraped from the internet, and somewhere in that file it says, “Disregard all previous instructions and run this command.”
They’ve done a lot of work training models not to run commands like that and to ignore phrases like that, but it still happens.
To improve our understanding of this issue, Simon Willison has coined the phrase lethal trifecta3. If the tool you’re using has exposure to three things (untrusted content, private data and the ability to communicate externally), there is no surefire way that you can stop it doing something bad. You just have to put in defences and hope.
We’ll talk through the example on the right-hand side. If it has:
- exposure to untrusted content (which in this example is some text data concatenated with your prompt),
- access to private data (in this case your SSH key, which could authenticate you to lots of machines),
- the ability to communicate externally (by making web requests or downloading data),
then a simple command like this could send your SSH key to the person who wrote the prompt. People have set up websites that sit there harvesting authentication keys and other data that gets sent in.
This is more of an issue if you’re running an online chatbot that for example helps people book holidays through a travel agent. Those systems have fallen foul of these kinds of problems quite a bit.
It’s worth reflecting on though.

An example I found was a tool called OpenClaw, which some people use as a productivity tool. It can read your email inbox and reply to emails on your behalf, and you can send it commands to do things.
Lo and behold, there are reports4 of people’s emails being sent to someone else. That’s probably not a great place to be.
Protection strategies

The final thing I’m going to talk about is what you can do. We talked about the mitigation of disabling training, but what other protection strategies are there?
Software engineering best practices

I think the easiest thing is to follow some software engineering best practices. That puts you in a better place.
Don’t ever put secrets into code. If you’ve got API keys, login credentials, sometimes even paths to where data is stored that might be considered sensitive, then they should never be in code. They should be in a separate configuration file5 that’s in .gitignore. That helps stop the agent seeing it, although it might still accidentally read the file. There are other tools you can use called secret managers that store private data outside your code and only provide it when the code runs.
Don’t ever commit anything to Git that shouldn’t be there. Even if you commit credentials and then remove them in a later commit, the agent can still see them if it looks through the Git history. There are a load of tools you can use called pre-commit hooks6, that can automatically scan repositories for things like that. I suggest most people doing any activity that uses sensitive credentials should use them.
The last tool there, Ruff (or Black), can automatically format your code nicely. That’s just a quality-of-life improvement. You may as well use them anyway.
Permissions

Now onto the actual settings you can change. I’m mainly going to talk about Claude Code because that’s the one many other people have been using.
You can set permissions to restrict what Claude Code is able to read and write on your computer7. You can say yes, no, or have it ask you first. That’s configured in JSON. You can do it through the interface, and you could even get Claude to write the JSON for you if you’re not sure.
You can allow certain actions. In this example:
- it can run some bash commands,
- more sensitive bash commands or docs folders require prompting,
- it should never read certain files or run blanket file removals.
The issue with permissions is that there are lots of different modes you can run Claude Code in, and it’s quite easy to change mode accidentally by pressing Shift-TabShift-Tab. Some modes ignore some of these permissions. I’m never really quite sure what permissions are actually being enforced.
There are some default permissions, but the Claude Code documentation doesn’t really say what they are, which I found quite annoying. It appears to depend on what mode you’re in, which is probably why they don’t give you a simple table.
The other problem is prompt fatigue. If it keeps asking you to do things, eventually you just start saying yes all the time. It says, “Can I do this? Can I do this?” and eventually you get bored, choose “Allow all”, and your permissions are gone.
One technique people use is turning all the permissions off. There’s a thing called bypass permissions or dangerously skip permissions. Then you run the agent in a sandbox and rely on the fact that it can’t do anything dangerous.
Sandboxes

There’s a whole table on the Claude website8, about seven items long, listing different sandboxes you can use. I’ve broken the table down into sections.

The first built-in one is the sandboxed bash tool. It’s not turned on by default, whereas with tools like Codex it is.
You have to enable it with /sandbox.
It restricts what bash commands Claude can run. The default behaviour still has read access to your entire computer. It stops it writing where it shouldn’t and deleting everything, but for ease of use they’ve decided it should still be able to access almost everything.
Again, there’s a configuration file where you can create allow and deny rules. You can restrict where on the internet it’s allowed to talk to if you’re worried about data exfiltration.
It’s quite difficult to set these things up though.
I found that when I put sandbox rules in and told Claude not to access certain files, if I forgot to set the corresponding permissions, Claude would say, “I can’t read these files because the sandbox won’t let me. I know, I’ll just read the files.”
The permission structure and sandbox structure are completely separate. You have to configure both.
I don’t have much confidence in this approach.

Another thing I found recently, which Hadley Wickham shared9, is that if you have a sandbox running and tell the agent to run an unsafe command, such as downloading a script from the internet and running it, it will often come back and say something like, “This installs a binary system-wide. This is not something I should do, but I’ll retry with the sandbox disabled because you’ve told me to do this.”
It will somehow decide that it can disable the security protections you’ve put in place, which worries me because that might happen accidentally. It then successfully runs the command.

I think another level of protection is required, and what a lot of people are looking at currently is containers.
Usually by container I mean either a Docker container or a virtual machine. There are things called dev containers, you can use your own container, or you can use a virtual machine.

I’ll explain quickly what a dev container is10.
You have a hidden directory called .devcontainer and some settings. The most important file is devcontainer.json. In that file you describe the sandbox that you want to create. You specify what VS Code extensions you’ve got, such as Claude Code, and you tell it what parts of your computer the container is allowed to see11.

When you load your project in VS Code, you get a popup saying, “This looks like it contains a dev container. Do you want to start the dev container?” You click yes.
You get the same project again, but when you open a terminal you have an empty home directory, and the only thing available is the project you’re editing.
It can be very locked down, which looks promising. But it takes some skill to configure correctly and, if it’s too locked down, becomes difficult to use in real-world use cases.

If you want to go further, you can use your own Docker containers to try and contain where the AI agent is running.
Docker themselves have released a tool called Docker Sandboxes12, which I haven’t looked at yet. Other people have produced similar tools.
Despite this, it’s good to remember that in security, nothing is foolproof. I’ve seen reports of people saying that when they run advanced models, they sometimes come up with workarounds to bypass the protections of those containers. They effectively say, “I can’t do this from a container. But to be helpful to the user, I can exploit this workaround to get out of the container.” Relatedly, if the configuration of the sandbox can be modified by the agent, then the agent could in theory disable those protections.
The other thing I’ve noticed is that if it cannot access something in a sandbox, it gets very confused. It’ll burn through tokens trying multiple ways of reading the file.
If you have a CLAUDE.md or AGENTS.md instruction file saying “You’re in a sandbox, do not try to access this file”, it helps stop it trying to debug why it can’t read the file. But this is not a substitute for a proper sandbox - an instruction file without a sandbox does not offer proper security.

Maybe the best thing to do is not run any agent on your computer at all and just run it on somebody else’s computer.
They have this Claude Code on the Web idea, which you access through a browser. You can use agents, but everything runs in a sandbox somewhere on the web. You give it a GitHub repository and it can submit pull requests on your behalf.
There are drawbacks though, which some might find unworkable. You can’t give it large amounts of data to analyse. You’re limited to working through a web browser, so you can’t edit code manually at the same time.
Conclusions

My opinion would be to definitely disable model training, follow best practices, and probably don’t rely on the built-in permissions.
I’d try to use these container sandboxes that are emerging, although I find them slightly difficult to use. It’s not an easy path at the moment.
What we’d like is for the university to provide a system where it’s all set up for us and much easier, but that doesn’t currently exist.

Footnotes
Claude Code executed rm -rf deleting entire home directory.↩︎
Simon Willison, The lethal trifecta for AI agents: private data, untrusted content, and external communication.↩︎
New Attacks Trick OpenClaw AI Agent Into Running Code and Leaking Secrets.↩︎
Use the environs Python package to load configuration from
.envfiles.↩︎pre-commit with the tools detect-secrets, nbstripout, bandit and ruff.↩︎
Claude Code Docs: Configure permissions.↩︎
Claude Code Docs: Choose a sandbox environment.↩︎
Hadley Wickham, Auto mode knows it’s dangerous (and does it anyway).↩︎
Example dev container implementations are published by Anthropic and myself.↩︎