I think I and lot of others lack a clear instruction how to use it and what are the limitations. With docker/containers we are sold the promise of sandbox (not true, but regardless), and with existing systemd user level separation is really hard to grasp.
Interesting!
Personally I'm not eager to spend tokens on having my own harness, sandbox etc.
It's much much more feasible by spending tokens. But it does feel like a distraction and that I end up owning it. As in: a continuing distraction.
So I've been experimenting for these purposes with omnigent as a way to be less locked into a single provider and for its sandbox abstraction. And I've also tried openshell for sandboxing. Hoping those two will keep improving.
Super cool! I have a similar system that incorporates Forgejo and Pi into a system that functions a little bit more like a software team: https://bakedpi.dev
It's definitely more manual than some other solutions, but I prefer to know what my agents want to do before they do it.
It's not exactly isolated in their own VMs though. It's using Docker which isn't the safest solution out there but also... well, it is what it is.
Running coding agents securely within sandboxed environments on self-hosted infrastructure hits the sweet spot for sovereign and private development workflows. Great implementation.
I'll be sure to check this out but in the meantime, I'd like to ask: Isn't it generally a good practice to run coding agents in a VM? It provides a security boundary so that the agent is restricted to the VM.
How does this compare? I see it's an "isolated sandbox", but what exactly does that mean?
I've been working on a personal alternative harness to Claude Code and a very very small part of it was doing precisely this :) The project has been loads of fun.
To be fair: I support sandboxing and microVMs.
I'm pretty sure there's thousands of us, all implementing our own harnesses. The era of truly personal computing!
It offers various abstractions for running agents in your cluster. It can be defined with simple inline text, or be your own agent image such as one that contains pi. It adds lots of useful features like APIs, UI, observability and Substrate integration.
With some levity, I'll also add that as a bilingual English/Spanish speaker, my brain seizes a bit at the name. Why? That 'pi' to a Spanish-infected mind pronounces it like 'pea'. Adding 'pod' then naturally takes me to a 'peas in a pod' association. Split-language hall of mirrors. There oughta be a word for this?
Anymore, more a me problem than a name problem. Just wanted to share.
Hi, I think I might be your target audience, I currently run pi on the server in my basement, in a minimalist docker container that gives it access to my code workspace and a config dir. Give me an idea what benefit I get on top of that by using the self-hosted pi pod?
For you the main benefit would probably be portability via the mobile app. The mobile app is not released yet but you can build it yourself from the swift/kotlin source in the github repository.
Otherwise, the main benefit here is for people to be able to clearly define all of their pi and agent config at the org, user, and project level. This most directly benefits teams of people working together, but I also find it helps me as a single user working on multiple projects on multiple devices.
I have a similar project you might be interested in if you use ACP via Zed or something; I need to make a Show HN post at some point.
The gist is that it pretends to be an ACP agent to Zed, but it actually spins up a Docker container on your PC and proxies the ACP connection over websocket to the agent. It supports bind-mounting your Pi config like you're doing (as well as copying files, so changes don't propagate to your host). It also has early plugin support for your ACP connection. It's basically an nginx proxy for ACP, I just made a plugin that will automatically kill a session if it detects OpenRouter secrets in agent output, tool calls, file access or shell commands.
It doesn't currently support remote targets like your server, but it's on the list.
Another notable difference is that it creates a new container every time you connect so you get an idempotent environment, though I am also working on a persistent style. Everything supports it, I just need to make the host-side binary support finding a container to use before it tries starting a new one.
I am developing tarvis.io that allows you to self host apps without managing the servers. In addition to that it also comes with cloud dev env where you can do coding and spawn dev servers and you can host same app in your tarvis workspace as a self hosted app. Tarvis takes care of security, storage, backups and ssl domains. Each app auto gets ssl domain and only you can access your apps. Give it a try!
I am running deepseek harness as self hosted app on tarvis for myself. You can run pi as self hosted harness there too by telling the agent to set it up and then access on the go from the browser.
The barrier to create this myself is so low that 1: I can do it and 2: bad actors can do it. I’d like to use a shared tool that will get iterated on, but I just don’t want to risk it at this point given I can get “good enough” doing it myself.
It depends on your threat model but I think it's a good reminder either way. I recently discovered SmolVM <https://github.com/smol-machines/smolvm> and it looks like a pretty good sweet spot for usability and isolation. There are so many sandboxing technologies to pick from and understand.
I've been using (rootless) Podman which gives me some basic assurances that it will stay in its designated directory and not run tools on my system directly but I have no limits on the network and with an internal UID/GID of 0:0 I have not done myself any favours. This is the same level of protection one would implement to keep a poorly written bash script from wreaking havoc and that's about it.
I focus on being the batteries-included approach for microVMs. So network is off by default, and you can allow specific hosts (DNS is filtered too), so an agent can reach its model API and nothing else.
And then I also put a lot of work in the jailer-style hardening around the VMM process itself: seccomp allowlist, Landlock, separate unprivileged uid per VM, and cgroup limits.
So that users have security & knobs right out of the box.
fwiw i used to operate an AWS service using firecracker.
not op but firecracker exists, as does docker sbx (not container!),
firecracker is a lot more of a headache to setup than docker sbx (not container!) but if you can get it going it probably “feels” the best
of a different variety some people feel better using stuff like bubblewrap/fire jail but idk if these are still microvm as opposed to the above
but it’s my opinion (perhaps completely criticizable), that sandboxing for personal home use machines is 1) somewhat overkill 2) somewhat theatre 3) psychologically sometimes exhausting and unrealistic always assuming the worst is going to happen and 4) not really worth the time and creates quite a bit of friction that there wasn’t previously. In the end I spent more time tinkering with the sandbox to get it “just” right that it drained my time actually using the agents so …
The main benefits for using pi pod versus running pi on a server is the native mobile apps and environment composability. The cli experience has also been made really nice with pi pod, the tui is rendered locally so user key input does not lag over the ssh connection.
I use a T3 Code pod in a kubernetes Deployment (sadly no official Docker image), with a PVC so configs and cloned repos survive pod restarts. Seems to work quite well with my Claude subscription (no GPU in my homelab yet)
I spent lots of time on this topic, and had a similar path. Meanwhile the tool also supports a quick way to add custom or local providers to agents: https://vibepod.dev/news/vibepod-cli-0-24/
This looks very interesting, letting people run these in any box they own. I very much agree with the sentiment that there are no proper tools that let you run any coding agent without being locked to a single provider. I just want to run opencode or pi somewhere in a box without having to spin up all of them in my machine, let alone being able to trigger them from within another prouduct as a background agent.
Would pi-pod allow me to standardize pi config on a team/project level so that other people in my org can also use them on the same private infra?
it has all the options you would ever want, and is actually deeply aware of the kernel capabilities.
every day someone comes with a new sandbox solution because they are too lazy to read one page documentation.
So I've been experimenting for these purposes with omnigent as a way to be less locked into a single provider and for its sandbox abstraction. And I've also tried openshell for sandboxing. Hoping those two will keep improving.
It's definitely more manual than some other solutions, but I prefer to know what my agents want to do before they do it.
It's not exactly isolated in their own VMs though. It's using Docker which isn't the safest solution out there but also... well, it is what it is.
How does this compare? I see it's an "isolated sandbox", but what exactly does that mean?
To be fair: I support sandboxing and microVMs.
I'm pretty sure there's thousands of us, all implementing our own harnesses. The era of truly personal computing!
https://kagent.dev/docs/kagent/0.x/concepts/agents/
It offers various abstractions for running agents in your cluster. It can be defined with simple inline text, or be your own agent image such as one that contains pi. It adds lots of useful features like APIs, UI, observability and Substrate integration.
Anymore, more a me problem than a name problem. Just wanted to share.
Otherwise, the main benefit here is for people to be able to clearly define all of their pi and agent config at the org, user, and project level. This most directly benefits teams of people working together, but I also find it helps me as a single user working on multiple projects on multiple devices.
The gist is that it pretends to be an ACP agent to Zed, but it actually spins up a Docker container on your PC and proxies the ACP connection over websocket to the agent. It supports bind-mounting your Pi config like you're doing (as well as copying files, so changes don't propagate to your host). It also has early plugin support for your ACP connection. It's basically an nginx proxy for ACP, I just made a plugin that will automatically kill a session if it detects OpenRouter secrets in agent output, tool calls, file access or shell commands.
It doesn't currently support remote targets like your server, but it's on the list.
Another notable difference is that it creates a new container every time you connect so you get an idempotent environment, though I am also working on a persistent style. Everything supports it, I just need to make the host-side binary support finding a container to use before it tries starting a new one.
https://abyss.scurry.io/
I am running deepseek harness as self hosted app on tarvis for myself. You can run pi as self hosted harness there too by telling the agent to set it up and then access on the go from the browser.
I've been using (rootless) Podman which gives me some basic assurances that it will stay in its designated directory and not run tools on my system directly but I have no limits on the network and with an internal UID/GID of 0:0 I have not done myself any favours. This is the same level of protection one would implement to keep a poorly written bash script from wreaking havoc and that's about it.
I focus on being the batteries-included approach for microVMs. So network is off by default, and you can allow specific hosts (DNS is filtered too), so an agent can reach its model API and nothing else.
And then I also put a lot of work in the jailer-style hardening around the VMM process itself: seccomp allowlist, Landlock, separate unprivileged uid per VM, and cgroup limits.
So that users have security & knobs right out of the box.
fwiw i used to operate an AWS service using firecracker.
firecracker is a lot more of a headache to setup than docker sbx (not container!) but if you can get it going it probably “feels” the best
of a different variety some people feel better using stuff like bubblewrap/fire jail but idk if these are still microvm as opposed to the above
but it’s my opinion (perhaps completely criticizable), that sandboxing for personal home use machines is 1) somewhat overkill 2) somewhat theatre 3) psychologically sometimes exhausting and unrealistic always assuming the worst is going to happen and 4) not really worth the time and creates quite a bit of friction that there wasn’t previously. In the end I spent more time tinkering with the sandbox to get it “just” right that it drained my time actually using the agents so …
https://github.com/pkulak/opencrow
I spent lots of time on this topic, and had a similar path. Meanwhile the tool also supports a quick way to add custom or local providers to agents: https://vibepod.dev/news/vibepod-cli-0-24/
This looks very interesting, letting people run these in any box they own. I very much agree with the sentiment that there are no proper tools that let you run any coding agent without being locked to a single provider. I just want to run opencode or pi somewhere in a box without having to spin up all of them in my machine, let alone being able to trigger them from within another prouduct as a background agent.
Would pi-pod allow me to standardize pi config on a team/project level so that other people in my org can also use them on the same private infra?
Yep!