Rendered at 12:18:33 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
epage 9 minutes ago [-]
Been looking at sandboxing, both low level and higher level like this.
The API for their Rust mxc-sdk looks nice but
- their "sdk" has binaries and the build script has logic for them
- their build scripts do windows-exclusive work on all platforms
- not putting some of the backends behind features causes more build script work (and that work will break on future Cargo versions)
- at least some of the remaíning build script work doesn't need to be a build script
- it seems pretty dependency heavy
neobrain 1 hours ago [-]
Do any of these sandboxing solutions have a dynamic component to them that lets you grant permissions, starting with a minimal sandbox and asynchronously adding permissions as they become necessary? Harnesses try to do this when accessing non-project folders, but it's not always strictly enforced and generally not revocable. Harnesses also block agent execution until a decision is made, which requires constant monitoring to ensure progress can happen when the agent could easily proceed with an alternative method right away.
I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).
Does anything close to this exist yet?
0kk33 29 minutes ago [-]
If I understand you correctly https://nono.sh/ might go into that direction. It can add permission after the sandboxed command is terminated based on which blocks occurred.
Its not life though as you seem to describe
ghm2180 18 minutes ago [-]
I have faced this exact dilemma as well. It's not always clear what permissions are needed in advance for my pi sessions and it's child sub sessions. A simple example is when child sessions do a task they locally want to fire random docker commands to learn the state of my local docker devstack.
dannyw 3 hours ago [-]
This looks pretty decent actually. Sure, you could consider it a frontend/SDK for bubblewrap/seatbelt/processcontainer; but setting em up consistently is far from trivial; and hand rolling is a really bad idea (speaking from experience).
I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.
Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.
LeBit 2 hours ago [-]
It does look nice and easy.
Not sure I will give up on smolvm though.
I have scripts to launch an instance per task. Nono wraps my coding agent.
I provide the git clone for the specific task.
It works really well.
kernc 2 hours ago [-]
350,000 of mostly Rust SLOC [1] ... And the upstream sandboxes aren't even vendored!
I'd be way more confident building upon something I can grasp and understand. [2]
That site thinks this file has 2.9k sloc and doesn't seem to parse rust comments. In reality, there's only 1,465 sloc; and 635 loc of tests.
Definitely nowhere near 350k sloc.
--
As for your sandbox run: it's a single-contributor project, seems to have only have basic smoke tests, and has a few major/critical security issues:
* _generate_seccomp_filter compares newline-deliminated syscalls, against a multi-line blocklist, meaning the entire function doesn't block anything and is essentially a no-op.
* Main script invokes working directory's .env as shellcode, before switching into restricted filesystems and dropping capabilities. Attacker-controlled .env can run shellcode with full privileges.
* Lots of race conditions which I haven't verified, but doesn't really matter.
I'd make PRs, but I don't think it's a good idea to try and DIY a sandboxing system in bash with minimal SLOC as the target in the first place. I'm also slightly concerned that most of your comments on HN seem to be promoting this repo?
IshKebab 2 hours ago [-]
And it's 600 lines of dense Bash. I trust 350k lines of Rust way more than that!
its-summertime 1 hours ago [-]
I feel a better metric is `lines changed / time` as that affects what will be audited as time goes on
That being said, a month of MXC has more line changes than 2-3 years of runc
mintflow 48 minutes ago [-]
Seems aws also announced a sandbox solution
I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far
Why we need this layer of complexity? Or its mainly for big company that need control ?
joshuanapoli 42 minutes ago [-]
If you have a custom agent in a product, then it needs isolation to be sure to protect the customer data.
stingraycharles 14 minutes ago [-]
People are running custom agents that are able to run custom code in production just like that ?
Why is everyone making their own code execution agent runtime engines I have an entire project built on top of openshell already, why not first come up with a sandboxing policy design, like unix did, and then build on top of that.
Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.
I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.
I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.
lifeisloving 2 hours ago [-]
I saw a tweet that said:
"Im really fkin worried we're all building the same thing"
Everyone has been building a harness/sandbox the last 6 months. Ive seen dozens and dozens shared in discords.
Even companies are totally stuck focused on the same paradigms.
The previous iteration of this was RAG/Chat interfaces. See PewDiePie's project. Last month it was briefly everyone building the same classifier.
Peter Thiel, gave a lecture about this same phenomenon 15 years ago likely because he observed the same things going on during other hype cycles. Everyone building the same things. Its called something like "Dont build the obvious thing"
This is why Im moving towards hardware for personal projects, it forces me to be much more creative and think outside the "How can I make something AI adjacent/powered" trap thats so easy to fall into in pure software right now.
fg137 1 hours ago [-]
My guess is that Microsoft thinks this can be deployed with standard, company-wide policy across platforms (mostly) with their IT management tools which poses a unique advantage.
In reality, however, knowing how much difference there is between OSes, how tricky it is to configure these things to make them actually useful, and how bad Microsoft products are, I'm not enthusiastic about this project -- there are so many others on the market already, and I'll wait to see if this gains traction.
(Note that on MacOS it only supports seatbelt? That's not nearly the same as microvm.)
torginus 3 hours ago [-]
The problem with OS level sandboxes, and the reason why WebAssembly's being explored in this space (and Electron is so popular), is that relying on OS/hardware features means your TAM shrinks to a fraction of total, and it's historically well known you set yourself up to lose.
History is littered with tons of super cool OS features that didn't manage to gather enough market share and ended up as cool futures, and fodder for 'we invented the future 20 years ago' style articles.
hobofan 3 hours ago [-]
Different use-cases have different requirements.
e.g. this one puts multi-platform support as a high requirement, a requirement that OpenShell doesn't fulfil (and likely won't given it's architecture/goals).
rock_artist 3 hours ago [-]
That’s exactly it.
There should be some permission logic for delegating.
But as always, there are rivals trying to set their tone on what’s the standard. We all wish there was one unified agreed concept that will work but I guess the most common one will eventually survive.
Just as Microsoft in a sense embraces Linux with WSL and also Apple has their virtualization framework.
I hope we’ll eventually get unified model management system to include also permissions designed properly
booster-rooster 3 hours ago [-]
[flagged]
chneu 4 hours ago [-]
Right you are, Ken!
nizbit 3 hours ago [-]
There it is! :)
arj 2 hours ago [-]
Would this allow a sandboxed container on windows to still run commands in wsl?
rfgplk 1 hours ago [-]
They have a sandbox escape in there. Likewise capability ordering is wrong. Exactly what you should expect from Microsoft.
Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.
zbentley 15 minutes ago [-]
Could you link to issue reports to back that up (or maybe file them if these are novel findings)?
If that’s too much of an ask, at least reference the code you found for things like “mxc sandbox escape” or “bubblewrap setuid is wrong”. Those claims require evidence.
9 minutes ago [-]
singularityisne 16 minutes ago [-]
[flagged]
t_privos 59 minutes ago [-]
[flagged]
smitty1e 3 hours ago [-]
Asked Grok the difference between mxc and flatpak:
"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."
zenapollo 1 hours ago [-]
Saw the M stands for Microsoft and immediately closed the tab. 1 it’s unnecessary - communicates nothing but look-at-me branding. 2 toxic company.
The API for their Rust mxc-sdk looks nice but
- their "sdk" has binaries and the build script has logic for them
- their build scripts do windows-exclusive work on all platforms
- not putting some of the backends behind features causes more build script work (and that work will break on future Cargo versions)
- at least some of the remaíning build script work doesn't need to be a build script
- it seems pretty dependency heavy
I like the idea of a minimal sandbox that protects against accidental `rm -rf` and against personal data leakage, but such a setup then often gets in the way of the specific task to be done. Ideally the sandbox would be able to aggregate blocked accesses and then expose them in an external TUI dashboard, where I can then enable access (without blocking any running agent on this, since that's prone to "press okay" fatigue).
Does anything close to this exist yet?
I like the ‘learning’ mode for figuring out what perms/config a runtime needs, the MIT license, the clear optional telemetry disclosures, and somewhat light and still readable documentation.
Regardless of your views on Microsoft, this looks quite useful; serves a clear purpose, and from a quick glance, looks like a high quality project even if it’s just the first version.
Not sure I will give up on smolvm though.
I have scripts to launch an instance per task. Nono wraps my coding agent.
I provide the git clone for the specific task.
It works really well.
I'd be way more confident building upon something I can grasp and understand. [2]
[1]: https://ghloc.dev/microsoft/mxc [2]: https://github.com/sandbox-utils/sandbox-run
https://ghloc.dev/microsoft/mxc?branch=main&locsPath=%5B%22s...
That site thinks this file has 2.9k sloc and doesn't seem to parse rust comments. In reality, there's only 1,465 sloc; and 635 loc of tests.
Definitely nowhere near 350k sloc.
--
As for your sandbox run: it's a single-contributor project, seems to have only have basic smoke tests, and has a few major/critical security issues:
* _generate_seccomp_filter compares newline-deliminated syscalls, against a multi-line blocklist, meaning the entire function doesn't block anything and is essentially a no-op.
* Main script invokes working directory's .env as shellcode, before switching into restricted filesystems and dropping capabilities. Attacker-controlled .env can run shellcode with full privileges.
* Lots of race conditions which I haven't verified, but doesn't really matter.
I'd make PRs, but I don't think it's a good idea to try and DIY a sandboxing system in bash with minimal SLOC as the target in the first place. I'm also slightly concerned that most of your comments on HN seem to be promoting this repo?
That being said, a month of MXC has more line changes than 2-3 years of runc
I used agent over 1 year and basically always give codex full permission on each thread, do not get issue so far
Why we need this layer of complexity? Or its mainly for big company that need control ?
I’d personally opt for SELinux in such cases
Currently all project do tend to agree on what and how they work but certain things being different makes porting tedius, if all of them have a bare minimum subset common amongst them it would be much easier to switch, and validate security surface area.
I feel like there are more vulnerabilities in this vibe coded slop sandboxes, and it's more likely everyone one of us trusting them to build projects around them will shoot our foot off once a cve is hit in one that's common in all of them but since they are all slop copies someone will have to figure out how they apply to all others and then manually fix it properly, and if one of them makes a CVE public it will leave dozens of these runtimes open to exploits.
I wish the best to my future self with regards to security I feel like we are completely screwed. Since we can no longer depend on upstream for security.
"Im really fkin worried we're all building the same thing"
Everyone has been building a harness/sandbox the last 6 months. Ive seen dozens and dozens shared in discords.
Even companies are totally stuck focused on the same paradigms.
The previous iteration of this was RAG/Chat interfaces. See PewDiePie's project. Last month it was briefly everyone building the same classifier.
Peter Thiel, gave a lecture about this same phenomenon 15 years ago likely because he observed the same things going on during other hype cycles. Everyone building the same things. Its called something like "Dont build the obvious thing"
This is why Im moving towards hardware for personal projects, it forces me to be much more creative and think outside the "How can I make something AI adjacent/powered" trap thats so easy to fall into in pure software right now.
In reality, however, knowing how much difference there is between OSes, how tricky it is to configure these things to make them actually useful, and how bad Microsoft products are, I'm not enthusiastic about this project -- there are so many others on the market already, and I'll wait to see if this gains traction.
(Note that on MacOS it only supports seatbelt? That's not nearly the same as microvm.)
History is littered with tons of super cool OS features that didn't manage to gather enough market share and ended up as cool futures, and fodder for 'we invented the future 20 years ago' style articles.
e.g. this one puts multi-platform support as a high requirement, a requirement that OpenShell doesn't fulfil (and likely won't given it's architecture/goals).
But as always, there are rivals trying to set their tone on what’s the standard. We all wish there was one unified agreed concept that will work but I guess the most common one will eventually survive.
Just as Microsoft in a sense embraces Linux with WSL and also Apple has their virtualization framework.
I hope we’ll eventually get unified model management system to include also permissions designed properly
Since this apparently wraps bubblewrap (another incompetent action on behalf of Microslop), did a quick sweep of that codebase too. Setuid is wrong, capability dropping is wrong, bubblewrap does _not_ protect against compromised/vulnerable kernels (and it should fyi), wrote up a full bubblewrap/mxc sandbox escape too.
If that’s too much of an ask, at least reference the code you found for things like “mxc sandbox escape” or “bubblewrap setuid is wrong”. Those claims require evidence.
"So MXC is a cross-platform “what may this workload touch?” layer aimed at agents. Flatpak is a Linux app format whose sandbox happens to share a backend with MXC on Linux."