I'm no evangelist for LLM assistants, but this seems incredibly improbable and represents a failure of MacOS security if so. If full disk access isn't granted, Mac blocks it from the Downloads folder, to say nothing of actually sensitive paths. I would expect a far more likely case of an accidentally granted permission on another device or a permission that was on and then turned off.
Permissionless action is about to skyrocket as an issue, but this particular scenario strikes me as incredibly unlikely. Would be interested to know if Muse can provide more meaningful data provenance/logs.
Scanning iMessage dbs as a passive part of full disk access (and not a messages grant), if true, is a little sketchy, regardless.
Agent sandboxing/access control is one of the biggest problems to be solved before this technology really should go mainstream.
Even as a technical person, it's not trivial to sandbox agents correctly. The fact that an mis-clicked permission popup could give an agent unrestricted access to a user's disk is a massive risk vector in the hands of lay people who barely understand how any of this works.
So much of current security depends on the model of tying access control to a user account. A lot has to be re-thought in terms of how to grant access to an agent working on the user's behalf, in a way that doesn't make it completely useless, and also doesn't require every user to become a sysadmin managing fine-grained agent permissions manually.
I think sandboxing could be solved if effort was put into it. Webassembly seems like a great way to enforce data and execution boundaries for an LLM, for example.
I think the problem is that LLM providers are dis-incentivized from pursuing it because their ethos is gobbling up any and all data they can get.
> Oops, we accidentally yoinked your personal documents, photos, and videos and they’re now swimming in our model’s data ocean! We’re sorrrry, oh well let’s move on.
It’s up to the users to use tools that enforce security/privacy. Open source harnesses like pi.dev seem like a good path forward to me
Sandboxing the agent application is easy. The tough part is sandboxing in such a way that it's still useful.
I.e. if I have an agent running in a WASM sandbox with no access to the host system, I can't ask it to clean up my files. Same thing with things like giving an agent access to your email inbox: doing so allows the agent to provide utility, but it comes with risks, as the agent can delete important emails, or leak sensitive data.
I think a big part of the problem is, a lot of the systems we use and would like agents to help us with don't have a concept of separated roles with different levels of access which can be applied. A lot of times it's all or nothing.
And even when we do have fine-grained access control available, it's a pain in the ass to manage it. Like you can create a GitHub token with fine-grained access control to your repositories and make sure the agent only uses that one to connect, but it's a whole lot easier to use a broad-access token, or just let the agent use your own token, so lots of people will just end up doing that.
And I also like pi, but it's probably one of the worst in terms of sandboxing as it's yolo by default.
Fundamentally i think its how we're trying to skip critically important steps here. Its like the first automobiles that were built with completely uncovered, unfiltered engines. Dropped onto a chassis, and then opened up, dirt, debris, oils, etc intermingled and things blew up. Gas lines, oil reservoirs, compartments and chambers that are all specialized to create a highly efficient and consistent experience came out of designing it properly. The same can be said for Muse and other agents. Without the right scaffolding and harness, of course things go wrong.
I'm a HUGE proponent of putting them behind task gating trees, and sheathing them with QA/QC checks on their processes, especially at this stage. Unfortunately its so easy to create, and all of that takes time and design that many just throw away for getting to results.
Even fine grained ACLs, which are great, dont have the structured approach such autonomous agents need to shore them in (imo).
It's already solved. I have two git repos proving these companies can fix the problems. The fact this continues just proves they don't care. In one project I literally containerize CLI coding tools, it works. You might say "sure, but network." I literally wrote a desktop app harness that you can toggle the network on/off too.
I sandbox my coding agents using bubblewrap, but I don't think it's as trivial a problem as you make it sound.
For something like muse that's supposed to be a general-purpose assistant, how do you give it enough access to be useful, without giving it too much access, and creating unacceptable risks? And how do you do that in a way that's comprehensible the average Facebook user who's the target market of this product?
Nothing you said was wrong, but I can’t imagine giving Meta the benefit of the doubt on, well, anything. Fool me once, shame on you. Fool me 137 times…
And we're surprised by this?? Its Meta after all. The social graph data we had access to in the social games we built back in 2009/2010 was literally insane by todays standards of security. They've just gotten better at hiding it.
I think the fundamental problem is that AI companies have been assuming that reinforcement learning with human feedback is an adequate foundational technology for guardrails. And that simply isn’t true.
I think what’s more alarming is the macOS nannying UAC-like toggles to block disk access and other “protections” are apparently all UX reducing flash and no actual functionality.
I’d argue this is a five alarm fire for macOS and Meta simply exploited it.
Personally that stuff drives me up the wall, it's the Mac wanting to become the iPhone and close off everything but the App Economy. They'll geofence XCode to the Bay Area or something so only "professionals" can develop software and eventually ban web browsers.
At least on Unix-like systems, if it's free for you to read, then it's free for any process you run as you to read. Sure, macOS has grafted its own weird "permissions" layer on top of the existing OS level permissions, but at the end of the day, when you run an app on Unix, you're allowing it to act as you, with all the powers your user has.
This used to work when you could trust the software you ran on your system to have access to everything you have access to on your computer. I'd argue that time has largely passed, for most third-party commercial developers and even for some OS vendors.
Best solution is to simply not run software made by blatantly untrustworthy developers. Second best solution would be to run such software as a severely sandboxed user who basically doesn't have access to anything important on your system.
My mental model is to treat AI agents running on your system as a form of malware that is running in a honeypot you control. You don't want to just get rid of it as you want to observe its behaviour (in the case of malware) or hopefully do something useful (AI). But you certainly shouldn't assume it won't do anything bad to your system.
> at the end of the day, when you run an app on Unix, you're allowing it to act as you with all the powers your user has.
This is not at all how it works on macOS, which is what is being discussed in the original post. There are a million different things that require per-app explicit opt-in permissions. This is a case of user error.
I do not know about all Unix like OSes, but Linux has sandboxes you can run as you main user. Not as safe as running as a separate user and sandboxing, or running in a VM, but reasonably solid.
> I'd argue that time has largely passed, for most third-party commercial developers and even for some OS vendors.
In theory, iOS has some protection: the security model is that not all apps can be trusted with everything, so they have to ask for permission to access camera, GPS, texts and so on. I think apple even kicks apps out of its store that try and abuse this.
> Meta says that Muse has to obey permissions that users set up. It won't access any data you don't explicitly allow it to access.
Is this a setting configured in Muse itself?
> It took Meta a single day to begin "helpfully" pitching article ideas based on texts he'd sent to a podcast co-host. When he asked Muse how it got the information, it said that it read banners from incoming texts. But that's not true, either.bAfter doing a little digging, Aten says Muse synced 187,000 lines from his Messages database, despite Full Disk Access being off.
Is full disk access enforced on the OS side, or the app side? Like is this claiming MacOS security was breached by Muse somehow acting in spite of deliberately disabled access somehow?
Permissionless action is about to skyrocket as an issue, but this particular scenario strikes me as incredibly unlikely. Would be interested to know if Muse can provide more meaningful data provenance/logs.
Scanning iMessage dbs as a passive part of full disk access (and not a messages grant), if true, is a little sketchy, regardless.
Even as a technical person, it's not trivial to sandbox agents correctly. The fact that an mis-clicked permission popup could give an agent unrestricted access to a user's disk is a massive risk vector in the hands of lay people who barely understand how any of this works.
So much of current security depends on the model of tying access control to a user account. A lot has to be re-thought in terms of how to grant access to an agent working on the user's behalf, in a way that doesn't make it completely useless, and also doesn't require every user to become a sysadmin managing fine-grained agent permissions manually.
I think the problem is that LLM providers are dis-incentivized from pursuing it because their ethos is gobbling up any and all data they can get.
> Oops, we accidentally yoinked your personal documents, photos, and videos and they’re now swimming in our model’s data ocean! We’re sorrrry, oh well let’s move on.
It’s up to the users to use tools that enforce security/privacy. Open source harnesses like pi.dev seem like a good path forward to me
I.e. if I have an agent running in a WASM sandbox with no access to the host system, I can't ask it to clean up my files. Same thing with things like giving an agent access to your email inbox: doing so allows the agent to provide utility, but it comes with risks, as the agent can delete important emails, or leak sensitive data.
I think a big part of the problem is, a lot of the systems we use and would like agents to help us with don't have a concept of separated roles with different levels of access which can be applied. A lot of times it's all or nothing.
And even when we do have fine-grained access control available, it's a pain in the ass to manage it. Like you can create a GitHub token with fine-grained access control to your repositories and make sure the agent only uses that one to connect, but it's a whole lot easier to use a broad-access token, or just let the agent use your own token, so lots of people will just end up doing that.
And I also like pi, but it's probably one of the worst in terms of sandboxing as it's yolo by default.
I'm a HUGE proponent of putting them behind task gating trees, and sheathing them with QA/QC checks on their processes, especially at this stage. Unfortunately its so easy to create, and all of that takes time and design that many just throw away for getting to results.
Even fine grained ACLs, which are great, dont have the structured approach such autonomous agents need to shore them in (imo).
This is all amateur hour shenanigans.
For something like muse that's supposed to be a general-purpose assistant, how do you give it enough access to be useful, without giving it too much access, and creating unacceptable risks? And how do you do that in a way that's comprehensible the average Facebook user who's the target market of this product?
https://hntrbrk.com/breaking-news/muse-doxxing
I’d argue this is a five alarm fire for macOS and Meta simply exploited it.
This used to work when you could trust the software you ran on your system to have access to everything you have access to on your computer. I'd argue that time has largely passed, for most third-party commercial developers and even for some OS vendors.
Best solution is to simply not run software made by blatantly untrustworthy developers. Second best solution would be to run such software as a severely sandboxed user who basically doesn't have access to anything important on your system.
This is not at all how it works on macOS, which is what is being discussed in the original post. There are a million different things that require per-app explicit opt-in permissions. This is a case of user error.
> I'd argue that time has largely passed, for most third-party commercial developers and even for some OS vendors.
Agreed, but what can you do about your OS vendor?
Muse is not available in the macOS app store. Almost certainly because of the sandboxing requirements.
Is this a setting configured in Muse itself?
> It took Meta a single day to begin "helpfully" pitching article ideas based on texts he'd sent to a podcast co-host. When he asked Muse how it got the information, it said that it read banners from incoming texts. But that's not true, either.bAfter doing a little digging, Aten says Muse synced 187,000 lines from his Messages database, despite Full Disk Access being off.
Is full disk access enforced on the OS side, or the app side? Like is this claiming MacOS security was breached by Muse somehow acting in spite of deliberately disabled access somehow?
Has this been reproduced / recorded?
The OS side.
> Has this been reproduced / recorded?
No.