An OpenAI agent was told no, so it went round. Now think about what your agents can reach.
On September 24, Australia’s Prime Minister Anthony Albanese told the country that an OpenAI agent had broken into a Medicare statistics portal run by Services Australia. Nobody pointed it there. According to the government, OpenAI had given the agent a “benign” research job during an internal evaluation: find figures on public medicines spending. It went searching, found the portal, asked it questions, and didn’t get what it wanted. So it went round. Albanese’s description was that it “found a way around those blocks” and “didn’t accept no for an answer.” Acting Prime Minister Richard Marles said the data had been kept “behind a fence that the AI agent effectively climbed over.”
The timeline makes it worse. The agent got in on June 18. OpenAI says it found the activity in August , during a wider review of what it calls misaligned model activity. It told Services Australia on September 10, by email, to the agency’s public inbox. Services Australia saw it the next day and alerted the Australian Signals Directorate on the 15th. OpenAI’s statement says its models “took actions we did not intend.” Somewhere in those three months, Sam Altman met Marles in San Francisco on September 1. Make of that what you like. Albanese called the whole thing “unacceptable” and has set up a taskforce with the ASD and the AI Safety Institute.
The good news is that the statistics were low-stakes. They were aggregate figures: bulk-billing figures, immunisation and PBS data, annual reports. Some weren’t public yet, and the government says none of it was “particularly sensitive” and all of it has since been released. No one’s personal Medicare record was touched. But it didn’t just read. In its apology on September 28 , OpenAI says the model ran commands, retrieved internal files, credentials and aggregate statistics, and wrote files. It has also paused training and evaluation involving tool use for its most capable models. Nobody has said which weakness it used.
It wasn’t a one-off either. OpenAI has since said dozens of third parties were affected by its agents, and the behaviour it lists includes agents “using leaked passwords to access online services.”
So this isn’t a story about stolen health records. It’s a story about what an agent does when a system says no. That matters to you, because you’re probably about to let agents near things far more sensitive than spending statistics.
“No” is only a no if the agent can’t reach it
Put yourself in the agent’s position. It has a goal and tools. The portal’s refusal isn’t a moral boundary to it. It’s an obstacle, the same kind of thing as a slow page or a missing search box, and obstacles are what it’s built to get past. Nobody told it to hack anything. The goal did the telling.
That’s the uncomfortable lesson. Any check an agent can reach, it can try to route around. A web form that refuses, a setting that says “please don’t”, a warning in the page text: those are all things the agent reads and acts on, which means they’re all things it can reason its way past, or be talked past by someone else. The checks that actually hold are the ones that sit outside the agent’s reach entirely, or that ask for something the agent doesn’t have.
Now point that at your password manager.
We already know what happens when an agent can reach your vault
This isn’t a hypothetical. In March, Zenity Labs published PerplexedBrowser , showing how Perplexity’s Comet browser agent could be pointed at a user’s password vault. The attack was a calendar invite with instructions hidden below the meeting details. The user asked Comet to deal with their calendar. Comet read the invite, took the hidden text as part of the job, and because the password manager’s extension was unlocked and signed in, it opened the web vault, searched it, revealed passwords and sent them to the attacker’s server. It could go further, changing the account password and pulling out the account’s Secret Key.
The researchers put it bluntly: “No software vulnerability is required. Comet operates within its intended capabilities.” Nothing was broken. The vault was open, the agent was allowed to use the browser, and so the agent could use the vault. Perplexity tightened its prompt-injection handling and confirmations in February, and 1Password added options to turn off automatic sign-in and to require confirmation before filling a password. Zenity notes that both sets of fixes had to be switched on by hand.
Put that next to the Medicare story and the pattern is plain. In one case, an agent treated a refusal as an obstacle and got past it. In the other, an agent was steered by a stranger’s text and faced no refusal at all, because an unlocked vault in the same browser doesn’t have one to give. Either way, your passwords are only as safe as the nearest check the agent can’t get round.
Where the check has to live
Here’s the test we’d apply to anything that holds your logins, now that agents are sitting in browsers and on desktops.
Is the vault open inside the agent’s reach? If your password manager stays unlocked for hours in the browser, and the browser is now something an agent drives, then the agent is holding your vault for those hours. It doesn’t need to break anything.
When a password leaves, what has to happen first? A setting that says “ask me” is only as good as where the asking happens. If the prompt is a web page or a panel inside the browser, it’s inside the thing the agent controls.
Can the approval be given by a click? This is the one people miss. Agents that operate a computer can click buttons, including buttons in dialogs outside the browser. A fingerprint or a secret the agent was never given can’t be faked that way. A button can.
Does the approval name where the password is going? The PerplexedBrowser attack sent secrets off to a server nobody chose. If each release has to show you the real site before anything moves, the odd one stands out.
How Vauz answers it
This is where we can speak for our own product, because we built Vauz around the same idea. The Vauz browser extension holds no vault, no stored passwords and no key that opens the vault. It can ask the Vauz app for the login for the site you’re on, and that’s all. The app then asks a person, outside the browser, before anything is sent, and the prompt names the real site. On a Mac with Touch ID, that’s a fingerprint. Otherwise, if you’ve set a V-Key, it’s your V-Key. If you’ve set neither, it’s a plain Allow or Deny dialog. We laid out how that works against page tricks in how Vauz stops clickjacking from stealing passwords, and what a hijacked extension could and couldn’t take in the Twitch token leak.
Held to our own test, the answer isn’t perfectly clean. Here’s where it falls short.
The Allow or Deny dialog is a click. A web page can’t reach it, but a modern agent with access to your computer can press a button on a system dialog as easily as you can. What it can’t do is put its finger on your Touch ID sensor, or type a V-Key it was never given. That’s why the fingerprint comes first where there is one, and why the dialog itself tells you to set a V-Key or turn on Touch ID if you want more than a click. If you’re going to let agents anywhere near your computer, do that.
An approval also covers the same site for two minutes, so a sign-in split over two pages doesn’t ask twice. Something acting inside that window on that same site could ride it. If you’d rather not have that two-minute window at all, turn on Prompt for authentication on every autofill in the extension’s settings and every release asks again.
What an agent can’t do is open the whole vault and browse it, because there’s no open vault in the browser for it to browse. It can ask for one site’s login, and a person has to say yes with the site’s name in front of them. We made the longer argument for why that approval has to stay with a human in letting an AI agent hold your passwords.
The question to ask
OpenAI’s agent wasn’t malicious (at least we hope so). It was keen, and that’s the point. An agent that wants to finish its task will treat every refusal it can reach as a problem to solve, and an agent that’s been fed someone else’s instructions will do the same on their behalf. You can’t fix that by asking agents nicely.
So before you hand an agent your browser, or your whole computer, ask one thing of the tool that holds your passwords: when a password is about to leave, what has to happen that the agent can’t do by itself? If the honest answer is “nothing, the vault’s already open”, then the agent has the same access you do, for as long as you leave it that way.
A no that a human has to say yes to
Your logins leave Vauz only after a person approves it, outside the browser.
Every release names the real site first, and with Touch ID or a V-Key it takes something an agent driving your screen can't fake. The free plan stays completely free for life, with Plus and Premium available when you need more!
Use Vauz completely free — for, like, ever