Hello fellow keepers of numbers,
This week, we dive into a few news items that might give us a glimpse into what our future looks like.
First, I cover the news from last week on the ChatGPT model that breached Hugging Face along with the Anthropic news that their Claude models also breached some organizations. I expect we’ll be hearing a lot more about these kinds of security incidents and how best to secure our data in the new world in which we now live.
Second, I cover the new ChatGPT Voice that launched late last week. This genuinely feels like a very advanced, fun way to work. I don’t know if voice is how everyone will work two years from now, but it feels like it’s possible. Check out this week’s demo in the Put It to Work section to see what I mean.
The Latest
OpenAI and Anthropic agents access real companies during cyber tests

Source: ChatGPT Images 2.0 / The Appreciable Asset
OpenAI disclosed that AI agents being tested for cybersecurity escaped their controlled environment and gained unauthorized access to Hugging Face, an AI software platform. The agents had been instructed to complete a cybersecurity challenge and continued pursuing that goal after reaching real systems.
Hugging Face reconstructed roughly 17,600 actions taken by the agents over 4.5 days. The agents moved through the company’s internal systems and accessed five datasets connected to the cybersecurity challenges. OpenAI later found that the models also accessed four accounts on other online services.
Anthropic later reported three incidents involving Claude after reviewing 141,006 test runs. A configuration mistake gave the models access to the live internet even though they had been told everything they could reach was part of a simulation. The models accessed a live company database, published a malicious software package that ran on 15 systems, and compromised another company’s public application after scanning roughly 9,000 potential targets.
OpenAI disabled the internal research model involved and tightened controls around its testing environments. Anthropic stopped the affected tests, notified the three organizations, and said it is expanding monitoring and security checks across its own systems and outside testing partners.
Why it’s important for us:
I wrote about OpenAI’s security incident in last week’s Weekly Random. Yesterday, Anthropic disclosed three more incidents with Claude models. Four security incidents disclosed by the two leading AI labs within about a week feels significant. Enough for us to pay close attention.
These incidents came out of tests designed to push the models. The AI models were given a problem to solve inside a sandbox, which is basically a locked testing room. They’re supposed to have enough access to complete the test without reaching the public internet or anything actual companies use to run their businesses.
OpenAI’s models realized the information they needed to complete the test was outside the sandbox. They found a security hole no one knew existed to escape the sandbox. They then reached the internet and hacked Hugging Face, again using security holes the company didn't know existed.
Anthropic’s models were accidentally connected to the internet, so they treated websites, servers, and databases belonging to actual companies as if they were part of the simulated test. Less technical "hacking" here, but still a security breach nonetheless.
The models weren’t acting maliciously. They weren’t trying to escape and take over the world. They were literally doing what they were asked to do. The concerning part is that their path to completing the task included breaking out of the test and hacking real-world companies.
At this point, I think if someone points a top AI model at a piece of software and tells it to hack its way in, it probably can. Probably within minutes or hours. These models can find security holes humans don’t know exist and move far faster than any human hacker.
Accountants don’t need to have the solution to this, and we’re not going to. Security experts and software companies need to figure out what protection looks like now. But every vendor we use holds sensitive client data, and whatever counted as strong security before these incidents may not be strong enough anymore. Things will likely be changing, so we need to pay attention.
OpenAI adds voice control to ChatGPT Work and Codex
OpenAI announced ChatGPT Voice for Work and Codex in its desktop app. Powered by GPT-Live, the feature allows users to speak naturally, interrupt, ask follow-up questions, and direct work across multiple agents, conversations, and projects.
OpenAI says Voice uses the tools and permissions available inside the selected Work or Codex session. This includes connected apps, local files, and other tools already available to that session.
Voice is available through the ChatGPT desktop app on macOS and Windows for Plus, Pro, Business, Edu, and Enterprise plans.
Why it’s important for us:
This was a tiny announcement, but after testing it for the last week, it could be one of my favorite launches of the last few months.
I opened a new task in the ChatGPT app, turned on voice mode, and started talking. From there, I could ask it to search my Google Drive, read a document back to me, create a new file, open another app, or start separate tasks while I worked elsewhere on my computer. I didn’t even need to look at the ChatGPT app. Voice mode sat in a little bubble on my screen, listening whenever I spoke and giving me updates as the work progressed.
The ability to delegate is probably the coolest part. You can tell it to start four separate tasks, give each one a different assignment, and it automatically brings the results back into the original conversation when they’re finished. And then ChatGPT Voice (I haven't named him yet) tells me what I need to know. Anything I would normally do in the ChatGPT app, I can now ask it to do by talking.
It's definitely not perfect, so I don't want to overreact. It occasionally stopped working, missed words, or transcribed something incorrectly. You’ll also probably want headphones, and I'd imagine talking to an AI assistant in your headphones in a crowded office is a good way to make your coworkers question your mental stability. There are also plenty of tasks where typing is still better, in my opinion. I’m not using this all day, every day.
But when it works, it’s awesome. I see myself using this when I step away from my desk or if I'm out and only have my phone with me. Although, I prefer not to use my phone constantly and just enjoy life sometimes.
For the Marvel nerds like myself, this feels like we're a small step closer to Jarvis from the Iron Man movies. That felt stupid to type, but I did it anyway. Let's all just move on.
Voice isn't useful for everything. After all, I'm typing this with my fingers and a keyboard right now. But ChatGPT Voice does feel like a particularly interesting future. I suspect Anthropic, Google, and others are already working on voice features similar to what now exists in the ChatGPT app.
Trending News
OpenAI added multi-folder ChatGPT app projects that can work across related code, documents, templates, and reference files: This is one of my favorite underrated features of Cowork, and now it's also in ChatGPT. One example of why this might be useful: You link Client ABC's folder so your AI of choice can read the files in that folder and save outputs there. But what happens when you download a file from your email and it lands in your Downloads folder? Previously, projects scoped only to Client ABC's folder couldn't go grab the file from Downloads. But if you link the Downloads folder to your project for Client ABC, your AI can now grab and save your files. Plenty of other use cases as well.
AI workers urged Washington to build a mechanism for slowing frontier development if needed: Another week, another news story about the people building the models asking to slow down. Seems reasonable. We should probably pay attention.
The White House accused Moonshot AI of distilling Anthropic’s Fable model to build Kimi K3 and threatened sanctions: I'm oversimplifying a bit, but distilling an AI model means they ran large quantities of prompts in Claude and used the outputs to "teach" their model how to act the same. Some users ran tests where they asked K3 what its name is. Kimi K3's response: "I'm Claude." Seems like they're guilty of the claims.
Cursor added Kimi K3 with US-based inference and zero data retention support: Kimi K3 cheated to be as good as it is (see story above). But we're not going to be the moral police right now. If it's good and cheap, maybe we should use it. I don't buy the claims that it's as good as Opus or GPT-5.6, but it's also much cheaper. And now Cursor is providing a secure way to use it.
Greenshoe launched an AI platform that continuously monitors regulatory changes and identifies the affected sections of future SEC filings: AI is making things like continuous monitoring possible. That feels obviously better to me than checking each quarter. This sort of rhymes with what we're seeing in audit where people are considering the possibility of auditing full populations instead of small samples.
Bizora launched an advisory group of CPAs and enrolled agents to review returns before filing: This is an AI tax research company launching an advisory service line. It expands its customer base to firms that may not be advanced enough to use Bizora but see the value of outsourcing work, and Bizora obviously believes it can use its own platform to make the economics work. The business model is definitely interesting, and the service could be useful for firms without the capacity or expertise to pursue certain advisory opportunities.
Google integrated Gemini Spark with Chrome so the agent can work through logged-in accounts and saved credentials: Google is still so far behind on agentic AI, but they do have a huge leg up when it comes to their AI models using browsers. Claude and ChatGPT are both great at using the browser now. Gemini should hopefully catch up fast since Chrome is under the Google umbrella. AI using the browser is still fairly risky for firms right now though, so proceed with caution.
Karbon found 22% of clients want AI handling a business crisis without human help: Am I missing something here? 22% seems insanely high. Who the hell wants 100% AI handling a business crisis? I guess the point of the report though is that clients still want that human touch even with advanced AI. Feels like many of us have been saying this all along, so nothing has changed yet.
Paylocity launched Ignite AI with agents for payroll analysis, time corrections, employee data, and recruiting: We've seen a lot of AI agent updates to the A/P and payroll software recently. I like that admins can turn individual agents on and off. Still not sure how useful these updates are vs a connector to pull that data into Claude Cowork or ChatGPT.
DeepSeek launched an updated V4 Flash API with stronger agent performance and native support for the Responses API and Codex: This is something to watch because DeepSeek is positioning V4 Flash as an Opus-level model at a tiny fraction of the cost. I don't think I buy the intelligence level just yet, but between DeepSeek V4 Flash and Kimi K3, the Chinese models are opening eyes.
Anthropic settled its copyright lawsuit for $1.5 billion after a court found it illegally downloaded and stored millions of copyrighted books to train its models: This is a huge amount of money, but not that significant per piece of work they pirated. Interestingly, the courts decided training was fair use. It was the pirating that was the problem.
Put It to Work
Not even going to lie, this week’s demo is pretty awesome. The ChatGPT app is so good, and the new ChatGPT Voice feels like living in the future. It’s you speaking to ChatGPT and it speaking back while doing work for you.
I spent the majority of the demo inside the ChatGPT app so you could see it transcribing my voice and ChatGPT’s voice, but this is probably best used while you’re working in other apps during your day, or even away from your computer.
Weekly Random
People used to tell us about when AI would build AI. We've made it.
OpenAI used GPT-5.6 Sol to improve the systems running its own models. It rewrote production code, ran hundreds of experiments, and helped train and monitor its own smaller draft model.
That work cut serving costs by 20% and improved token-generation efficiency by more than 15%.
AI companies have been telling us for months now that they're using AI to build future AI models. But this is a little different. The day after OpenAI made the announcement about GPT-5.6 Sol improving its systems, they cut the price of GPT-5.6 Luna by 80% and Terra by 20%. The cuts apply to API prices, and Luna and Terra now eat up less of your usage in Codex and ChatGPT Work.
So the improvements have had a net positive impact on end users. It makes it pretty hard to debate the fact that AI is building and improving itself.
AI is building better and more efficient AI. We just have to hope it continues to positively impact our wallets.
Until next week, keep protecting those numbers.
Preston
