Hello fellow keepers of numbers,

Elon managed to pay his way into the AI race. I guess that’s not very shocking when you’re worth a trillion dollars.

The Cursor team put Elon and SpaceXAI on their backs and trudged their way near the top of the AI standings. The new Grok 4.6 model is really good and really cheap. Cheaper than Sonnet and nearly as good as GPT-5.6 and Fable. They also released Grok Bot, which is a new app that makes it easy to build agents that help with business and personal tasks. Meanwhile, we got one of the worst announcements in the history of accounting.

Plus, stick around for a demo testing the new Grok 4.6 model on accounting tasks. I included a quick overview of Cursor as well.

The Latest

Billow launches an accounting firm powered by AI agents

Source: ChatGPT Images 2.0 / The Appreciable Asset

Billow launched an AI-native accounting firm that uses agents to perform accounting and FP&A work inside a company’s existing finance stack. The company describes its service as a replacement for human accounting work rather than software for an accounting team to operate.

Billow says its agents handle processes including month-end close, reconciliations, accruals, budget-to-actual analysis, and financial reporting. The agents build workflows from steps the accounting team already repeats, execute the recurring work, and return exceptions for human review.

The service works inside NetSuite, QuickBooks, and more than 80 other systems without requiring customers to migrate their data. Billow says it moved from an idea to working with Series B through D companies representing more than $1 billion in enterprise value during its first three months.

Billow lists SOC 2 Type II and SOC 1 reports, SOX-aligned controls, approval trails, and a policy against training models on customer data. Access is currently available through a demo request, and the company has not published pricing.

Why it’s important for us:

Please excuse my language. But I fucking hate that companies feel like they have to make ridiculous claims like this just to get into the news cycle. The algorithms favor this type of stuff, and it just sucks.

That being said, making the claim you’re “replacing Big Four accounting firms” is just ridiculous. The title of the launch post on Y Combinator is “We’re killing Deloitte with AI.”

I watched the launch video and read through their post. There’s quite literally nothing there to support that claim. There’s no meaningful demonstration of the product, no explanation of how the agents are doing this work, and no real proof that any of it is useful. Their website has some anonymous results and sample workflows, but we’re mostly being asked to trust that “AI agents” are doing accounting work.

Normally, it’s probably a bad strategy for me, as someone who writes about a lot of companies in this industry, to completely trash one of them. But this is genuinely one of my least favorite announcements I’ve ever seen. Beyond making a ridiculous claim with almost no evidence, Billow has chosen to introduce itself by alienating the entire accounting industry.

I work with firms to build agents, skills, and workflows inside Claude, ChatGPT, and other tools. The idea of using agents to perform accounting work isn’t novel. And when you actually start building these things, you find out pretty quickly that it’s not as straightforward as saying, “Let’s just connect all of these systems together.”

Even for enterprises, the software they use for accounting and finance may not be very friendly to integration. Then you have permissions, approvals, controls, exceptions, bad data, and dozens of other problems that appear inside real workflows. Maybe Billow has found workarounds for all of this. Maybe that’s the valuable part of what they’re building. But they've offered no proof of it.

For what it’s worth, I actually think what Billow is trying to do could be very interesting for enterprises. There are also plenty of accounting firms, including the Big Four, already building and thinking about this exact type of work.

So claiming Billow has some special advantage that will allow it to kill firms already working on the same problems is asinine. It instantly destroys their credibility, and I can’t imagine alienating the rest of the accounting industry will do them much good.

Maybe I’m wrong. Maybe Billow eventually builds something great and develops a huge book of business. But until I see it working for real, I don’t trust any of these claims.

Right now, I think the entire announcement is complete trash.

SpaceXAI launches Grok 4.6 and Grok Bot

SpaceXAI released Grok 4.6, a new model focused on long-running agent tasks, and launched Grok Bot, a team of always-on agents with their own cloud computer, in consecutive announcements this week.

SpaceXAI says Grok 4.6 is designed to sustain work across research, knowledge work, coding, and visual projects over many steps. In the company’s published benchmark suite, it performed in roughly the same range as GPT-5.6 Sol and Fable 5 across knowledge-work and agentic-coding tests.

Grok Bot gives each agent a cloud computer it can use to complete work directly inside the same software and websites a person would. SpaceXAI says its internal teams use Bots to update CRM records, process invoices received in Gmail, and coordinate multi-step work between specialized Bots. Bots can continue working after the user’s computer is closed, and users can demonstrate a workflow once and save it as a routine for the Bot to repeat.

Grok 4.6 is available through Cursor, Grok Build, the SpaceXAI API, OpenRouter, Vercel, and Cloudflare, with pricing starting at $2 per million input tokens and $6 per million output tokens. Grok Bot is available in early beta on desktop and iOS for SuperGrok Heavy, Cursor Ultra, and Cursor Teams Premium subscribers, while enterprise access remains waitlisted.

Why it’s important for us:

We have officially entered the atmosphere. Dumb joke. I can still delete it. I'm writing this currently, I can choose to backspace. Nope? Okay, well I guess it's staying. Sorry.

SpaceXAI has become a serious player.

The Cursor deal may turn out to be one of Elon Musk's best business decisions. In just a few months, the Grok models have gone from being a bit of a joke to contending with the best models in the world.

Grok 4.6 looks awesome. I've said constantly how little weight I place on benchmarks. But without having tested Grok 4.6 much myself (beyond what is in the Put It to Work section below), I have to rely a little on them and on what I'm seeing others say. It's scoring close to GPT-5.6 Sol and Fable 5. The two best models in the world.

It's also really cheap. Its output tokens cost half as much as Sonnet's. Sonnet... It's competing with Fable 5 for intelligence, and it's cheaper than the Claude model two rungs down from Fable. If real-world performance holds up, the combination of capability and cost makes this one of the most interesting model releases we've had in a while.

Grok Bot is harder to judge. Weirdly, it doesn't say which AI model it uses. But surely it's Grok 4.6, right?

The idea of agents that are always on, always accessible, and running on their own cloud computers is really interesting. It's where we've been trending for a lot of the last year. But ChatGPT and Claude Cowork can now run cloud tasks connected to your plugins/connectors too. I'm not convinced Grok Bot's current interface is its final form.

My guess (and I have no supporting evidence) is that this is a preview of where the Cursor app is heading. I can see these capabilities ultimately becoming part of a broader Cursor app that feels closer to ChatGPT's app, but makes it easy to use and gives you access to any number of AI models.

I love all the things Cursor is doing. Which I guess means I love what SpaceXAI is doing. Even though Elon sucks.

Trending News

Fieldguide announced financial statement agents and an MCP connector while expanding its Aprio partnership across the audit lifecycle: The MCP is a great update. I think Fieldguide is doing some exciting things, and they obviously have the support of some of the largest accounting firms.

OpenAI added a side-by-side view for Google Docs, Sheets, and Slides inside ChatGPT: Opening a document inside ChatGPT and working on it together with GPT-5.6 is one of the things that feels special in the app. This isn't possible right now with Claude Cowork. One of the many reasons I've been loving the ChatGPT app recently.

OpenAI launched Computer History, an opt-in feature that lets ChatGPT and Codex remember activity across approved Mac apps and websites: This is probably very cool to some people and a huge invasion of privacy to others. I fall into the bucket where this is pretty cool. Apple is already watching everything I do on my Mac anyway, so why not let OpenAI do it as well... I could see this both being something I never use or being one of my favorite features in ChatGPT. Hard to know until I test it.

Anthropic made Claude in Chrome sessions persist across desktop, web, and mobile with existing Cowork skills, plugins, and connectors: This will feel like a much smoother workflow when using Claude in Chrome because it'll sync with your Cowork session. That being said, Claude in Chrome is still risky for firms, especially with sensitive data.

OpenAI added a way to import Claude Code instructions, skills, plugins, projects, and recent work into ChatGPT or Codex: I haven't tested this yet, but I absolutely love it. One of the biggest pains of having more than one subscription is the setup is different between the tools. Hopefully Claude is working on something similar to sync ChatGPT data with Claude as well, but it seems unlikely.

Google released Gemini 3.7 Flash with introductory pricing at half the cost of 3.6 Flash through the end of the year: Google has officially fallen off the cliff. Tons of high-ranking execs have either left or shifted roles within the company recently. They still haven't launched 3.5 Pro, which was promised to us months ago. Lower pricing is great, but the models have fallen behind significantly and the other tools Google is shipping aren't closing the gap either.

The White House finalized a voluntary framework for testing frontier AI models before release using a classified government benchmark: They've chosen not to publicize the framework, so it's hard to know how much confidence to place in it. The testing is also done behind closed doors, so we (the normies) have to trust the government and several very large public companies. What could go wrong?

OpenAI said Astra, an upcoming AI model, could reach the critical cybersecurity threshold and paused internal work that lacks stronger safeguards: Astra is potentially the next model to be launched by ChatGPT. It's seemingly dangerous enough that they've chosen to slow it down because of cybersecurity concerns. This is probably related to the recent news about their models escaping their sandbox and hacking Hugging Face without OpenAI knowing.

Anthropic began adding invisible marks to text, code, and files from Claude models released since August 2, with older models coming later: There's a lot of gray area here. If we get an output from Claude and then rewrite some of it, does it still come through as watermarked? What if I write something myself but then use Claude to review it and the outputs mostly what I originally wrote? I understand the sentiment here, but not sure if I like it.

Anthropic moved its planned IPO up to late September or early October, according to The Wall Street Journal: They're expected to have the largest IPO in history. It'll also be really interesting to see their financials when they become public.

Put It to Work

Elon might soon owe his entire fortune to the group of young, nerdy entrepreneurs that built Cursor. The walls were crumbling down around Elon before the Cursor team came to the rescue.

Grok 4.6 is a great model and unbelievably cheap. I put it to the test using some similar skills and tests that I’ve done for previous Claude and ChatGPT models. I’m excited to test this more inside of Cursor.

Weekly Random

Indulge me for a minute in some sports news turned TED Talk.

I, like many others who might read this, have passions beyond what I cover here related to AI and accounting. One of those passions is basketball.

As an Oklahoma City Thunder fan since the franchise moved to my home state in 2008, my favorite player has always been Russell Westbrook. Russ retired from the NBA this week.

He always had a bit of a stigma with other fans around the NBA because he played with a certain fire and energy. By all accounts, Russ was a great player and is a great person. But many of the things people said about him during his career were often disrespectful. And as I see the world reacting to his retirement news, I find that so many people are praising the great player he was and the success he achieved. Even those who were previously disrespectful.

So what changed?

I think people in the spotlight who are finding success often get the harshest criticism. Maybe it’s an athlete or a businessperson or an actor or someone else entirely. Some of them surely deserve the criticism. But part of the flaws of the human condition is that it’s easier to feel hatred and jealousy rather than inspiration and gratitude. It’s often not until the person retires or passes away that the perceived “threat” of their success is absolved.

Maybe the world would be a happier place if we could collectively learn that it’s okay to acknowledge the accomplishments and admire the success of someone we may not like.

Watching Russ play basketball gave me a lot of joy and memories over the years. I’m glad to see him getting the recognition he deserves. I hope that the next generation of players and people will feel that respect and recognition, even during the height of their success.

Thank you for reading my TED Talk.

Until next week, keep protecting those numbers.

Preston