In partnership with

Hey friends 👋

Big week for "we swear our numbers mean what we say they mean." xAI and Cursor shipped a new model together, and OpenAI gave ChatGPT the ability to talk and listen at the same time instead of taking turns like it's on a walkie talkie. Let's dig in.

xAI Calls Grok 4.5 "Opus-Class." Its Own Chart Disagrees.

xAI dropped Grok 4.5 this week, built jointly with Cursor using trillions of tokens pulled from actual developer sessions inside the editor. Musk's framing on X was that it's an Opus-class model, just faster and cheaper. His own launch page tells a messier story. On Terminal-Bench 2.1, Grok 4.5 does edge out Opus 4.8. But on SWE-Bench Pro, it lands behind both Opus 4.8 and Anthropic's Fable 5. On the DeepSWE benchmarks, it trails GPT-5.5 by double digits in one version.

Where it actually wins is efficiency. xAI says Grok 4.5 uses roughly four times fewer output tokens than Opus 4.8 to solve the same SWE-Bench Pro tasks, and it's priced at $2 per million input tokens against $5 for the frontier competition. It's now the default model in Grok Build and it's live across Cursor for every plan.

There's also an asterisk buried in Cursor's own post: an earlier snapshot of the Cursor codebase got accidentally folded into Grok 4.5's training data, which inflated its score on Cursor's internal benchmark. That metric got pulled from the public comparison because of it. So the "cheap and fast" pitch holds up. The "beats Opus" pitch depends heavily on which chart you're looking at, and Musk is choosing the generous one.

Porkbun is the domain name registrar you need.

Still using GoDaddy or Namecheap? There’s a better way with Porkbun!

Porkbun is the domain registrar trusted by creators, developers, entrepreneurs, and folks who want low prices without the nonsense.

Why people are choosing Porkbun:
• Most domains sold at cost
• Low, transparent registration and renewal pricing
• Free features like WHOIS privacy and SSL certificates
• Powerful web and email hosting options
• Real human support 24/7, 365 days a year
• Named the #1 domain registrar by Forbes Advisor and USA Today

For launching a business, building a personal brand, starting a side project, or creating your first website, Porkbun makes it easy.

Get $1 off your next domain registration with Porkbun now.

ChatGPT Can Now Talk Over You, on Purpose

OpenAI rolled out GPT-Live this week, a new voice architecture that processes what you're saying while it's still speaking, instead of waiting for you to stop. That sounds like a small technical tweak but it's actually the thing that made every voice assistant feel like a phone tree. The old Advanced Voice Mode had to detect silence to know when to respond, so a cough or a pause got read as "your turn," and the model would barge in at the wrong moment.

The more interesting part is what's happening under the hood. GPT-Live itself isn't the smart model. When you ask it something that needs a real web search or actual reasoning, it quietly hands the question to GPT-5.5 in the background and keeps the conversation going while it waits, backchanneling with "mhmm" and "got it" so it doesn't feel like dead air. It's rolling out globally now to ChatGPT users on iOS, Android, and the web, with an API for developers coming later through a signup form.

It's a good idea dressed up as a small feature. Splitting "sound like a person" from "actually be smart" from the same model is probably where every voice assistant ends up eventually. OpenAI just got there first.

A few more things worth knowing

  • Cognition released SWE-1.7 for Devin, built by doing more reinforcement learning on top of Moonshot's already heavily RL'd Kimi K2.7 model. The result jumped from 9.4% to 42.3% on Cognition's own FrontierCode benchmark, which is close to GPT-5.5's 43% and about 4 points behind Opus 4.8. Cognition's pitch is cost: about $1.97 per task, run on Cerebras hardware at 1,000 tokens per second. The finding that RL keeps paying off even on an already-optimized model is the more interesting story here than the benchmark score itself.

  • All three of these launches happened within about 24 hours of each other in early July, which says something about how compressed the release cycle has gotten. Nobody's waiting for a quiet week anymore.

That's what stood out to me today. Reply and tell me what caught your eye, I read everything.

Talk tomorrow,
Hatman 🎩

Keep reading