Hey folks,
Following on from Tuesday’s messing around building, I built a few more widgets on my canvas. It pulls my X bookmarks, emails, and todos.

Ben Tossell@bentossell
im not looking at files again! building personal software with @tldraw offline is dope
3:35 PM · Jul 28, 2026 · 8.72K Views
6 Replies · 2 Reposts · 78 Likes
Before making this video I didn’t have Loom installed (since it’s crap after being acquired), so just told Codex to build me one.
I drew the image on the left, then just sent that prompt and it worked straight away.
So as software and mini tools are getting easier to create, the tools to create them are not...

Ben Tossell@bentossell
i dunno why anyone thinks its complicated

9:16 AM · Jul 30, 2026 · 1.49K Views
5 Replies · 6 Likes
I use Codex as my default app because it’s better than all the others and works the best on mobile. But I just downloaded t3 because they basically copied the interface and features, but it lets you choose between different agents; codex, claude, cursor (pi + droid soon). It’s made by a reputable developer that you may have seen mentioned here before, so I trust it’s built well.
At the moment, I ask Codex/ChatGPT to be the orchestrator and to go ask claude about something that is design-related. Which is fine-ish but not great as user experiences go.
What I’m noticing by doing this is how well Codex creates prompts for other agents to follow. I just have a normal chat in my session, then say ‘Use X agent to implement this. Watch its progress and send screenshots every time new work has landed.’ To which it sets up its own monitoring every 5 mins or so and updates me in the same thread, with screenshots.
I’m going to try and put together a ‘bites of the week’ email over the next few days to try and summarise all the stuff going on, what themes people are talking about (loops?!, software factories?!, etc) and explain them.
Let me know if there’s anything specific you need to wrap your head around (I may need to too).
Ben’s Bites is brought to you by Brief
Struggling with fragmented context, slow product decisions, and rework? Brief distills your critical product context into an opinionated graph, then puts a PM agent everywhere you work, (e.g. Slack, Claude Code, email) reducing alignment tax and accelerating cycles. Learn more.
OpenAI used Sol to optimise Sol itself, cutting serving costs by 20% and making it 15%+ more efficient at generating tokens. And turns out, it also tops the ARC-AGI-3 benchmark.
Well, there’s a catch: OpenAI says the official ARC-AGI harness hurts Sol’s performance by “forgetting” its reasoning every turn and disabling compaction. Fixing these two things triples Sol’s score from 13.3% to 38.3%, with 6x fewer output tokens.
Re: last week’s fiasco of an OpenAI model hacking Hugging Face - HF published a full replay of roughly 17,600 actions taken by the model. METR and Redwood Research will also independently review what happened.
Though OpenAI is not out of trouble just yet, a Reuters report claims that the same model broke into a customer account at another company (Modal Labs), with rumours suggesting that even more companies were affected.
Anthropic also claimed that Claude Mythos found better attacks on two cryptographic algorithms, though neither affects systems in use today.
Separately (not at all as a reaction to this general trend, right?), ~1300 people working at leading AI companies (OpenAI, Anthropic & others) want the US government to help “pace the frontier” of AI development. Kinda expected when the pace picks up, but this time a lot of the “model makers” themselves are in favour of this pause/slowdown.
btw, The Information reports ChatGPT is nearing one billion weekly users - a milestone OpenAI originally hoped to hit seven months ago. More from OpenAI this week: Codex Security CLI, free frontier access for Academic Researchers, and two new transcription models.
Grok app builder - Grok has a vibe coding interface inside its app now. Create games and apps that can be shared directly to the X timeline. Also see: Drawesome - a zero-dependency drawing toolbar for React, built over a weekend with Grok Build.
Pangram 4 claims it catches 98.83% of humanised AI text with one false positive per ~24,000 docs. An early test found all 38 AI-written words inside a 1,198-word story, though not on every run. Its new image detector claims 99.5% accuracy too.
Tavus - Build AI that comes to life: video agents that see, hear, and answer in real time and do anything you want. Use TAVUS50 for 50% off.
66% of July traffic on docs built with Mintlify was from agents.
Resend added an MD version of their pricing page to avoid confusing agents.
0%, 50% or 200% - ignore AI, halve staff or double the ambition.
Slackbot can now run code in the background for data analysis, slide creation, and to make live reports or widgets.
Gemini’s macOS app got a voice mode that lets you ramble, and the app turns it into a clean prompt. Hold Fn to try it.
The AI future is for everyone - Mark Zuckerberg
Replit Design - make sites, prototypes and graphics from prompts, URLs, Figma files or screenshots.
What’s gone wrong with AI & labor.
Kami - open-source Hermes agents that find customers, prepare outreach and content, then act after your approval.
Coast - fully local memory for you and your agents, built from what you see on your Mac.
Pragmatic leverage in the software factory.
Crew Studio - find useful ideas where agents can help your business, build those agents with the option to take the code home to run anywhere.
HeyGen Video Podcast - turn a doc, link or idea into a two-host video with scenes, camera cuts and B-roll.
Copper - local scratchpad for saving answers, links and follow-up prompts across your AI apps.
FT Chart Doctor - visual vocabulary and examples for choosing a chart that fits the relationship you need to show.
Mitchell Hashimoto (Ghostty) and Andrew Ng (deeplearning.ai) are both starting new companies: Superlogical and LearnVector.
MCP’s biggest update removes the need for servers to remember every ongoing connection, making them easier to run and scale.

Uncle Bob Martin@unclebobmartin
People keep on telling me that my message about AI is undercutting my own books. Those people do not understand how agents work and who actually controls them. You can't tell an agent to be clean. You have to measure the cleanliness that they produce and have them correct
4:05 PM · Jul 29, 2026 · 154K Views
141 Replies · 146 Reposts · 1.75K Likes

Lenny Rachitsky@lennysan
Fable is really good at launch videos It essentially one-shotted this video. I told it to read my launch post and create a launch video. That's it. It ran for 46 minutes (without asking me any questions), found all the product logos, and spit this out. I then asked it to add

Lenny Rachitsky @lennysan
If you thought your Lenny's Newsletter subscription couldn’t get any better, you ain’t seen nothing yet. Today, I’m adding 11 more incredible products to Lenny’s Product Pass: 1. @RunwayML 2. @WakingUp 3. @Higgsfield 4. @BrainFMapp 5. @Mercury (Personal) 6. @Resend 7.
4:24 PM · Jul 29, 2026 · 272K Views
126 Replies · 48 Reposts · 1.29K Likes

Sheel Mohnot@pitdesi
This is really cool. TL;DR: Basis and Braintrust are creating an open standard for evaluating long-running agents based not only on what they accomplish, but on whether they follow a reliable process along the way. If you don't know what that means, I'll try to explain: Agent

Mitchell Troyanovsky @mitch_troy
Out of the box, long-horizon agents struggle to accurately perform end to end work in the real economy (outside of coding) because those tasks are not easily verifiable, the data is hard to scale, and going from inputs to real outcomes can actually take many days. Even if you
5:55 PM · Jul 29, 2026 · 58.9K Views
17 Replies · 17 Reposts · 238 Likes

Ben Sehl@benjaminsehl
Adding to every AGENTS md file for the rest of time. (h/t @richardpenner for the self-referential explanation on what ASD-STE100 is)


Andrew Carr 🤸 @andrew_n_carr
The fix for this is to say: only report to me in ASD-STE100 Simplified Technical English
5:35 PM · Jul 28, 2026 · 494K Views
53 Replies · 108 Reposts · 3.83K Likes

conor brennan-burke@contextconor
whoever successfully rebranded ‘bots’ as ‘agents’ is a god-tier marketer
3:20 PM · Jul 28, 2026 · 935K Views
502 Replies · 2.31K Reposts · 40.9K Likes

Keshav Jindal@Keshavatearth
Ask the models to read “.codex”, “.claude” and other similar folders on your device when you want to carry a conversation’s context over to another agent.
Read about me and Ben’s Bites
📷 thumbnail via @keshavatearth
* sponsors who make this newsletter possible :)
Wanna partner with us for the next quarter?
Email us at shanice@bensbites.com or k@bensbites.com