Hi folks,
I’ve not been building as much this week 😬 - hand, foot and mouth is back in our house (happens yearly with 3 young kids!), and I have to get prep done for my Stanford talk on the 28th - they invited me back!
I’ll be in SF that week, heading to OpenAI Dev Day on the 29th to see what goodies they have in store for us - I predict we’ll see some stuff today, and they’re bound to have a personal agent up their sleeves to release very soon (pure speculation though).
For my Stanford talk, I’ve actually gone back to revisit the reference manual/course I was making for months (and hated every version I made). But now, since doing my Friday posts, I feel like I’ve got a better feel for how to do the course. This time I’ve actually made lessons I’m happy with, which is already miles better than my last attempts.
I’ll try and whip up something for a post tomorrow, but we’ll see if time (and kids) permits.
Ben’s Bites is brought to you by Name.com
Ship domain integrations in hours with the name.com API. The API that powers Vercel, Lovable, and Netlify’s domain name services. Built for agents and human developers with OpenAPI spec and MCP support—integrate search, registration, and management without the engineering overhead. Start building.
Claude Cowork is merging with Claude Chat. No more separate tab for bigger tasks. All your connected apps, skills and context are available in a normal Claude.ai chat. It can also keep working after you close your laptop. OpenAI will likely follow this pattern soon and merge ChatGPT and ChatGPT Work.
Claude Artifacts is also getting dedicated products for Docs and Slides, with Claude Design moving into conversations too. Is Anthropic invading G Suite and Office?
Update on Claude Code’s latest experiment. It has a name → Claude Mods. Mods change how Claude Code itself looks and behaves. You can ask Claude to write these mods for you, so the app can fit how you work (well, Claude works and you play Tetris).
Gemini API has two new Live models: 3.8 Live and 3.8 Live Extended Thinking. These models take video and audio as input and output audio. They work great for use cases where you want the model to provide real-time assistance. I tried them in AI Studio for a couple of minutes, and didn’t notice any hiccups. It can be a great alternative to GPT-Live 1, since 3.8 Live is 7x cheaper.
New LLM killer model - Jev. Built by TypeSafe AI. The founder co-created ChatGPT. Jev doesn’t really generate text like LLMs. Instead, it creates probabilities for possible answers. Suitable for software that needs lots of quick judgments: what API to call next, which model to route a request to, or a trading bot. 5x cheaper inputs vs 5.6 Luna; free output tokens and 5.6 Terra-like performance on relevant tasks.
Union Alpha - new stealth model available through OpenRouter, Cloudflare and other providers. Outperforms 5.6 Sol on the DeepSWE benchmark at 5.6 Luna’s cost. Some rumours say it is a router; others say a GPT-6 variant (GPT-6 Luna??) or another GLM model (GLM-5.3-Flash launched this way). I had a good first chat with it; it spoke well, but then it revealed all my api keys in another session, and I’ve been getting ‘provider errors’ because it’s overloaded ever since….
Meta One - new subscription from Meta with extra AI features/usage in Muse and paid tools across Instagram, Facebook & WhatsApp.
Factory raised $200M at a $5B valuation. To celebrate, I updated my plugin for using Droid in the bb app. If I have to build anything ‘properly’ ie I care that it’s not vibe-slopped, I use Droid. And I’ve been using the Droid Core models a lot recently (open-source models) - they’re so quick, and not noticeably different from frontier models for most cases.
Inside OpenAI’s agentic software factory.
Can AI models build a T-shirt store? This was a cool read; also check out all the supporting docs she includes in the post.
Reception by ElevenLabs - an AI receptionist for small businesses.
How OpenAI plans to report models misbehaving.
Slack Code lets your team and coding agents plan, build and review together
Pion - agents for running fully autonomous companies.
Which models are best at searching the web?
Buzz - Your coding agent can ping your iPhone when it needs you.
Arrow 2 by Quiver - New SVG generation model. Their last one was way ahead of LLMs; curious how it compares with Astra now.
Code contracts - write down what your code must do, then have agents check it.
Self-updating docs from Mintlify. Now with ready-made automations.
An API to invent a dataset to train your models. (docs)
TanStack Markdown - a tiny parser for docs, blogs and streaming AI responses.
Building a company brain people actually want to use.
A benchmark to answer how often AI agents cheat.
Xiaomi’s AI team is streaming the RL run for their upcoming model.
What stays expensive when AI gets cheap?

Kevin Ngo@kevin_t_ngo
I made a website with 25 mini rooms, each with Claude keeping people company. Created with Claude Opus 5.

3:01 PM · Sep 16, 2026 · 285K Views
90 Replies · 119 Reposts · 2.41K Likes

Patrick Collison@patrickc
A quick EU founder survey: eusurvey.vercel.app.
12:22 PM · Sep 16, 2026 · 208K Views
58 Replies · 65 Reposts · 392 Likes

Shu@shuding
After one week, there are 5 more small models for specific tasks on the web: 1. gpu-time: natural language to JS date and time by @imarikchakma 2. gpu-query: natural language to structured filters by @cheatyyyy 3. neural-flexbox: model to centre a div (!) by @aaronvanston 4.…

Shu @shuding
I trained a small model to do syntax highlighting in the browser with GPU. Meet gpu-lexer from Vercel Labs: Small (27.5KB), fast (runs on WebGPU), and language-agnostic (model guesses the syntax). https://t.co/ELWO2Qtu0J It is experimental and built for learning!
7:54 AM · Sep 16, 2026 · 44.4K Views
20 Replies · 32 Reposts · 650 Likes

rachel@chaotictransfem
okay i made a quick skill to get the model to delete useless UI copy. surprisingly works quite well with opus. have fun! paste.rachel.systems/ukogokijaz.md

rachel @chaotictransfem
anybody have go-to skills for fixing bad labels/copy in web apps? cannot handle opus and astra adding unnecessary "it does [good thing] -- never [random error]" below everything and making all the labels weirdly verbose and self-aggrandizing
8:45 PM · Sep 16, 2026 · 6.87K Views
5 Replies · 1 Repost · 100 Likes

Yohei@yoheinakajima
the independence of AI evaluators is going to matter more... so i made it a "benchmark": evaluatorbench.com (research preview) source-linked directory of third-party evaluators of frontier AI. each one has an independence score you can take apart: change the weights, …
6:22 AM · Sep 16, 2026 · 8.4K Views
20 Replies · 3 Reposts · 60 Likes
Read about me and Ben’s Bites
📷 thumbnail via @keshavatearth
* sponsors who make this newsletter possible :)
Wanna partner with us for the next quarter?
Email us at shanice@bensbites.com or k@bensbites.com