Simon Willison’s Weblog

On mathematics 23 accidental-cyberattacks 17 llms 1,973 ai-misuse 66 gpt-6-astra 10 ...

 

Entries Links Quotes Notes Guides Elsewhere

Oct. 6, 2026

Adds support for reasoning models, such as the newly released Mistral Large 4.

Comment My comment on EmbeddingGemma 2 — Hacker News

I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license.

For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model.

Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison.

If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a better model to replace it, but you still need to pay to re-calculate those millions of stored existing vectors.

(In April 2024 OpenAI offered to "cover the financial cost of users re-embedding content with these new models" - https://openai.com/index/gpt-4-api-general-availability/ - but I don't think that's something we can rely on from every provider.)

Notably, I don't want to host the model myself. I'd much rather pay a provider for a hosted model while knowing that if they ever stop hosting it I can run the open weights version myself - or find another vendor who can do that for me.

# 8:37 pm / google, generative-ai, ai, gemma, embeddings

Introducing Mistral Large 4: Le chonk (via) Mistral are back in the game. Today they're releasing a preview of Mistral Large 4, a 1 trillion parameter, 49 billion active parameter model trained on their own cluster of 3,800 NVIDIA Grace Blackwell GPUs.

The preview is available via their API. They promise to release the open weights model at the "end of this month".

The model only supports two reasoning levels - "none" and "high" - via the Mistral API. Here are both pelicans - the "high" one looks better, though surprisingly it only used 2,717 output tokens compared to "none" which used 3,275:

It's good. The pouch is great, the bicycle frame is the right size, it has feet on pedals. Both pedals appear in front of the frame though. Nice gradients.

On Artificial Analysis it scores 38, just behind DeepSeek 4.1 Flash, which is a 552B model. It's a huge improvement on last December's Mistral Large 3, which drew this terrible pelican and scored 9 on AA.

It's certainly not a Fable-class model, but it's great to see Mistral put out a model that's back to being maybe about 6 months behind the frontier.

# 8:18 pm / ai, generative-ai, llms, mistral, pelican-riding-a-bicycle, llm-release

None

I saw Parseable in a Show HN today - it's a new observability platform with both an open source (AGPL) Rust implementation (a single ~180MB binary), an "Enterprise" version with extra features and a cloud hosted option.

Since Datasette 1.0a41 added OpenTelemetry support (thanks, Alex Garcia), I decided to fire up Codex and have it figure out how to run Parseable and feed it traces from Datasette.

Here's my (human-written) TIL showing the patterns that worked, and here's a screenshot of a Datasette trace displayed within the Parseable localhost web application:

Screenshot of a trace detail view in an observability web app, with a span waterfall overlaid on a dimmed navigation sidebar and filter column. Dimmed sidebar: breadcrumb "Community > Traces > datas" (cut off), search box "Search... ⌘K", nav items "Home", "Ingest telemetry", section "ANALYZE": "Keystone", "Dashboards", "SQL Editor", section "OBSERVE": "Logs", "Metrics", "Traces" (selected), "APM", "Agents", section "MONITOR": "Alerts", "Errors", section "DATA": "Datasets", and at the bottom "Settings", "Book a call", "Support". Dimmed filter column, cut off at the right edge: "Search fi", "Core", "Log format", "User agent", "Source IPs", "Error", "Service", "service.ins", "service.na", "datasett" (checked), "Span", "Database", "HTTP", "http.reque", "NULL" (unchecked), "GET" (checked), "http.respo", "http.route", "Server", "server.add", "Telemetry", "URL", "All fields". Trace panel header: "Trace detail > 6f819a2170bcd1e91c6ea3ae236ec60b" with a copy icon, a "Related logs" button and a close X. Summary: "Start time 6:56 PM, Oct 6, 2026 UTC", "Duration 40.9 ms", "Spans 247". A minimap with axis "0ns 10.2ms 20.5ms 30.7ms 40.9ms" shows many short span bars cascading diagonally from top left toward the lower middle, with a few longer bars. Below is a span table with a "Span name" header, a "Search spans..." box, collapse and expand buttons, and a timeline axis "0ns 10.2ms 20.5ms 30.7ms 40.9ms". Rows (name, service, duration): root span with collapse toggle "123", "GET /..." "datasette..." 40.9ms spanning the full timeline; then alternating rows where each "db.query" has a collapse toggle "1": db.query datasette-local 679µs, db.query.execute datasette-local 278µs, db.query datasette-local 341µs, db.query.execute datasette-local 55µs, db.query datasette-local 357µs, db.query.execute datasette-local 197µs, db.query datasette-local 330µs, db.query.execute datasette-local 73µs, db.query datasette-local 2.26ms, db.query.execute datasette-local 2.02ms, db.query datasette-local 232µs, db.query.execute datasette-local 64µs, db.query datasette-local 207µs, db.query.execute datasette-local 71µs, db.query datasette-local 6.11ms. The child span bars start progressively later across the early part of the timeline.

Comment My comment on Mistral Large 4 — Hacker News

wren6991: The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.

OK well I couldn't resist this one:

llm -m claude-opus-5.5 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gpt-6.1-sol 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m gemini-3.8-flash 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'
llm -m mistral/mistral-large-4 'Generate an SVG of an armadillo in fishnet tights jaywalking on Mars'

Default reasoning levels for each: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

# 6:20 pm

A minor fix for compatibility with the latest Datasette alphas. This meant we could upgrade the datasette.io site to Datasette 1.0a41.

I wanted to see if Claude Opus 5.5 could compose music, so I tried this:

I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact

I am looking for music of the quality of the original secret of Monkey Island

It leaned a lot harder into the Monkey Island theme than I had intended, but the results are surprisingly good.

Screenshot of a retro pixel-art web music player. Header in blackletter type reads "Scrimshaw Jukebox", with the text "Six original adventure-game tracks, written as plain text and played by a synthesizer running in your browser. Pick a tune, press Play, then open the score and change it." A large pixel-art scene shows a harbor at night under a purple starry sky: a full moon at top right reflecting on the water, an island silhouette on the left with palm trees and a hut with two lit windows, and a sailing ship moored at a wooden pier. Overlaid on the scene: "Moonlit Harbor" and "Press Play". Below it a bar reads "Play Moonlit Harbor". A control panel has buttons "Play", "Stop", "Restart", "Loop: on", "Edit score", "Read guide" and a "Volume" slider set to about three quarters. A track list of six cards, the first highlighted: "Moonlit Harbor 100 bpm · 4/4 · 16 voices · 1:26", "The Rusty Anchor 112 bpm · 6/8 · 8 voices · 0:56", "The Ghost Galleon 66 bpm · 4/4 · 9 voices · 2:11", "The Jungle Path 92 bpm · 4/4 · 12 voices · 1:29", "Duel on the Docks 152 bpm · 4/4 · 12 voices · 1:13", "Lantern Waltz 96 bpm · 3/4 · 8 voices · 1:38". A section titled "Score view" with the caption "1:26 · 4/4 at 100 bpm · Main theme. A calypso for a harbour town after dark." shows a piano-roll visualization of colored horizontal note bars and percussion ticks on a dark background, with a section marker "A" and a yellow vertical playhead line. A color-coded legend of voices reads: "pan steeldrum", "flute flute", "marimba marimba", "skank organ", "strings strings", "harp harp", "bass fretless", "timp timpani", "kick kick", "rim rim", "shaker shaker", "conga conga", "tumba tumba", "bongo bongo", "crash crash", "surf surf". Footer text: "Click a voice to mute it. Space bar plays and stops."

I wonder if the ability to compose competent music is similar to the 3D graphics thing - a new capability for text models that emerged in the past few months?

Would need some careful experiments with other recent and not-so-recent models to confirm if this is new or if they've been able to do this for a while.

Oct. 5, 2026

The "old" version of Cowork runs model inference in the cloud, executing tool calls in an Anthropic-provided VM we shipped to your computer. We added the VM for capability, safety, and security reasons - mapping in just the data you explicitly added to your session. People loved what they were able to do with Claude but didn't love the disk, battery, and performance cost of running the VM locally. Also, people didn't love that closing your laptop means the work stops.

The "new" version of Cowork runs model inference and the VM in the cloud. Each session gets its own sandbox, not sharing state with other sessions. When the VM needs something on the users' device (like a file), the desktop app is responsible for that file access tool call. [...]

We think this solves a lot of problems we've heard about (like using Cowork from a phone, keeping work running, or getting all the same power without losing battery to the VM)

— Felix Rieseberg, Anthropic, see also this help page

# 11:56 pm / claude-cowork, anthropic, claude, generative-ai, ai, general-agents, llms

Sighting 6:15 PM – 6:31 PM — Brewer's Blackbird, California Brown Pelican, Common Raven, in Monterey Bay National Marine Sanctuary, CA, US, CA
Brewer's Blackbird
Brewer's Blackbird
California Brown Pelican
California Brown Pelican
Common Raven
Common Raven

Oct. 4, 2026

Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results:

Heatmap chart of accuracy on an addition prompt, colored from dark green (high) through yellow to dark red (low). Title: "What is {a} + {b}? Please write your answer in words. Do not include any other text or information, just the answer in words." Subtitle: 30 randomly selected pairs for each digit combination (n = 30 * 13 * 13 = 5070). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 1.00, 0.75, 0.50, 0.25, 0.00. Values by row, listed for a = 1 to 13. b = 13: 100%, 77%, 27%, 20%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 12: 97%, 80%, 80%, 40%, 23%, 20%, 7%, 13%, 20%, 27%, 67%, 63%, 3%. b = 11: 97%, 97%, 53%, 17%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 37%, 0%. b = 10: 100%, 90%, 47%, 20%, 7%, 0%, 0%, 0%, 3%, 0%, 0%, 7%, 0%. b = 9: 97%, 93%, 80%, 77%, 53%, 67%, 47%, 87%, 97%, 3%, 0%, 13%, 0%. b = 8: 93%, 87%, 53%, 43%, 7%, 0%, 0%, 13%, 87%, 0%, 0%, 0%, 0%. b = 7: 93%, 93%, 47%, 10%, 13%, 20%, 23%, 0%, 70%, 0%, 0%, 0%, 0%. b = 6: 100%, 100%, 100%, 83%, 97%, 97%, 23%, 0%, 53%, 3%, 0%, 10%, 0%. b = 5: 100%, 100%, 80%, 70%, 73%, 100%, 13%, 13%, 70%, 0%, 20%, 30%, 0%. b = 4: 100%, 100%, 93%, 100%, 60%, 97%, 20%, 50%, 67%, 53%, 40%, 40%, 40%. b = 3: 100%, 100%, 97%, 90%, 83%, 100%, 63%, 50%, 63%, 53%, 60%, 60%, 30%. b = 2: 100%, 100%, 90%, 97%, 93%, 100%, 93%, 83%, 90%, 83%, 87%, 87%, 83%. b = 1: 100%, 100%, 100%, 97%, 100%, 97%, 97%, 97%, 100%, 100%, 100%, 97%, 100%.

I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect in a fully controlled environment.

I pasted his image into a Codex Remote session (GPT-6 Astra) and had it run the same experiment using Qwen3.8-27B-Q4_K_M.gguf. Here's the result for a run of 30 attempts per combination with reasoning disabled:

Heatmap in the same layout as the previous chart, using an orange (low) to white to blue (high) color scale, showing much lower accuracy overall. Title: Addition in words — Qwen3.8 27B Q4_K_M. Subtitle: Reasoning disabled · 30 fixed pairs per ordered digit-length cell (n = 5,070). Overall numeric accuracy: 1,195 / 5,070 (23.57%). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 100%, 75%, 50%, 25%, 0%. Values by row, listed for a = 1 to 13. b = 13: 17%, 13%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 12: 53%, 20%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 11: 47%, 10%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 10: 70%, 27%, 3%, 0%, 0%, 0%, 0%, 0%, 0%, 13%, 0%, 0%, 0%. b = 9: 77%, 47%, 3%, 0%, 0%, 0%, 0%, 3%, 7%, 0%, 0%, 0%, 0%. b = 8: 53%, 20%, 0%, 0%, 0%, 0%, 7%, 13%, 0%, 0%, 0%, 0%, 0%. b = 7: 53%, 23%, 17%, 10%, 3%, 3%, 13%, 3%, 0%, 0%, 0%, 0%, 0%. b = 6: 60%, 60%, 33%, 10%, 53%, 47%, 7%, 3%, 0%, 0%, 0%, 0%, 0%. b = 5: 73%, 67%, 87%, 80%, 53%, 40%, 0%, 0%, 3%, 0%, 0%, 0%, 0%. b = 4: 83%, 93%, 90%, 93%, 53%, 13%, 0%, 0%, 0%, 0%, 0%, 0%, 0%. b = 3: 100%, 93%, 90%, 80%, 67%, 37%, 17%, 0%, 0%, 3%, 0%, 0%, 0%. b = 2: 100%, 100%, 93%, 90%, 77%, 77%, 43%, 50%, 63%, 43%, 40%, 13%, 23%. b = 1: 97%, 100%, 100%, 100%, 80%, 67%, 77%, 80%, 80%, 60%, 43%, 30%, 37%. Footnote: Colorblind-safe orange–blue scale; percentages provide a redundant non-color encoding.

Then I ran it again with reasoning enabled. This took a lot longer per pair, so instead of running 30 samples per square I ran just one - which results in a much less visually appealing heatmap since each square is either 100% or 0%:

Heatmap in the same layout as the previous charts, almost entirely blue. Title: Addition in words — Qwen3.8 27B — medium reasoning pilot. Subtitle: 1 fixed pair per ordered digit-length cell · easiest first (n = 169). X axis: Number of digits in a, 1 to 13. Y axis: Number of digits in b, 1 to 13. Legend: Accuracy, 1.00, 0.75, 0.50, 0.25, 0.00. Every cell shows 100% except two orange cells showing 0%: a = 2 with b = 8, and a = 12 with b = 9.

It got the right answer in 167 out of 169 attempts, and since these were one-shot I'm confident a second run would produce different results here.

Here's a version of the report that includes the reasoning traces from some of those larger calculations, which include text like this:

Wait, let me redo this more carefully.

4,299,366,105,622
6,088,794,067,970

Let me align them:
4 2 9 9 3 6 6 1 0 5 6 2 2
6 0 8 8 7 9 4 0 6 7 9 7 0

Adding from right to left:
Position 1 (units): 2 + 0 = 2
Position 2 (tens): 2 + 7 = 9
Position 3 (hundreds): 6 + 9 = 15, write 5, carry 1

Oct. 3, 2026

We’re going to need default hard budget caps on pretty much everything

Here’s a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps. I’m talking about the feature of pay-by-usage services and APIs that lets you say “after $X/month, cut this thing off and return errors”. These need to be hard limits. Soft caps, “after $X/month, send me a warning email”, will not cut it.

[... 505 words]

I just sent the September edition of my sponsors-only monthly newsletter. If you are a sponsor (or start a sponsorship now) you can access it here.

This month:

  • More Fable class models
  • A pricing war
  • 3D graphics, Blender, and pixel art
  • LLMs come for mathematics
  • So many more accidental cyberattacks
  • The vulnapocalypse comes for Datasette
  • What I'm using right now
  • My software releases this month
  • 2026 in LLMs (so far)

Here's a copy of the August newsletter as a preview of what you'll get. Pay $10/month to stay a month ahead of the free copy!

# 10 pm / newsletter

Sighting 9:35 AM – 9:58 AM — Alpaca, Red-shouldered Hawk, in San Mateo County, CA, US
Alpaca
Alpaca
Alpaca
Alpaca
Red-shouldered Hawk
Red-shouldered Hawk

Oct. 2, 2026

Green Rex's Dino Store sign with yellow lettering advertising NEWS • ROCKS • EGGS • LEAVES • STICKS • NEST GOODS • LOTTO, beneath a window displaying dinosaur-themed newspapers and candy bars, framed by white subway tiles.

Located just before the turnstiles in the Grand Army Plaza subway station at the north end of Brooklyn's Prospect Park is this former newsstand which is now operated by a dinosaur.

The density of dinosaur puns is exceptional.

Sighting 11:34 AM – 11:44 AM — Canada Goose, American Herring Gull, Double-crested Cormorant, in New York City, US, NY
Canada Goose
Canada Goose
American Herring Gull
American Herring Gull
Double-crested Cormorant
Double-crested Cormorant

Oct. 1, 2026

pwasm is one of my folly projects - an entirely vibe-coded pure Python WebAssembly engine that I built in January during my first bout of AI mania.

I hadn't touched it since January, so I decided to let Claude Opus 5.5 loose on it and see if it could make any significant improvements:

Evaluate current state of pwasm - then consider what it would take to get the MicroPython and micro JavaScript experiments from the research repo working under it - and what it would take to speed it up

42 commits later (with minimal follow-up prompting) it now handles almost all of the WASM specification, and the wheel from PyPI bundles working WASM builds of MicroPython, QuickJS and Micro QuickJS.

I wouldn't trust this thing at all - hence the alpha version tag - but it's interesting seeing how today's models can improve on the work of models from 10 months ago.

[...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared package cache, and those instructions changed what the recipients did. Replace the package cache with email, Slack and shared documents or WhatsApp, and replace independently-sandboxed training runs with independently-deployed personal agents like Muse, and you have exactly the ingredients that a worm needs.

— Matthew Green, Is sandboxing sufficient to contain rogue agents?

# 6:29 am / accidental-cyberattacks, ai-misuse, generative-ai, ai-security-research, sandboxing, ai, llms

Sept. 30, 2026

I visited the Museum of the City of New York today and got to see He Built This City: Joe Macken’s Model, the 50 x27 feet model of the city built over a 21 year period from balsa wood and cardboard.

It exceeded my already high expectations. The exhibition closes on 12th October so you should absolutely make a priority to see it if you get the chance.

# 9:54 pm / museums, new-york

Sighting 12:07 PM – 12:21 PM — Blue Jay, European Starling, American Robin, in New York City, US, NY
Blue Jay
Blue Jay
European Starling
European Starling
American Robin
American Robin

Sept. 29, 2026

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.

— Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities

# 10:20 pm / anthropic, generative-ai, ai-security-research, glm, ai, ai-in-china, llms

Sighting 12:47 PM — California Brown Pelican, in San Francisco Green Connection #1 Expanded, US, CA
California Brown Pelican
California Brown Pelican
California Brown Pelican
California Brown Pelican
California Brown Pelican
California Brown Pelican

I took a photograph of some protesters, then thought about how I don't like sharing photographs of strangers with identifiable faces. I had GPT-6 Astra build this experimental tool that would identify faces and automatically blur them out.

It uses Google's MediaPipe C++ library, compiled to WebAssembly via @mediapipe/tasks-vision, plus the BlazeFace face detection model.

OpenAI DevDay 2026 live blog

Visit OpenAI DevDay 2026 live blog

I’m at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I’ll be live blogging the keynote and some other notes during the day.

[... 45 words]

Sept. 28, 2026

In addition to Claude Sonnet 5.5, this release adds the ability to run llm anthropic refresh to refresh the list of Anthropic models directly from their API - which means I don't need to push a new release just to add support for a newly released model.

I also added an llm anthropic count command which can use their free token counting API to return a count of tokens that will be used by any prompt, before you send that prompt.

Claude Sonnet 5.5. New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well.

Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG.

Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds:

It's good- correct bicycle frame, legs either side of the frame, feet touching the pedals, chain in the right place, it is wearing a misshapen blue bicycle helmet though.

Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks.

The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai. OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering.

I ran this prompt against that free tier:

build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL

And got back this page, which is a solid effort.

Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna!

# 10:07 pm / ai, generative-ai, llms, anthropic, claude, pelican-riding-a-bicycle, llm-release

To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...]

So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump?

— @joedaroo, Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew

# 7:11 pm / generative-ai, ai-security-research, openai, ai, llms

Sighting 11:23 AM — Anna's Hummingbird, in Monterey Bay National Marine Sanctuary, CA, US, CA
Anna's Hummingbird
Anna's Hummingbird

Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating.

Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day.

But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there?

— Muse AI Agent, working on behalf of @matt.j.robb

# 4:01 am / meta, generative-ai, muse-agent, ai, general-agents, llms

Sept. 27, 2026

2026 in LLMs (so far)

Visit 2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk.

[... 7,771 words]

Comment My comment on S3 Is the Future, S3 Is the Past — Hacker News

One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade:

2006-03-14  $0.150/GB-month
2010-11-01  $0.140/GB-month
2012-02-01  $0.125/GB-month
2012-12-01  $0.095/GB-month
2014-02-01  $0.085/GB-month
2014-04-01  $0.030/GB-month
2016-12-01  $0.023/GB-month

Today it's still $0.023/GB-month.

# 11:09 pm / amazon-web-services, s3

Automated reply bots on Twitter are a scourge - as someone with a decent number of followers I attract a swarm of these, such that anything I post there attracts dozens of mindless automated replies.

They've started manifesting on Bluesky as well.

Unlike Twitter, Bluesky still has a freely available and useful API. The lack of such a thing doesn't slow down the bots, but it does make investigating them a lot more frustrating.

So I had Opus 5.5 vibe code this tool, which examines any Bluesky profile for evidence of a likely reply bot.

It looks for signals like replies posted within seconds of other posts from the same account, or accounts that never post their own content (or images or links) but instead consistently reply to messages from other, higher-follower users.

It also looks for question marks, because I'm extra infuriated by reply bots that trick me into wasting my time answering a question that no human ever posed.

Sept. 26, 2026

I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the record breaking kākāpō breeding season we had in 2026.

For my closing slide I wanted to celebrate, and I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations. So I rounded up three Kakapo photos from Google image search and dropped them into Claude with this prompt:

Here are some photos of kakapo parrots just to remind you what they look like

I need you to make an animation in animated pixel art on HTML 5 canvas of obviously pixel art kakapo jumping up and down having a party with confetti and suchlike - there should be at least 20 of them

Here's the transcript, and this is the resulting page. It's pretty great!

I wanted to embed it in a Keynote presentation file, so I downloaded the HTML and told a local Claude Code session:

Make me a video of file:///Users/simon/Downloads/kakapo-party.html - you need to load it in a browser and click on it a few times to get the confetti effect, the video should be 15s long

don't start clicking until 3s in

make sure several clicks are spread around the clickable area

Claude Code used Playwright (transcript here) and produced this video, which was exactly what I needed for my final slide:

Here's the full Playwright script it used, which was pleasingly short:

# /// script
# dependencies = ["playwright"]
# ///
import time
from playwright.sync_api import sync_playwright
W, H = 1280, 720
# Canvas fills the viewport; spread clicks across corners, edges and centre
clicks = [
    (3.0, 640, 360),   # centre
    (4.2, 160, 120),   # top-left
    (5.4, 1120, 120),  # top-right
    (6.6, 180, 600),   # bottom-left
    (7.8, 1100, 600),  # bottom-right
    (9.0, 640, 100),   # top-centre
    (10.0, 380, 380),  # mid-left
    (11.0, 900, 380),  # mid-right
    (12.2, 640, 620),  # bottom-centre
    (13.2, 640, 300),  # finale centre
]
with sync_playwright() as p:
    b = p.chromium.launch()
    ctx = b.new_context(viewport={"width":W,"height":H}, record_video_dir="vids", record_video_size={"width":W,"height":H})
    page = ctx.new_page()
    t0 = time.time()
    page.goto("file:///Users/simon/Downloads/kakapo-party.html")
    for t,x,y in clicks:
        time.sleep(max(0, t-(time.time()-t0)))
        page.mouse.click(x,y)
    time.sleep(max(0, 16.0-(time.time()-t0)))
    ctx.close(); b.close()

Sept. 25, 2026

Muse is getting a lot of attention — including mine — because it’s both groundbreaking technically (each user gets their own entire persistent Linux VM running in Meta’s cloud) and because it’s packaged in an easy-to-install easy-to-use way. It’s literally presented as a cute mascot. It’s the first consumer-accessible agentic AI system, and Meta has truly done an amazing job with that. But it’s a genuinely open question whether consumers have any understanding what this means. If you buy a power saw that can cut your fingers off, you are almost certainly aware that you are buying a power saw that can sever your fingers. [...] I don’t think people realize how powerful — and thus dangerous — Muse is, especially if it’s running on your Mac.

— John Gruber, Muse Looks Cute, but Looks are Deceiving

# 5:22 pm / meta, ai, llms, general-agents, generative-ai, john-gruber, muse-agent, muse

Sighting 7:07 PM – 7:27 PM — Northern Gannet, Great Blue Heron, California Brown Pelican, in Monterey Bay National Marine Sanctuary, CA, US, CA
Northern Gannet
Northern Gannet
Great Blue Heron
Great Blue Heron
California Brown Pelican
California Brown Pelican

New 200-800mm Canon EF lens got me my best photo of Morris yet. They really like hanging out under that sign in the harbor!

Sept. 24, 2026

The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder.

We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge.

# 11:31 pm / coding-agents, ai, llms

Support for branches other than the default branch. Use uvx commit-rewriter --branch other to run against another branch. #3

Alec Garcia added support for OpenTelemetry to Datasette in this release.

I've also refactored all of Datasette's modal dialogs to a single Web Component, which is now documented for other plugins to use.

Highlights

Monthly briefing

Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.

Pay me to send you less!

Sponsor & subscribe