30 Sep 2026

feedDjango community aggregator: Community blog posts

Weeknotes (2026 week 40)

Weeknotes (2026 week 40)

I have been at Django on the Med πŸ–οΈ and already wrote a lengthy post about that.

I did a lot of work on a DEP for adding import map support to Django which is currently also being discussed on the forum. Apart from that I'm not going to repeat anything from the post linked above, so check it out if you want to know more.

Motivated by a discussion I had at the sprint I also improved my release process. I now have a make-release script in my dotfiles which updates the CHANGELOG with the version, bumps the version itself in the repo and commits and tags the release. The rest is handled by trusted publishing. I'm now finally also properly handling patch releases so that you don't have to check the history to know what's in a patch release. I already did that for minor and major version bumps, but was a bit too lazy. Now I can be even lazier and still more correct, and it feels great.

Next, I refactored the static site generator script for this blog to be much faster. I now do not have to wait when saving before refreshing the browser. Much nicer.

Releases from the last three weeks

django-js-asset

django-js-asset 5.0a1 implements the API proposed in the DEP. This is an alpha release because the DEP is still being discussed and I don't want to break people's code again and again if, during discussion, it appears that the API should be different.

django-json-schema-editor, django-prose-editor and django-content-editor

The three alpha releases django-json-schema-editor 0.15a1, django-prose-editor 0.28a2 and django-content-editor 9.1a1 depend on django-js-asset 5.0a1 mentioned above and implement the necessary changes for the new import map definition style.

django-authlib

django-authlib 0.20 adds support for specifying the tenant when using Microsoft Entra ID.

feincms3-downloads

feincms3-downloads 0.6 includes translation fixes, uses different error codes when pdftocairo or convert are missing, and changed the PATH environment variable handling to be less annoying for local development.

30 Sep 2026 5:00pm GMT

Django on the Med

πŸ”— Links

πŸ“š Books

πŸŽ₯ YouTube

30 Sep 2026 3:00pm GMT

feedPlanet Python

Python Insider: Python Language Summit 2026

The 2026 Python Language Summit was hosted in KrakΓ³w, Poland as part of EuroPython 2026. There were 15 talks covering free-threading, Rust, garbage collection, type annotations, and more.

30 Sep 2026 12:00pm GMT

Python Insider: Lightning Talks (Python Language Summit 2026)

Lightning talks on a one-time ABI break, safer interruptions, EktuPy (Scratch but Python), an AGENTS.md file for CPython, and a call to read PEP 836.

30 Sep 2026 12:00pm GMT

Python Insider: PEP 827: Type Manipulation (Python Language Summit 2026)

Michael J. Sullivan presents PEP 827 and discusses a key design decision: how to store type annotations?

30 Sep 2026 12:00pm GMT

feedDjango community aggregator: Community blog posts

I don't write codebase documentation anymore

Hello everyone πŸ‘‹

Confession time: in 10+ years of writing software, I have never kept documentation up to date. Not once.

And I've tried! Confluence spaces, GitHub wikis, a docs/ folder, READMEs that start strong and stop being true three sprints later. It always goes the same way. Someone writes a nice page, the code moves, nobody touches the page, and six months later a new person reads it, believes it, and loses an afternoon. Updating docs always felt like a chore I owed someone, and I pay chores about as reliably as you'd expect.

A while back I started reading some really cool wikis that a paid AI service had generated for a few projects. Architecture overviews, flow diagrams, a page for every subsystem, all built from the code. I loved them. What I didn't love was that they lived on someone else's platform, and I couldn't shape what they said or how they said it. So I thought: I want my own. In my repo, in plain markdown, maintained by an agent, updated on every push.

So I built it. One of my projects now has a 95-page wiki: about 134,000 words and 67 Mermaid diagrams. I didn't write a single one of those pages, and it updates itself every time something lands on main.

This post is the full shebang: what the wiki looks like, how it works, every part that broke along the way, what it's still bad at, and the two files you need to copy to get it in your own repo (both are at the end of the post, complete).

Quick caveat before we start, because the title is doing some heavy lifting: what I stopped writing is codebase documentation, the pages that describe what the code does. There's still a small set of docs the wiki doesn't touch, and I'll get to those near the end.

What it looks like

The wiki lives in docs/wiki/ inside the repo, and it's two levels deep, never more:

docs/wiki/
README.md the index: every reader starts here
architecture-overview.md root pages: stuff that spans more than one section
getting-started.md
api/
README.md a section "hub": directory map, a diagram, a table of pages
app-and-routes.md a "leaf": one seam of the code
auth-and-sessions.md
...
web/
README.md
...
workers/
operations/
.outline.json which source files each page owns
.wiki-state.json the commit the wiki was last generated from

(The real one has eight sections, but they're very specific to the project, so I'm keeping this one generic.)

The project is a client one, so I can't show you the real pages. Here's the top of a leaf page, lightly trimmed:

> Auto-generated by the wiki skill from commit `20c2b08` on 2026-09-28. Do not
> edit by hand; changes will be overwritten.

# App Bootstrap, Routes and the Error Envelope

The FastAPI application is assembled in `backend/main.py`: `create_app` builds
the app with the `ClerkAuthMiddleware` and the aggregated router, and the
`lifespan` function verifies Clerk settings, constructs every provider-backed
service that was not injected for tests, and fills an empty Clerk mirror before
serving. [...]

## Key files

| Path | Role |
|--------------------------+---------------------------------------------------------|
| backend/main.py | `create_app` and `lifespan`: middleware, router, ... |
| backend/routes/errors.py | Registers the three exception handlers that render ... |
| backend/exceptions.py | `ErrorCode` vocabulary and the `ApiError` family ... |

After that you get sections on how the app gets built, the route families, the error format, and a Mermaid diagram of the startup sequence. At the bottom there's a ## Related section linking back to the hub and to the sibling pages it depends on. Every page has the same shape, and that turned out to be a big deal for the readers.

Who actually reads this?

Humans and agents.

For humans, it's been great for onboarding. When someone new joins the project I send them to the index and the "Reading order" list at the bottom of it, and they get a tour of the codebase that matches what's on main right now. Much better than whatever Confluence page someone last updated when they felt guilty.

But honestly, the agents are the heavy users. The coding agents working on this repo read it all the time. Instead of grepping around for ten minutes to figure out how auth works, an agent reads the index, jumps to the right hub, and loads the one leaf page it needs. It's a shortcut into the codebase, and a cheap one, since a leaf page is small enough to load whole.

The index even has a section written just for them:

## For agents

The lookup path is this index, then a section hub, then a leaf. Each hub's
directory map names the page owning each part of its tree, and
`docs/wiki/.outline.json` maps every source path to the page that describes it.

How it works: two files

The whole thing is two files:

File What it is
.claude/skills/wiki/SKILL.md The instructions: layout, page templates, size rules, three modes, a self-check. About 230 lines of prose, zero code.
.github/workflows/wiki.yml The plumbing: triggers, model choice, the commit step, and a check for runs that got cut off.

I split them on purpose. The skill never touches git: it reads code and writes markdown, and that's it. The workflow never decides what a page says: it runs the skill and commits whatever changed. Because of that, you can drop the skill into any repo, or run it by hand with /wiki full in Claude Code, and it doesn't care who's calling it.

On every push to main, GitHub Actions starts a headless Claude Code session with exactly one prompt: /wiki incremental. The session reads the diff since the last wiki run, figures out which pages describe the files that changed, rewrites those pages from the current code, and exits. Then a plain shell step commits docs/wiki/ back to main.

Pages follow seams

The first big decision was how to cut a codebase into pages. The skill calls the unit a seam: a boundary the code already has, like a package, a service, a pipeline stage, or a bunch of modules that always change together. Pages follow seams, and never file types or a fixed template like "one page for models, one for views".

The skill also has to find the seams on its own, from git ls-files and the code. There's no hardcoded list of directory names anywhere. On my repo it came up with eight sections, including some splits I never asked for: it broke the backend into the core service, identity, the answer pipeline, and the test suites.

The two state files

These two files are what make incremental updates possible.

.outline.json is the page map. Each page lists covers (the globs that make the page suspect when they change) and seeds (one to five files to start reading from):

{
 "file": "api/app-and-routes.md",
 "title": "App Bootstrap, Routes and the Error Envelope",
 "covers": ["backend/main.py", "backend/routes/**", "backend/exceptions.py", "..."],
 "seeds": ["backend/main.py", "backend/routes/__init__.py"]
}

The rule is that every tracked file belongs to exactly one page, and the most specific glob wins. So when a file changes, there's always exactly one page to re-check. If a file doesn't match any page, that's a hole in the outline, and the run has to fix it by extending a page or adding a new one.

.wiki-state.json is one line:

{"last_generated_sha": "a60e3827...", "generated_at": "2026-09-29T14:11:31Z", "mode": "incremental"}

The skill always writes this file last, and that one rule is the whole crash-safety story. If a run dies halfway (timeout, API error, whatever), the old SHA is still there, so the next run computes the same diff and does the work again. The work gets done a bit later, but it gets done.

Three modes

incremental runs on every push. It diffs last_generated_sha..HEAD, maps each changed file to its page, and rewrites each affected page from scratch using the current code. The diff only tells it where to look, and the old page is just a checklist of topics to re-verify, so a page always reads like it was written today (no "this was changed to…" edit logs). Then it rewrites the hubs and index tables that changed.

audit runs every Monday at 03:23 UTC. Ten incremental runs can each correctly decide "nothing to change here" and still, together, leave a page wrong. So once a week the audit re-derives the whole outline, fixes the structure (splits, merges, new pages, deleted pages), and checks every page's main claims against the code, oldest page first. A page that passes keeps its banner untouched, so a clean audit produces no diff at all.

full runs when I ask for it, or automatically when the state files are missing. It writes the outline first, then the leaves, the hubs, the root pages, and the index last, so every level describes pages that actually exist.

Incremental also bumps itself up to an audit when the diff touches more than ~40% of the tracked files, or when the recorded SHA doesn't exist anymore (someone rewrote history).

Oh, and the 03:23 is on purpose. GitHub's scheduler gets hammered at the top of the hour, so an odd minute dodges the delay.

Keeping it honest

The skill's top priority is literally written as "accuracy beats coverage". A wiki that confidently describes code that doesn't exist anymore is worse than no wiki, because humans and agents both trust it. So a big chunk of the skill is rules about that:

Before it writes the state file, every run does a self-check: links and anchors resolve, navigation works both ways (leaf to hub, hub to index), every file has exactly one owner, the outline matches what's on disk, and a few greps make sure there are no em dashes, no "now / recently / no longer", no ticket references, and no bold-label bullets.

Yes, I banned em dashes in my generated docs. I have strong feelings about em dashes.

The workflow

Here's the flow:

 push to main Monday 03:23 UTC manual dispatch
β”‚ β”‚ (pick a mode)
β–Ό β”‚ β”‚
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚
β”‚ dispatch β”‚ gh workflow run wiki.yml β”‚
β”‚ job │──────────────┐ β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β–Ό β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ wiki job β”‚
β”‚ checkout (full history) β”‚
β”‚ claude-code-action: /wiki <mode>β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β–Ό
did THIS run write the state file?
β”‚ β”‚
yes no
β–Ό β–Ό
commit docs/wiki, rebase, commit partial pages,
push to main fail the job (red run)

Some of these details took way more effort than they look like they should.

Re-dispatching the push

claude-code-action doesn't accept push events. It throws Unsupported event type: push and that's the end of that. So a tiny dispatch job (it runs for a few seconds) re-fires every push as a workflow_dispatch, which is one of only two event types the default GITHUB_TOKEN is allowed to trigger.

Then there's a second wall: the action refuses to run for bots by default, and a run dispatched by GITHUB_TOKEN shows up as github-actions[bot]. So the workflow sets allowed_bots: github-actions to let exactly that one bot through.

Avoiding the infinite loop

The wiki commits to main. That's a push to main. Which would trigger the wiki… πŸ€”

It doesn't, for two separate reasons. paths-ignore: docs/wiki/** means a push that only touches the wiki doesn't trigger the workflow, and the commit is pushed with GITHUB_TOKEN, which GitHub never lets trigger push-based runs. Either one would be enough. I like having both.

The newest run wins

A concurrency group cancels a run in progress when a newer push comes in, so the wiki is always generated from the latest main. This is safe thanks to the state-file-last rule: whatever the cancelled run didn't finish, the new run redoes.

Different models for different modes

Incremental runs happen on every push, so they should be cheap. Audits and full runs make the structural decisions, so they get the stronger setup. The models and the audit effort level are all repo variables, which means switching models is a settings change and not a commit.

Everything goes through Lazer Proxy (that's the ANTHROPIC_BASE_URL line in the workflow). Right now both slots run GLM 5.3. Audits and full runs get max effort, and incremental runs use the model's default effort (which, now that I'm writing it down, I should probably bump to high lol).

The commit step

The commit step is plain shell, and it always runs, even when the Claude step fails or times out, so generated pages never get thrown away. It commits as claude[bot] with [skip ci], then rebases onto the latest main and retries up to three times, because main has probably moved during a run that can take over an hour.

It also checks whether the run actually finished. It compares the state file's SHA and timestamp with the time the run started, and if this run didn't write the state file (and the diff wasn't empty), it still commits the partial pages, since the next run redoes that diff anyway, but then it exits 1 so I get a red run.

That check exists because of the worst bug in this whole project.

What broke along the way

Three trigger designs in one day

The very first day was all about getting it to run on push at all. I started with claude-code-action, hit the push rejection, and switched to calling the CLI directly (claude -p "/wiki incremental") in one job with a normal push trigger. That worked, and honestly it's simpler! But it meant owning the CLI install and version pin myself, and losing the action's run reports and GitHub integration. So the same day, I went back to the action and built the re-dispatch job.

There's a research doc in the repo comparing every option I looked at: workflow_run chaining, repository_dispatch, schedule-only, the raw CLI, and dispatch plus allowed_bots. The workflow_run option was sneaky. My CI workflows only run when their own project's files change, so a push that only touched the README wouldn't trigger any of them, and the wiki would silently skip that push. The workflow header links to that doc, so the next person who opens wiki.yml and thinks "why is this so weird?" gets an answer.

The first layout was flat

The first version of the skill wrote one flat level of pages. Later I restructured it into index, hubs, and leaves, and added "outline_version": 2 to the outline. An old outline without that field forces a full regeneration automatically, so the upgrade happened on the next push without me doing anything.

That regeneration was also the last time I touched the wiki by hand. Of the 111 commits to docs/wiki/, 105 are from the bot. My 6 are merge commits, one feature commit that happened to touch the directory, one hand edit on day one, and that regeneration.

Eleven green runs that did nothing

My favorite bug, in the "I want to throw my laptop into the sea" sense.

A big batch of changes landed at once, and the diff grew past ~100 files. At that size, the model decided (on its own, nobody asked it to) that the smart move was to hand the page rewrites to subagents. Which is very reasonable! In an interactive session that's exactly what you'd want.

But in a headless run, subagents launch asynchronously, and the process exits when the main turn ends. So every run went like this: spawn a bunch of agents, then end the turn with some version of "the agents are still running, I'll resume when they complete". The process exited, the job reported success, the commit step committed whatever was written so far (between 1 and 12 files per run), and the state SHA never moved.

This happened eleven times in a row, and all eleven runs were green πŸ™ƒ The diff snowballed to about 245 files before one run happened to do everything inline and finish.

The fix was three changes:

  1. --disallowedTools Agent in the workflow, so the model can't spawn subagents at all.
  2. A line in the skill: do every step in your own turn and never hand page writing to subagents or background tasks, because a headless run ends when your turn ends.
  3. The "did this run write the state file?" check from above, so a run like that shows up red.

The third one is the one I care about most. The subagent thing was a bug, sure, but what really bothered me was eleven green checkmarks telling me everything was fine.

Nineteen refused commands

The next problem showed up on an audit that split a page in two and then couldn't delete the old one.

Without an explicit allow rule, a headless Claude Code session refuses any shell command it can't prove is safe. That includes git -C, anything with a pipe, small Python helpers, and the rm that deletes a page dropped from the outline. That audit lost 19 commands this way. The skill has an allowed-tools list in its frontmatter, but those rules match on command prefixes, so rm docs/wiki/foo.md was allowed and the exact same delete with an absolute path wasn't.

Two fixes here. The skill now tells the model to always run commands from the repo root with relative paths, never with cd or git -C in front. And the workflow allows Bash outright. That second one is a judgment call: it's fine for me because the prompt is a constant string, the repo is private, and the job token can only write repo contents, so there's nothing untrusted for the sandbox to protect against. If your repo is public, think twice before you copy that line.

The numbers

Measure Value
Pages 95: 1 index, 4 root pages, 8 section hubs, 82 leaves
Words ~134,000
Mermaid diagrams 67
Commits to docs/wiki/ 111 (105 by the bot)
Workflow runs 136: 101 succeeded, 28 cancelled, 7 failed

Most of the cancellations are by design: a newer push cancelled an older run.

Here's what a normal incremental run looks like. The diff touched 24 of the repo's 1,029 tracked files. The run took 23 minutes and 162 turns, cost $5.29, and changed 25 files in the wiki: 14 pages rewritten, 4 new pages (existing pages had grown past 2,500 words and had to split), plus the hubs and the index. At the end it left a summary in the CI log, including a list of pages that need the next audit to clean them up.

On models: an audit on Claude Opus got killed at the 50-minute timeout I had back then without finishing. I switched to GLM 5.3 at max effort, and the first audit on that setup finished in 13 minutes for about $4.30.

What I don't have is a monthly total, sadly. The wiki shares its API key with the Claude workflows that review our PRs in CI, so the spend on that key is both of them mixed together, and I can't tell how much of it is the wiki. The cost also depends on how big each diff is and how often you push, so your mileage will vary a lot. If you want a rough number for your repo, multiply a few dollars per run by your pushes to main, then add one audit per week. I'm giving the wiki workflow its own API key so I can track exactly what it spends, and I'll update the post when I have real numbers.

What it's still bad at

It's not perfect, so here's where it falls short.

Walls of text

This is the big one, and it's what I'm tuning next. The skill says paragraphs are at most ~120 words and hubs stay under ~600. My architecture overview has a paragraph that's 524 words long πŸ˜… Several hubs are between 870 and 1,050 words, and a dozen leaves are past the 1,500-word "you should split this" line.

My theory: the model follows the hard rules that the self-check measures (it always splits a page past 2,500 words) and treats the soft ones as suggestions. Agents don't mind a wall of text. Humans very much do. I think the fix is turning more of the soft rules into self-check failures, because everything the self-check measures gets fixed, and everything it doesn't measure slowly drifts.

It's sometimes wrong

I've caught pages getting things wrong. I don't lose sleep over it, because it fixes itself: the next time anyone touches those files, the page gets rewritten from the code, and even if nobody does, the Monday audit re-checks every page. When a human-written page is wrong, it stays wrong until a human notices. This one is on a timer.

I also want to be upfront about what I haven't done: nobody has fact-checked the wiki sentence by sentence. The one independent check I ran confirmed the structure was right: the outline matched what's on disk, the links worked, and no page was behind the files it covers. That tells me the wiki is structurally sound, but it doesn't tell me every sentence is true. For the content, I'm trusting the verify-against-code rule, the self-check, and the weekly audit. So far that's been good enough for me, but it's a bet, and you should know it's a bet.

And one that made me laugh: the run summary at the end of each CI log (the one thing the self-check doesn't grep) is full of em dashes and bold-label bullets. Apparently the rules only count when someone is checking.

What the wiki doesn't write

Back to the caveat from the beginning.

The wiki describes what the code does right now. It can't know intent: why we picked this design, what we decided in a meeting, how an operator should run a data refresh. So the repo still has a small set of docs that live outside the wiki:

I don't write these by hand either, to be clear. I write them with AI. The difference is that I'm driving: the decision or the procedure comes from me (or the team), and the AI helps me turn it into a doc. The wiki doesn't need me at all.

The handbook also has a rule, backed by an ADR: if you change an admin page or the data refresh workflow, you update the handbook page that describes it in the same pull request. That rule lives in the agent rules too, so the coding agents follow it.

My favorite detail: the wiki has a page about the handbook, explaining what it is, why it's maintained outside the wiki, and which wiki pages describe the code behind each handbook doc. The wiki documents the docs it isn't allowed to touch, which I find hilarious.

Stop writing the docs a machine can write

OK, opinion time.

Docs that describe code are basically a build artifact. We don't hand-write compiled binaries or minified JS, we generate them from the source every time the source changes. For "how does this work" docs, the source is the code. And now we have something that can read the code and write decent pages about it for a few dollars a run. Asking a person to keep those pages in sync by hand is how you end up with a Confluence graveyard again.

Docs about decisions, intent, and procedure are different. Those come from people, so a person should be in the loop, with AI helping. And that set turned out to be way smaller than I expected: a handful of files and two directories, versus 95 pages I never have to think about.

Show me the code

Both files are below, complete. To use them in your repo:

  1. Copy SKILL.md to .claude/skills/wiki/SKILL.md and wiki.yml to .github/workflows/wiki.yml. Make sure docs/wiki/ isn't git-ignored.
  2. Set up model access. My workflow uses a LAZER_PROXY_API_KEY secret and a LAZER_PROXY_BASE_URL variable because everything goes through our proxy. If you're calling Anthropic directly, put your API key in a secret, point anthropic_api_key at it, and delete the ANTHROPIC_BASE_URL line. If you use another gateway, point that line at it instead. The model and effort variables are optional. Without them, the workflow uses Sonnet for incremental runs and Opus for audits.
  3. If main is protected, let the workflow push (a GitHub App or a ruleset bypass), or change the commit step to open a PR.
  4. If your repo is public, rethink two choices I made for a private repo: show_full_output: true (it dumps the whole transcript into the public run log) and allowing Bash outright.

Then push something! The first run finds no state file and switches to full by itself. On a big repo that first run takes a while, so go get a coffee β˜•

You can also skip CI entirely: open Claude Code in your repo and type /wiki full.

SKILL.md

---
name: wiki
description: Generate and maintain an agentic codebase wiki in docs/wiki/ (browsable markdown pages with Mermaid diagrams). Use this skill whenever the user asks to build, update, sync, audit, or regenerate the project wiki, codebase documentation, or architecture docs, or whenever it is invoked as /wiki. Also use it when asked "document this codebase" or "keep the wiki up to date".
allowed-tools: Read, Grep, Glob, Write, Edit, Bash(git log:*), Bash(git diff:*), Bash(git show:*), Bash(git ls-files:*), Bash(git rev-parse:*), Bash(date:*), Bash(wc:*), Bash(ls:*), Bash(cat:*), Bash(grep:*), Bash(find:*), Bash(jq:*), Bash(python3:*), Bash(rm docs/wiki/:*)
---

# Codebase Wiki Generator

Maintain a browsable markdown wiki describing this repository in `docs/wiki/`. The wiki is machine-owned: every run may rewrite any page, so treat existing pages as prior output, not as human work to preserve. Write files only. Never run `git add`, `git commit`, or `git push`; committing is the caller's job (CI or the human).

Accuracy beats coverage. Every claim in a page must be something you verified by reading the code in this checkout. A wiki that confidently describes code that no longer exists is worse than no wiki, because readers (human and agent) trust it as context.

Readers are humans browsing on GitHub and agents loading pages as context. Both navigate the same path: index, then section hub, then leaf. Every rule below exists to keep that path short and every page on it accurate.

## Invocation and mode selection

The invocation is `/wiki [mode]` where mode is `incremental`, `audit`, or `full`. Rules:

1. Run `full` regardless of the requested mode when any of these hold: `docs/wiki/.wiki-state.json` does not exist; `docs/wiki/.outline.json` is missing, unparseable, or lacks `"outline_version": 2`.
2. If no mode is given, run `incremental`.
3. In `incremental` mode, if `last_generated_sha` is missing or is not a commit in this repo (`git rev-parse --verify <sha>^{commit}` fails, for example after a history rewrite), there is no diff base: escalate to `audit` and say so in your summary. Audit verifies every page against the current code, which covers whatever the lost diff would have shown.
4. In `incremental` mode, if the diff since the last run touches more than ~40% of tracked source files, escalate to `audit` and say so in your summary.

Runs are time-bounded and the caller may cancel one. Write `.wiki-state.json` last, so a cut-off run leaves the previous SHA in place and the next run re-processes the same diff.

Do every step in your own turn. Never hand page writing to subagents or background tasks: a headless run ends when your turn ends, so work still running elsewhere is lost while the run reports success. Run shell commands from the repository root with repo-relative paths, never prefixed with `cd` or `git -C`: tool allowlists match command prefixes, so `rm docs/wiki/<path>` is permitted where the same removal by absolute path is refused.

## Layout

The wiki is two levels deep: root and sections.

| Path | Role |
|---|---|
| `docs/wiki/README.md` | Index. The only page every reader starts from. |
| `docs/wiki/<page>.md` | Root page. Spans more than one section. Always present: `architecture-overview.md`, `getting-started.md`. |
| `docs/wiki/<section>/README.md` | Hub. The landing page for one seam. GitHub renders it when a reader browses into the directory. |
| `docs/wiki/<section>/<page>.md` | Leaf. One seam within the section. |
| `docs/wiki/.outline.json` | The page map: the contract that makes incremental updates possible. |
| `docs/wiki/.wiki-state.json` | Run state. |

A **seam** is a boundary the code already has: a package, a service, a subsystem, a pipeline stage, a bounded set of modules that change together. Pages follow seams, never file types or a generic template.

Sections are earned. A section exists when one seam yields two or more leaves. A repo whose whole outline is three to six leaves has no section directories: the index is the only hub and the leaves sit beside it at the root. A seam that fits on one page gets `<section>/README.md` alone, hub and leaf in one file. There is never a third level: if a section wants one, raise the abstraction of its leaves instead. Decide all of this from `git ls-files` and the code, never from a fixed list of directory names.

## `.outline.json`

```json
{
 "outline_version": 2,
 "root_pages": [
 {
 "file": "architecture-overview.md",
 "title": "Architecture Overview",
 "covers": ["README.md", "AGENTS.md"],
 "seeds": ["README.md"]
 }
 ],
 "sections": [
 {
 "dir": "api",
 "title": "API Service",
 "covers": ["api/**"],
 "pages": [
 {
 "file": "api/auth-and-sessions.md",
 "title": "Auth and Sessions",
 "covers": ["api/src/auth/**", "api/src/middleware/session.*"],
 "seeds": ["api/src/auth/service.*"]
 }
 ]
 }
 ]
}
```

- `covers` on a leaf or root page is the set of paths/globs whose changes make that page suspect.
- `covers` on a section is the fallback for the whole seam: any file in the section's tree that no leaf claims (configs, lockfiles, READMEs, one-file directories) belongs to the hub.
- `seeds` are the 1 to 5 files to start reading from when writing the page.
- A section with a single page has `pages: []`; its `README.md` is written from the section's own `covers` and `seeds` (add `seeds` on the section in that case).

**Ownership rule:** every non-excluded tracked path resolves to exactly one page. Resolution is most-specific-glob-wins (the longest matching pattern). Leaf `covers` inside a section must not overlap each other; a path matching two leaves is a defect to fix in the outline, and the self-check reports it.

**Order of writes:** a page is written to disk before its entry is added to `.outline.json`, and a page dropped from the outline is deleted from disk in the same step. A cut-off run must never leave an outline entry without its page, or a page without its entry.

Two shapes the outline takes. A small single-service repo:

```
docs/wiki/
 README.md
 architecture-overview.md
 getting-started.md
 request-pipeline.md
 persistence.md
 background-jobs.md
```

A monorepo with two packages, one of which fits on a page:

```
docs/wiki/
 README.md
 architecture-overview.md
 getting-started.md
 web/
 README.md hub: directory map, diagram, page table
 routing-and-shell.md
 auth.md
 rendering.md
 tooling-and-tests.md
 worker/README.md hub and leaf in one file
 operations/
 README.md
 ci-workflows.md
 agent-configuration.md
```

`.wiki-state.json`:

```json
{"last_generated_sha": "<full sha>", "generated_at": "<ISO 8601 UTC>", "mode": "incremental"}
```

Get the SHA with `git rev-parse HEAD`. Rewrite this file at the end of every successful run; audit and full runs rewrite it even when no page changed. The caller reads its SHA and `generated_at` to tell a finished run from a cut-off one.

## Page sizing and splitting

A page is one seam, sized so a reader finishes it in one sitting and an agent can load it whole:

- A leaf covers typically 3 to 15 project-authored files and runs 300 to 1,500 words. Over 1,500 words is a split candidate; over 2,500 words splits in the same run, whatever the mode, with the new leaf added to the outline and the hub.
- A leaf has at most 7 H2 sections, and its H2s share one concern. Concerns are stack-neutral: bootstrap and configuration; auth and access control; request handling and routing; domain logic; data model and persistence; UI and rendering; external integrations; safety and compliance; tooling and tests. A page whose H2s straddle two concerns splits along that line. Read the concerns off the code; the list above is a vocabulary, not a template.
- A paragraph is at most ~120 words. Inventories (test files, routes, config keys, commands, environment variables) are tables.
- A section with more than ~8 leaves splits into two sections. A one-file directory is a row in the hub's directory map, never a page.
- Hubs run under ~600 words. The index runs under ~800.
- Vendored and generated directories get one row in the hub's directory map and one sentence on how they are produced, plus the local maintenance policy when the repo documents one (in its agent rules or README). Read that policy; never assume one.

## Page templates

Every page starts with this banner, values filled in:

```markdown
> Auto-generated by the wiki skill from commit `<short sha>` on <YYYY-MM-DD>. Do not edit by hand; changes will be overwritten.
```

**Index** (`docs/wiki/README.md`), in order:

1. Banner, H1, overview: what the project is and how it runs, two paragraphs at most.
2. One table per section (and one for root pages) with columns `Page | Summary | Key paths`. Summary is one sentence; key paths are the two or three directories the page is about.
3. `## Reading order`: a numbered list of 4 to 6 pages for a newcomer.
4. `## For agents`: two sentences stating that the lookup path is index, hub, leaf, and that `.outline.json` maps source paths to pages.

**Hub** (`docs/wiki/<section>/README.md`), in order:

1. Banner, H1, orientation: what the seam is and where it lives, one paragraph.
2. `## Directory map`: a table `Path | What lives there | Page` with one row per top-level subdirectory and config file of the seam. Vendored and generated directories are rows too, marked as such. The Page column links the leaf that owns the row, or says "this page".
3. One Mermaid diagram of the seam: module dependencies or the main flow through it.
4. `## Pages`: a table `Page | Summary`.
5. `## Cross-cutting`: links to the root pages and other sections this seam touches.
6. A final line linking back to the index: `Back to [the index](../README.md).`

**Leaf** (`docs/wiki/<section>/<page>.md`, and root pages), in order:

1. Banner, H1, orientation: one paragraph on what this seam does and where its code lives.
2. `## Key files`: a table `Path | Role` of the 3 to 10 files that matter most, each path a relative link into the repo (one `../` per directory level between the page and the repo root, so from a section directory it is `../../path/to/file`).
3. H2 sections describing purpose, structure, key flows, and interactions. A Mermaid diagram wherever the page describes a flow, a state machine, or a handoff between components; a sentence wherever a sentence is enough.
4. `## Related`: links to the hub, the sibling pages this page references, and the pages in other sections it depends on. Root pages link the index here instead of a hub.

## Page conventions

- Reference code by path and symbol (`src/auth/service.py`, `SessionStore.refresh()`), never with large pasted code blocks. Snippets over ~10 lines defeat the purpose; the reader has the repo.
- Link pages with relative links (`[Auth](auth.md)`, `[Worker](../worker/README.md)`), anchors allowed (`auth.md#session-refresh`).
- Describe what the code does today, in the present tense, as if the page were written fresh this run. Roadmaps, intent, and history belong in commits and tickets. If behavior looks like a bug, describe the behavior, not your guess about what was meant.
- **Single source.** A fact lives on exactly one page; other pages link to it. When two pages both need a fact, the owner is the page whose `covers` includes the file that defines it.
- Never copy values of secrets, tokens, API keys, connection strings, or `.env` contents into a page, even values found committed in the repo. Name the config key, never the value.

Writing style: plain, specific, low ceremony. Concretely:

- Punctuate with commas, colons, semicolons, periods, and parentheses. Em dashes are banned; the self-check greps for them.
- Say what the thing does: "X does Y" over "X serves as / is responsible for Y". Show that something is simple rather than asserting it.
- Plain vocabulary: delve, leverage, robust, seamless, streamline, comprehensive, and "plays a crucial role" are banned, as are bold-label bullets (`**Performance**: ...`) and negative parallelism ("it's not X, it's Y").

## What to exclude

Skip vendored dependencies, generated code, lockfiles, build output, fixtures/snapshots, minified bundles, and `docs/wiki/` itself. Use `git ls-files` as the source of truth for what is tracked, then apply judgment: if a directory is clearly machine-written (codegen output, migrations dumps, registry-pulled components), document that it exists and how it is produced, not its contents.

## Full mode

1. Inventory: `git ls-files`, apply exclusions, read the README and obvious entrypoints (main modules, app factories, CLI definitions, CI config) to understand what the project is.
2. Write `.outline.json`: find the seams, decide which earn sections, assign `covers` so the ownership rule holds, pick `seeds`.
3. Write each leaf and single-page section: start from its seeds, then explore with Read/Grep/Glob (follow imports, find callers, check config defaults) until you can describe the seam's purpose, structure, key flows, and interactions. Verify claims against code you actually read. Apply the sizing rules as you go; split before writing rather than after.
4. Write each hub from its finished leaves, then the root pages, then the index last, so each reflects the pages that exist.
5. Remove any `docs/wiki/**/*.md` not in the new outline (`rm docs/wiki/<path>`); an emptied directory may stay. Run the self-check. Write `.wiki-state.json`.

## Incremental mode

1. Read `.wiki-state.json` and `.outline.json`. Compute changes: `git diff --name-status <last_generated_sha>..HEAD -- . ':(exclude)docs/wiki'`. If the diff is empty, update nothing (you may still rewrite `.wiki-state.json`) and report a no-op. Commits after `last_generated_sha` that touch only `docs/wiki/` are not evidence that their diff was processed: a cut-off run commits partial pages without advancing the state file. The state file is the only record of what was processed; never narrow the diff by reasoning about those commits.
2. Resolve each changed path to its owning page via the ownership rule. A path that resolves to no page means the outline has a gap: assign it to the best-fitting existing page (extend its `covers`) or, if it belongs to a genuinely new seam, add a page or section.
3. For each affected page: read the diff for its files (`git diff <last_sha>..HEAD -- <paths>`) to learn where to look, then rewrite the whole page from the current code, using the old page only as a checklist of topics to re-verify. A page is a fresh description, never an edit log. Renames and deletions are reflected; a page whose entire subject was deleted is removed from disk and from the outline. Apply the sizing rules: a page that grew past its bounds splits now.
4. Rewrite the hub of every section whose leaves changed (directory map and page table included). Rewrite the index tables whenever any page was written, added, removed, or retitled. Both are short; keeping them current on every run is what stops the index going stale.
5. Run the self-check. Write `.wiki-state.json`.

## Audit mode

The weekly safety net. Incremental runs can each correctly conclude "no page change needed" while ten of them together leave a page wrong. Audit exists to catch that drift plus structural rot.

1. Re-derive an outline from the current tree as if running full mode, but don't write it yet. Compare with `.outline.json`: seams that grew enough to deserve their own page or section, pages whose subject shrank or vanished, tracked paths no page owns, leaves whose `covers` overlap, pages outside the sizing bounds. Apply the structural fixes (add/split/merge/remove pages; new or split pages are written as in full mode). Structural decisions are sticky: a split or merge made by an earlier run stands unless the code moved or a hard bound is exceeded, so audits do not oscillate on word counts alone.
2. For every surviving page, oldest banner SHA first: confirm each path in `covers` still exists, then verify the page's main claims against the current code (entry points it names, flows it describes, config keys it references). Rewrite what drifted. A page that passes keeps its banner untouched, so a clean audit produces no churn; say which pages passed.
3. Rewrite hubs and the index if the page set changed. Run the self-check. Write `.outline.json` and `.wiki-state.json`.

## Self-check

Run before writing `.wiki-state.json` in every mode. Fix what fails; anything you cannot fix goes in the summary.

| Check | How |
|---|---|
| Links resolve | For every `](<path>.md` target, Glob the file relative to the linking page; for anchors, confirm a heading in the target produces that slug. |
| Navigation is two-way | Every leaf is in its hub's page table and links the hub in `## Related`; every hub and root page is in the index; every hub links the index. |
| Present tense | Grep pages for `-`, for `\b(now|recently|no longer|previously|used to)\b`, for ticket and PR references (`\b[A-Z]{2,}-\d+\b`, `#\d+\b`), and for bold-label bullets (`^- \*\*[^*]+\*\*:`). Rewrite each hit. |
| Sizes | `wc -w` on every page against the bounds in Page sizing. |
| Ownership | Every tracked, non-excluded path resolves to exactly one page. Enumerate with `git ls-files <glob>` per pattern; list gaps and overlaps. |
| Outline matches disk | Every page in `.outline.json` exists on disk, and every `docs/wiki/**/*.md` (index and hubs included) is in the outline. Delete strays with `rm docs/wiki/<path>`; never leave a redirect stub in place of a deletion. |
| Banners | Every page written this run carries `git rev-parse --short HEAD` and today's date. |

## Reporting

End every run with a short summary: mode actually run (and why, if escalated); the page tree with word counts; pages created, updated, deleted, and verified unchanged; self-check results (gaps, overlaps, unresolvable links); whether `.wiki-state.json` was written; and anything that needs a human (huge uncovered directory, suspected bug, unparseable state). Keep it to a screen; it lands in a CI log.

wiki.yml

# Generated codebase wiki (docs/wiki/), maintained by the /wiki skill.
#
# Runs anthropics/claude-code-action. The action doesn't accept push events,
# so a push to main re-enters this workflow as a workflow_dispatch (one of
# the two event types the default GITHUB_TOKEN may still trigger), and
# allowed_bots lets that GITHUB_TOKEN-dispatched run pass the action's
# human-actor check. Full rationale and the failure history:
# docs/agents/research/claude-automation-on-push.md
#
# Triggers:
# - push to main, re-dispatched (incremental: diff since last wiki commit)
# - Mondays 03:23 UTC (audit: re-derive outline, catch drift)
# - manual dispatch with a mode override
#
# Loop safety (both hold, either alone is enough):
# - paths-ignore: pushes that only touch docs/wiki/ don't trigger this
# - the commit step pushes with the default GITHUB_TOKEN, and GitHub never
# creates push-triggered runs for events caused by that token
#
# Setup per repo:
# 1. Copy .claude/skills/wiki/ and this file into the repo, and make sure
# docs/wiki/ is not git-ignored (this repo ignores docs/* with exceptions).
# The skill writes a two-level tree (index, section hubs, leaf pages) and
# escalates to a full run on its own when docs/wiki/.outline.json is
# missing or predates the current outline format.
# 2. Set the LAZER_PROXY_API_KEY secret and the LAZER_PROXY_BASE_URL Actions
# variable (org-level recommended so repos don't each need a copy).
# Optional model overrides: LAZER_PROXY_WIKI_MODEL for incremental runs
# (cheap, every push) and LAZER_PROXY_WIKI_FULL_MODEL for full and audit
# runs (structure and verification decisions, worth a stronger model).
# LAZER_PROXY_WIKI_FULL_EFFORT (low, medium, high, xhigh, max) sets the
# effort level for full and audit runs; unset leaves the model default.
# Installing the Claude GitHub App is optional: without it the action
# falls back to the job's default GITHUB_TOKEN.
# 3. If main is a protected branch, allow this workflow to push (GitHub App
# or Actions bypass in the ruleset) or switch the commit step to a PR.

name: Wiki

on:
 push:
 branches: [main]
 paths-ignore:
 - "docs/wiki/**"
 schedule:
 - cron: "23 3 * * 1" # Mondays 03:23 UTC; odd minute to dodge the top-of-hour delay
 workflow_dispatch:
 inputs:
 mode:
 description: Generation mode
 type: choice
 options: [incremental, audit, full]
 default: incremental

jobs:
 # claude-code-action validates the event type and fails on push
 # ("Unsupported event type: push"), so a push re-enters this workflow as a
 # workflow_dispatch. Dispatching with the default GITHUB_TOKEN works:
 # workflow_dispatch and repository_dispatch are the two events exempt from
 # GitHub's no-retrigger rule for that token.
 dispatch:
 if: github.event_name == 'push'
 runs-on: ubuntu-latest
 permissions:
 actions: write
 steps:
 - run: gh workflow run wiki.yml --ref main -f mode=incremental
 env:
 GH_TOKEN: ${{ github.token }}
 GH_REPO: ${{ github.repository }}

 wiki:
 if: github.event_name != 'push'
 runs-on: ubuntu-latest
 # A newer run cancels an in-progress one, so the wiki is always generated
 # from the latest main. Cancelling mid-run is safe: the skill writes
 # .wiki-state.json last, so the replacement run re-processes the same
 # diff. Job-level (not workflow-level) so the seconds-long dispatch job
 # doesn't churn the group.
 concurrency:
 group: wiki-generate
 cancel-in-progress: true
 timeout-minutes: 85
 permissions:
 contents: write
 id-token: write
 env:
 WIKI_MODE: ${{ inputs.mode || (github.event_name == 'schedule' && 'audit') || 'incremental' }}
 steps:
 - uses: actions/checkout@v7.0.1
 with:
 # Full history: incremental mode diffs against the SHA recorded
 # in docs/wiki/.wiki-state.json, which can be arbitrarily old.
 fetch-depth: 0

 # The commit step compares this against the state file's generated_at
 # to tell a state file this run wrote from one left over from before.
 - id: start
 run: echo "at=$(date -u +%Y-%m-%dT%H:%M:%SZ)" >> "$GITHUB_OUTPUT"

 - uses: anthropics/claude-code-action@v1.0.235
 # Below the job's 85-minute limit so the commit step (if: !cancelled())
 # still has time to push whatever was generated when a run hits this
 # ceiling. An audit of the current tree needs more than 50 minutes.
 # Committing a truncated run is safe: the skill writes
 # .wiki-state.json last, so a killed run leaves the previous SHA in
 # place and the next run re-processes the same diff.
 timeout-minutes: 75
 with:
 anthropic_api_key: ${{ secrets.LAZER_PROXY_API_KEY }}
 # Dispatched runs are initiated by GITHUB_TOKEN, which the action's
 # human-actor check sees as the github-actions bot.
 allowed_bots: github-actions
 # Stream Claude's full transcript into the run log. By default the
 # action prints only the init and final-result messages, so a
 # multi-minute generation looks stalled. The prompt is a constant
 # string against our own repo, and the repo is private, so there is
 # no untrusted output to hide.
 show_full_output: true
 prompt: "/wiki ${{ env.WIKI_MODE }}"
 # Incremental runs happen on every push and only touch the pages a
 # diff points at; full and audit runs decide the page structure and
 # verify every page, so they get the stronger model.
 #
 # Agent is disallowed because subagents launch asynchronously and a
 # headless run ends with the main turn: the model would hand page
 # rewrites to subagents, end its turn to wait for them, and the run
 # would exit "success" having committed almost nothing. Eleven runs
 # did exactly that on 2026-09-15.
 #
 # Bash is allowed outright. Without an allow rule the action runs in
 # default permission mode, where a headless session refuses any
 # command it cannot statically clear: git diffs prefixed with cd or
 # -C, python and node helpers, pipelines, and the rm that removes a
 # page dropped from the outline (an audit on 2026-09-16 lost 19
 # commands this way and could not delete a split page). The prompt
 # is a constant, the repo is private, and the job token can only
 # write repo contents, so there is nothing for the sandbox to guard.
 claude_args: |
 --model ${{ env.WIKI_MODE == 'incremental' && (vars.LAZER_PROXY_WIKI_MODEL || 'claude-sonnet-5') || (vars.LAZER_PROXY_WIKI_FULL_MODEL || 'claude-opus-5') }}
 --allowedTools "Read,Write,Edit,Glob,Grep,Bash"
 --disallowedTools Agent
 ${{ env.WIKI_MODE != 'incremental' && vars.LAZER_PROXY_WIKI_FULL_EFFORT && format('--effort {0}', vars.LAZER_PROXY_WIKI_FULL_EFFORT) || '' }}
 env:
 # Routes all inference through Lazer Proxy (org-level Actions variable).
 ANTHROPIC_BASE_URL: ${{ vars.LAZER_PROXY_BASE_URL }}

 - name: Commit wiki updates
 # Run even when the Claude step fails or times out, so generated
 # changes aren't dropped; the job still reports the step failure.
 if: ${{ !cancelled() }}
 env:
 RUN_STARTED_AT: ${{ steps.start.outputs.at }}
 run: |
 # The skill writes .wiki-state.json last, so a state file this run
 # did not write means the run was cut off before it finished
 # (timeout, API error, or the model ending its turn early). The
 # only legitimate skip is an incremental run whose diff was empty.
 # The partial pages are still committed below, because the next
 # run re-processes the same diff, but the job fails so the gap is
 # visible instead of buried in a green run.
 head="$(git rev-parse HEAD)"
 previous="$(git show HEAD:docs/wiki/.wiki-state.json 2>/dev/null | jq -r '.last_generated_sha // empty')"
 recorded="$(jq -r '.last_generated_sha // empty' docs/wiki/.wiki-state.json 2>/dev/null || true)"
 generated="$(jq -r '.generated_at // empty' docs/wiki/.wiki-state.json 2>/dev/null || true)"
 state_written=true
 if [ "$recorded" != "$head" ] || [ -z "$generated" ] || [[ "$generated" < "$RUN_STARTED_AT" ]]; then
 state_written=false
 fi
 run_incomplete=false
 if [ "$state_written" = false ]; then
 if [ "$WIKI_MODE" != incremental ] || [ -z "$previous" ] \
 || ! git rev-parse --verify --quiet "${previous}^{commit}" >/dev/null \
 || ! git diff --quiet "$previous" HEAD -- . ':(exclude)docs/wiki'; then
 run_incomplete=true
 echo "::error::docs/wiki/.wiki-state.json records '${recorded:-nothing}' at '${generated:-no time}' but this ${WIKI_MODE} run started at ${RUN_STARTED_AT} from ${head}; the wiki run did not finish." >&2
 fi
 fi
 if [ -n "$(git status --porcelain docs/wiki)" ]; then
 # Author only; the push still uses GITHUB_TOKEN, which is what
 # the loop-safety rule above depends on.
 git config user.name "claude[bot]"
 git config user.email "209825114+claude[bot]@users.noreply.github.com"
 git add docs/wiki
 git commit -m "docs(wiki): update generated wiki [skip ci]"
 # Main may have moved during the (up to 75-minute) Claude run.
 # Rebase onto the latest main and retry; our commit only touches
 # docs/wiki, so conflicts are only possible against another wiki
 # commit, which the concurrency group already serializes.
 pushed=false
 for attempt in 1 2 3; do
 if git pull --rebase origin main && git push origin main; then
 pushed=true
 break
 fi
 git rebase --abort 2>/dev/null || true
 sleep 10
 done
 if [ "$pushed" != true ]; then
 echo "Failed to push wiki updates after 3 attempts." >&2
 exit 1
 fi
 else
 echo "No wiki changes."
 fi
 if [ "$run_incomplete" = true ]; then
 exit 1
 fi

Was it worth it?

Yes. This started as an itch: I liked someone else's generated wikis and I wanted one that was mine. Now it's the first thing I send to new people on the project, and the thing every agent in the repo reads before touching anything. It cost me one weird day of GitHub Actions trigger archaeology, eleven green runs that lied to me, and a few dollars per push.

I still need to fix the walls of text. When I do, I'll update the skill in this post.

If you set this up in your own repo, let me know how it goes! I'm really curious to see what seams it finds in codebases that aren't mine.

See you in the next one!

30 Sep 2026 5:00am GMT

28 Sep 2026

feedPlanet Twisted

Glyph Lefkowitz: What Would A Serious AI Product Look Like?

One of the issues that I have with the current generation of "AI" products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like.

Even before we get to the tremendous ethical problems with the frontier labs, it is this impression of their composition as a product that makes me feel, constantly, whenever I am interacting with them, that they are less a software product than that they are a grift, a scam designed to make me feel like I am interacting with a product that has capabilities that it simply does not, to try to lull me into a false sense of security that I can trust it.

The frontier labs are of course the worst offenders, but every criticism here applies just as much to Ollama, which (if anything, due to the obviously poorer quality of the available models themselves) needs these features even more than the frontier labs do.

Here, I will set down a few features that might convince me that an LLM-based product, particularly one focused on research or software development, was actually serious about helping me do useful things with it.

Make "Checking For Mistakes" A First-Class Feature

This is the biggest issue, and the major reason that I was inspired to write this post.

It is a truth universally acknowledged, that AIs cannot reliably provide information.

I could cite a ton of news articles and studies about this fact, but there is no need. Every single chatbot admits this, up front, in a fine-print disclaimer as a core part of their user interface. Gemini says "AI can make mistakes, so double-check responses", Claude says "Claude is AI and can make mistakes. Please double-check responses.1" ChatGPT says "ChatGPT can make mistakes. Check important info.".

Every time I see that last one, I wonder how I'm supposed to know what "info" is supposed to be "important".

All of these warnings are all small, gray text, painfully obviously included as legalese to push responsibility back onto the user rather than to help with anything. This is a core limitation of all these products. Checking their output is a part of the workflow for using them that:

  1. you absolutely cannot skip or skimp on without creating risks to yourself and whoever you are conveying its output to, and,
  2. it is very easy to skip or skimp on and you are encouraged at every turn to do so, because "just trust the output" is one of the quickest ways to save time.

A chatbot product that took this weakness seriously, as an actual consideration for using it, would put a checkbox next to every claim in its output. It would be a 2-column worksheet, where you've got the LLM output in the first column, and next to it, human notes in the second column, explaining what work went into checking this claim, and a big checkbox that you would only check off after you believe you'd checked its claims thoroughly enough.

Coding assistants would need to have some version of this as well. Right now, this is pushed off into code review, which means it is a dark pattern which subtly encourages the "author"2 to offload this work to their code reviewer without ever looking. Once again, "it's probably fine, I don't need to check" is the quickest way to save time and churn out those PRs faster.

It might even be useful for coding harnesses to have some affordance for checking code before it even runs tests. As the vendors themselves have admitted, it's not just expensive to burn tokens on your "AI", you also end up burning far more compute on the AI. Being able to check your diffs before sending them over to uselessly exhaust your testing compute cluster would be useful.

If your product tells me that it makes mistakes and I must be the one to check for the mistakes, but then gives me zero tools to check for mistakes, I cannot take it seriously.

More Citations to Check, And More Details

Most chatbots prefer to give an answer, rather than a citation. In my own personal use, I find that when asked to provide a list of citations with clearly marked sources for each one, they will appear to "get bored" halfway through the list and simply stop including citations at some point.

When the bots include citations at all, present them as inline annotations that say nothing but the domain name of the search result, in a font so small that it's barely legible, and an equally indecipherable icon that is fewer than 16 pixels on a side.

This is backwards.

Now, I am aware that these citations do come from somewhere, and in an attempt to reduce hallucinations, all of the major providers support some form of "grounding"3, and that those little barely-readable citation links are referencing actual structures in the RAG pipeline and not just potentially-hallucinated tokens, but I'm not talking about the underlying machinery in the model, I'm talking about the presentation to the user.

Plus, regardless of whether a snippet of text came from a RAG query, we know that LLMs can never provide an authoritative result; it's a fundamental limitation of the technology. They can still garble the results of RAG as much as they can misrepresent any other training data. This means that it must never present its results as authoritative.

If you ask an AI to do research queries, every result should be presented as a list of citations. Moreover, the presentation should display each citation as a large object of in its own right, with clearly identified metadata, including not just the site where it was found but its publication date and, if possible, the name of the author. The literal, unmodified quotation (not from RAG, not a summary: a quotation extracted with a regular program and not an LLM) should be front-and-center, larger than any AI-generated text.

If the AI product wants to editorialize or summarize (which should not always be necessary!), the AI-generated text should be presented as small text underneath the citation that has been found, de-emphasized as much as the disclaimer is right now, at the very least until the user has verified that the summary is accurate. Perhaps, for a research project, a "did you read the citation" checkbox might even be helpful.

If your product openly tells me that it will scramble, misrepresent, or omit its citations in its summaries, and I must read the original human-authored citations to be sure, but then gives me no tools to track my reading of those citations or even any way to find them, I cannot take it seriously.

No First-Person Output, No Apologies

There is no reason for a software development or research tool to use first-person language to describe itself. They should not do so. In fact they should not be allowed to do so.

There is also no reason that they should ever apologize. It is a waste of everyone's time; it's a waste for the chatbot to generate the apology, it's a waste for the user to read the apology, and it's a waste for the user to respond to the apology. Yet they unfailingly do this upon every correction.

The vendors of these tools know that they are routinely causing mental-health crises. In response, they have added non-functional "guard rails" that can still, in 2026, easily be bypassed.4

A product seriously interested in helping with productivity would correct this glaringly obvious flaw, focus on the task at hand, and stop emitting useless verbiage.

In the previous two sections, I tried to focus on ways in which the harness would be constructed differently even if the LLM technology is fundamentally impossible to improve; in this case, I have to assume that the labs have some control over the model itself. But unless they are truly incapable of influencing their output (and all their "benchmarks" and "capabilities" seem to indicate that they can control it very tightly) they ought to be building models that are much less verbose.

More Non-Natural-Language User Interfaces

Although natural language could hypothetically be a powerful interface for interacting with a computer system, the practical upshot of LLM natural language interfaces is that these interfaces are imprecise and repetitive, full of superstitions masquerading as "best practices". The inputs are a mess and the resulting outputs are a mess.

The general way of addressing this unstructured mess is to allow the chatbot to directly take action in response to the user's input; in other words to supply it with "tools" via an MCP server. But again, this is backwards. If we cannot even express our intent clearly in the first place, why are we trusting this system to take potentially destructive and harmful actions on our behalf?

Instead, I would expect a product that was seriously invested in helping me accomplish specific tasks, to have user interfaces specific to those tasks. Is it supposed to be able to be a security scanner that can discover OWASP top 10 bugs in a codebase? Have a button for that. Build that functionality into your harness, train it directly into the model, use smaller models that can satisfy that functionality more effectively than throwing it at the planet-sized brain of Fable or whatever.

I'm aware that there are small software startups that do something like this, but they are bolted on to the side of the main model providers' APIs, not integrated into the core of the product and not using their own models and AI systems to achieve consistent and repeatable results.

Strong Data Provenance Indicators

Chatbots produce data tables pulled from websites, from APIs, from MCP tools or from summarizing and scrambling the user's input. In order to provide the illusion of a seamless interface, this data is presented in-line regardless of where it comes from. But some of these outputs are produced mechanically via regular old API calls, for example, from the result of calling a tool or querying a website, but presented uniformly.

But there is a huge difference between an authoritative data source being inlined as part of a chatbot conversation, being treated as input by the chatbot, and some ad-hoc hallucinated data being treated as output of the chatbot.

If a product is trying to help me make accurate, empirically-grounded, data-driven decisions, the source of the data is critical.

Integrated into the "check for mistakes" and "verify citations" workflow I described above, there's a necessary "verify data programmatically" pass as well; to have tools that will treat portions of the output as a regular spreadsheet, allowing regular-old computer arithmetic to verify things and showing where such arithmetic was used, and how.

Better User Control of Reproducibility

Anyone familiar with the technical specifics of LLMs will know that they have a variable called "temperature" which controls the degree of randomness that the LLM uses to produce its outputs. But most users don't know this, because it isn't exposed as part of the user interface by default.

This leads to a subjective impression that you asked ChatGPT, and you got ChatGPT's authoritative answer.

You can't just set the temperature to zero and still get useful results - I am aware that it does more than just scramble the output at random, and there are perhaps good reasons that simply exposing just a temperature setting would not be that useful to users. But if we followed some more of my earlier recommendations for making more structured UI elements to solve specific problems rather than having long back-and-forth chats where each refinement depends on the previous response, perhaps those elements could also re-play the process so that users can see how reliable the bot is at a particular task and develop a sense of how the stochastic nature of the process actually affects it.

Similarly, if a user is trying to solve the same problem repeatedly with a chatbot, and the chatbot product has numerous computational tools that don't really have anything to do with the LLM, such as deterministic data-processing tools, then having a way to freeze the non-deterministic parts of the transcript but re-populate a particular data frame with updated information and fork / continue the conversation from there would be a way to avoid introducing pointless additional randomness when you already know what tool you're trying to use.

The fact that every conversation is presented as this flat chat prompt that doesn't let me interact with any of the widgets that were previously produced except through more chatting, really makes me feel like the whole product is just doing predatory social-media style "increase time on site" optimization, just trying to lure me into further repetitive and unreliable chats, rather than letting me get in, solve my problem, and get out.

Context Visibility

Managing the LLM context is the ongoing challenge facing organizations that are trying to use "agentic" workflows. Filling up the context with too much information causes well-known problems. In response, advanced LLM users attempting to solve larger problems must break up very long prompts into "skills", give access to lengthy information via "tools", and delegating sub-problems to "sub-agents" rather than simply extending a single prompt indefinitely.

All of these strategies have flaws, because even on the largest models, compared to the breadth and depth of knowledge-work problems, LLM contexts are quite small.

And yet, none of these products will show the context to the user by default. There are third-party addons that can show you a simple progress bar but for addressing the premier engineering difficulty with this technology, that is below the bare minimum.

This lack of visibility means that almost all of the tools for extending the context are flying blind. Rather than responding meaningfully to a full context, everyone just kind of guesses how much state they need by guessing and trying over and over again with progressively more elaborate skill and sub-agent layouts. Even managing context compaction ends up being an advanced API-driven workflow5.

A serious product that was trying to help the user understand would not only show "available context" but explain the impact of context compactions, make it easier to see harness-generated prompts, and so on. This would be a first-class feature, combined with the aforementioned reproducibility / replay tools, would allow users to do real experiments to develop an understanding about how to make good use of the context window.

A Sandbox That Actually Works

I've been focused on the chatbot interface here because it is the most immediately egregious upon looking at the UI. But the "agentic loop" tools used for coding are equally dangerous, if not more so. Coding tools keep destroying everyone's data, over the course of years.

These catastrophic incidents that become front-page news are relatively rare compared to the amount of coding-agent use out there. But they also aren't the only kind of sandbox violation. Coding models will so routinely edit test code instead of the system under test that there are "pro tips" articles all over the web giving you the flawed advice to simply ask the agent not to cheat. News write-ups of the catastrophic incidents themselves will also offer glib and wrong advice, like "use a docker container". That might prevent it from literally deleting your operating system, but it won't prevent it from destroying all the local work you have in your codebase (it needs access to a checkout, after all!)

There is a flurry of activity in the infosec space where people are rushing to plug the gaps left by these coding harnesses. Everyone's got their own version of an MCP approval gateway where you can optionally place a proxy between your agent and your production infrastructure.

In the best case, though, all these mitigations and proxies and prompts simply turn the user into an auto-approval automaton, hitting Y, Y, Y, Y over and over again, until you finally are driven mad and hit "yes to all", turn on full-auto mode and submit yourself to the void. With nothing between your personal vigilance and disaster, there are no workflows left beyond decrementing your own vigilance until there's nothing left and then hoping the disaster never arrives.

The fact that some mitigations exist that can be deployed by extra-cautious users does not change the fact that "agentic coding" is an unsafe-by-default technology deployed without concern or guidance. Every frontier lab has tied a spring-loaded shotgun to a dog; the fact that dog owners can publish thoughtful blog posts explaining how you can teach your dog the basics of gun safety or how you can have your dogs play in a bullet-proof room does not mitigate the fact that the product should not have been allowed in the first place, nor should it continue to exist without VERY strong security controls.

I might believe that a frontier lab were seriously interested in providing developers with a useful tool if they shipped something that had safety built-in.

That means tools in the harness, detached from any LLM, independent of the prompt, that could:

In the same way that I suggested above that research-based tasks should have a way of re-issuing prompts to determine how reproducible a result is, or whether other sources might be found, agent-based tasks should have a way of being executed against mock services for popular APIs, so that the verification can match both on the front-end (review the plan for making the API calls before they're executed) and the back end (review the API calls that were issued to the mock service and verify that they matched).

Instead, the frontier labs provide us products that are disasters out of the box, give us "best practices" to build massive and elaborate, as well as incomplete and error-prone, security perimeters of our own design. Then they blame "operator error" when it inevitably goes wrong. I cannot believe that these design choices are intended to help us be productive.

Bonus: Human Processes

Organizations deploying AI also frequently come across as unserious, for similar reasons. In 2023, naive exuberance could perhaps be forgiven. But today, as we near the close of 2026, there are several well-known problems, that have been extremely well-covered in the press. None of these things should be surprising, but most orgs deploying these tools are still just letting them rip and hoping it all works out.

Organizations deploying these tools would need at least three kinds of major modifications to their internal processes, if they wanted to be serious about using them safely:

1. Shift Rotations to Prevent Vigilance Decrement

There have been several high-profile incidents where software developers' gradual acquiescence to accepting LLM output have lead to serious economic consequences for the companies deploying them, perhaps best typified by Amazon's "millions of lost orders" due to a gradual decay of their engineering processes from LLM use.

These outages, and other AI-related failures, are due to the difficulty of maintaining focus on the same problems. In other words, as I described above, vigilance decrement is a constant problem, because AI outputs are most often correct, but continue to be incorrect in surprising and non-intuitive ways. As I have previously written, you cannot trust yourself to catch every bug with code review, and LLM output.

Aviation, for example, has very strict rules around rest requirements. There is also a specific rule that "No certificate holder may operate an aircraft without a second in command if that aircraft has a passenger seating configuration, excluding any pilot seat, of ten seats or more.". Other safety-critical professions have similar rules.

And yet, even in the age of the supposed "AI revolution", most software teams are still assigning every engineer a full feature load, not planning for any rest, and telling people to review code whenever they happen to have some "free time".

Maintenance of vigilance has to be your top priority. Regular, scheduled, inviolable rest periods where people do work without AI assistance, and are not exposed to any AI output for review or otherwise, would be crucial in order to stay mentally sharp enough.

The tools themselves should have this sort of thing built in. The mistake-review process described above should have a periodic spot-check mode where a second reviewer periodically reviews a chatbot log, doing their own independent verification of claims, to see if they spot the same errors. This could provide a feedback loop to determine how much rest is necessary to maintain continuous attention and actually spot hallucinations.

2. Skill Practice To Prevent Skill Loss

It is also well-known that AI use leads to AI reliance, and AI reliance leads to skill loss.

I like to use the analogy to dockworkers at a seaport6 adopting automation.

If you employ dockworkers to load and unload ships all day long, they are going to be getting tons of exercise. They will be able to lift heavy objects on demand, whenever. They might have plenty of health problems and injuries from this type of work, but "lack of exercise" will not be a problem.

With the development of standardized container ships and mechanized cranes, you are going to be changing their job description substantially: now they mostly spend all day sitting in a small cubicle moving a control lever back and forth, not lifting heavy stuff. They will get worse at lifting heavy objects.

In this analogy however, the cranes are not all that reliable. We know they break, and they drop their payloads sometimes, and the stuff needs to be manually moved. But this only happens a few times a week, at most. If you need whoever is driving the crane to be able to jump out at any moment and still move stuff around manually, then you need to make an affordance for that. You need to give them time to go to the gym and do some lifting for practice, or every crane failure is going to be a major emergency.

An organization doing an AI transformation would also need a massive increase to learning & development budget, both in terms of resources and in terms of schedule. If your people are going to lose skills because they've lost regular practice in the incidental course of doing their duties, then they are going to need deliberate, intentional, non-incidental practice of those skills to keep them sharp.

But rather than trying to accommodate new workflows and give time for people to adjust, most AI mandates are simply dropped on workers like a ton of bricks, with no time to adapt and no affordance for maintaining their skills. Operate the crane and stay fit and healthy and ready to switch back to manual lifting at any time and then get back in the crane cockpit right afterwards. Don't mess up.

Then an accident happens and everyone is surprised, as if this process weren't practically designed to produce a terrible result.

3. Mental Health Resources to Deal with Mental Health Risks

AI psychosis often begins with practical problem-solving, and beyond that, it can start specifically at work. Not to mention the more pedestrian condition of "AI brain fry".

If you are mandating your employees to use a hazardous tool that may seriously and directly damage their mental health, you need trainings and resources. You need in-house therapists and you need to be making sure to check in with people actively to make sure that this is not happening.

Again, the tool itself ought to have some way of dealing with this. An occasional "take a break" popup is easily dismissed; they need a user-visible AI personal dosimeter so you can see your cumulative usage over time.

I don't even know if "usage over time" is a sufficient metric to gauge risk. Maybe if your work chatbot start to talk about resonance too much, unless you literally work as an acoustic engineer, that should be flagged for someone.

We are, again, years into dealing with these tools, and we know these risks exist. Yet no serious mitigations are provided. Not even any way of measuring the risk exposure.

And More

There are also many other risks associated with the technology. There are intellectual property risks with the foundation models, due to recklessness with their training data. There are existential financial risks associated with the infrastructure build-out. The extent to which most "open" models are simply derivatives of frontier models is an open question.

What I Think

If any one of these things were regularly overlooked by AI vendors or users, that would be a totally normal product oversight. Room for improvement for the next version, but nothing catastrophic.

Shipping without any of them doesn't seem like lean product management, it seems like a careless attitude towards risk and a product design philosophy oriented entirely towards short-term demos, with no regard for how to realize actual productivity gains.

Furthermore, being available for years without anything like these features, despite hundreds of incidents demonstrating the risks, with hundreds of billions of dollars of funding, makes it seem to me like if they were to add all the features that would make their product actually safe and hypothetically useful, these features would reveal that it is actually not an improvement to productivity.

In the year since I first wrote about measuring the cost/benefit ratio of AI, I have heard from numerous people who have shown this to management to try to illustrate why their AI initiatives - like almost all AI initiatives - were either failing or burning out their engineers.

I've also heard from lots of people that have told me that it's obviously useful and they don't need to measure so carefully, because they are getting lots of work done that they couldn't have otherwise.7

I have yet to hear from a single person who has said "yeah, we measured according to your methodology8, and it turns out that our AI work is going great and that our ratio is 0.75".

Obviously, I cannot say for sure why this is; absence of evidence is not evidence of absence. But at this point I think the null hypothesis is that AI tools provide, in aggregate, zero value. They make mistakes too often, and the externalities they produce are so bad and so difficult to control that even before we get to the places where they are just physically poisoning people, even the negative effects on their direct users end up cancelling out whatever benefit to they provide to their organizations.

If I were wrong, then including tools to measure an AI's effectiveness at the tasks their users are actually trying to accomplish, rather than meaningless benchmarks, would show big productivity gains. The frontier labs would be champing at the bit to add such features, and crowing about their fantastic results.

I think the labs know that if they did that, it would present a grim picture to their users. Such tools would let their users see that it's making mistakes much more often than they realized, that they're spending much more time with it than they want to be, and that it's just generally not fit for purpose.

If they prove me wrong by adding in all of these safety mechanisms, and in the process, they make all of their AI technology less harmful, I'll be thrilled to be debunked.


Acknowledgments

Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!


  1. It is also interesting that for the next section, sometimes it seems that Claude's disclaimer is "Please double-check cited sources." instead. ↩

  2. ... by which I mean the "prompter", since authorship is not what's happening here. ↩

  3. Claude has the "citations API", Google has various different kinds of "grounding" against its own APIs, and I guess Microsoft can check OpenAI's homework if you want. ↩

  4. Given the relatively slow speed of the justice system and the mainstream press around the world, we probably will not hear about whether people are managing to incidentally break through these guard rails to self harm right now, but there are no shortage of stories still being reported right now where people were still doing just that, such as in this story where the effect of the much vaunted "guard rails" in 2025 was that if you wanted it to write you a suicide note, it would refuse twice but acquiesce on the third try. I don't see any reason to believe this fundamental issue has been addressed in the meanwhile, since it had been happening for years at that point. ↩

  5. In this tutorial we can also see an incredibly rosy scenario presented, where a long-running workflow effortlessly compresses all of the necessary information into the new context, even if it uses a lower-fidelity model to do so, rather than the tangled and gnarly problem of problems which really are too big to fit in the context, which is to say, "most real-world problems". This presents the context limit instead as a minor speedbump to be worked around rather than the fundamental flaw in LLM tooling. ↩

  6. A heavily fictionalized seaport. This is not how actual dockworkers work. This is a simplistic metaphor about incidental benefits of instrumental tasks, it is not supposed to delve deeply into the mechanics of maritime shipping. In particular I know that cranes are more reliable than this and this is not actually how you would respond to a crane malfunction anyway. Feel free to share fun facts about maritime shipping if that is your special interest but please do not @ me to correct this metaphor. ↩

  7. To my knowledge, none of their publicly-traded employers have posted a measurable improvement to efficiency outside the margin of error. ↩

  8. Or any similar methodology. I don't need people to adopt the exact practice that I proposed there. ↩

28 Sep 2026 4:27am GMT

25 Sep 2026

feedPlanet Twisted

Glyph Lefkowitz: Who Is Open Source About?

Open source is, at least in part, about you, where "you" refers to the user.

Open Source Is Not About "Open Source Is Not About You"

In other words: Rich Hickey was wrong when he wrote "Open Source Is Not About You" and I'm tired of pretending otherwise.

Of course he's not completely wrong, or his famous post would not have resonated quite so much in the first place. Obnoxious users who demand their personal use-cases be immediately addressed by volunteer maintainers for free should indeed be viewed as the pariahs that they are. Similarly, corporate users who want free support from the community that supplies their infrastructure to lower their costs. As should those who profit from this type of externalization by their own customers.

But the exchange of "open source" (or even "free software") is not as simple as "I have prepared some software for you, please enjoy it, you have no right to complain", and maintainers ought to have a precise understanding of the costs and benefits - as well as the ethical implications - of that exchange.

Right now we barely even articulate that the exchange exists, let alone that it establishes a long-term, subtle, and implicit relationship between maintainer and user.

Let's fix that.

A Brief Aside about Meta-Ethics

When we talk about "obligations" and "rights", of "shoulds" and "musts", we are constructing an ethical system. The purpose of such a system is to develop social expectations and social consequences. There is not much use in me telling you that you are transcendentally evil for failing to follow some arbitrary recommendation that I have. But I am implying that I believe there should be consequences for your behavior. I am also implying that there probably already are some consequences, and they're just not written down anywhere yet.

Therefore, a post like this, where I say that we should view our social obligations in a certain way, that is the beginning of a broader social conversation. I think there should be some consequences, so I am gesturing towards that possibility. Exactly what consequences?

For now, I'm not sure. Let's figure it out.

What Are We Doing When We Do An Open Source?

Hickey, and his many acolytes in the years since his fateful post, asserts that the process of "open source" goes like this:

  1. Maintainer makes a thing, and makes it available to users as a gift.
    1. Maintainer may "love working with the team".
    2. Maintainer may be "proud of the work we do".
  2. Users accept the gift, and extract utility from it.
    1. (Users MUST be grateful for this.)
  3. A tiny fraction of users reciprocally contribute to the thing.
    1. (Maintainers may be grateful for this.)

He makes various oblique references to the specific activities of his company, which does things vaguely related to his projects for money1. These activities are exclusively characterized as for "customers", however, a subset of the aforementioned users so tiny ("fewer than 1%") as to nearly be an entirely distinct group.

Breezing past this process in an essay about obnoxious users demanding things they are not entitled to, one might nod along, as this sounds mostly sensible. Giving gifts is nice. I too love working with good teams and taking pride in things.

Examined more closely, however, it starts to logically fall apart. If you have consulting clients and that's where all of your money is coming from, why are you bothering (as he repeatedly insists) "doing [things] for the community"? What was the point of releasing this code in the first place? You could love working with your team and be proud of the work that you do in a lot of different contexts; why bother implicating this horde of entitled and obnoxious people, if that's all you're getting out of it? What's in it for you?

If we've left out something as fundamental as "why is the maintainer doing this", perhaps this story leaves out some other important bits as well.

Why Are You Doing This?

There are many possible motivations for releasing and maintaining open source software. They are often subtle, often overlapping, and rarely clearly stated. Maintainers are not a monolith and not everyone does it for similar reasons. But let's review a few reasons that someone might want to contribute.

Reputation

One reason that you might want to release some open source software is advertising. The most common form of this is self-promotion; if you are a visible, prominent contributor to an open source project, it stands to reason that you will have an easier time finding work in the domain of that project.

If you operate a consultancy, as Rich Hickey did at the time of his famous rant, then this reputational currency translates into advertising for your services. It's a practical demonstration of the skills of your team.

The trade in this benefit is most like the traditional "gift economy" that open source has been compared to. You give the code to your users, which has some value, but the users give you back some reputation, in the form of their attention, their esteem, and possibly even their money if they become customers or employers.

Influence

Infrastructure is the most popular type of open source for a good reason. Programmers working on a problem are often hemmed in by sclerotic architectural choices which prevent them from solving problems in the way that they'd prefer to solve them. Major infrastructural investments are difficult to justify in a planning process, as their benefits are hard to prove. Sometimes the benefits are highly personal; different engineers have different aesthetic preferences about what types of equally-valid solutions they'd prefer to work with.

If you can develop your preferred type of solution and release it as open source, then you can influence how everyone else solves this type of problem. As an individual, such a position of influence can allow you to have some transferable expertise between employers. You know how to use the tool you developed, so you can be very quick and effective with it, and you can shape it to your ongoing taste over time.

If you're an employer, and you can get everyone else to use your open source thing2, this can reduce both your hiring and training costs. Potential employees can read the code, see that it's good, and want to work at a place that produces good code like that. They can also read the code and become familiar with it in advance of coming to work for you, which means that you have a ready supply of developers who already know how your internal systems work.

The trade in this benefit is more like "soft power" than a gift economy. You give the code to your users, which has some value, but the users give you back the ability to dictate their technological agenda. You gain both the ability to influence their initial direction, and, as part of ongoing maintenance, to dictate their behavior over time.

Improvement

As an engineer, you might want to improve your own skills. Writing something proprietary and commercial cuts against this in two ways.

First, you will want to build something that already exists within your skill set, so that it will attract commercial interest and actually be competitive. Within the context of a larger team, you will want to personally be able to be immediately effective for similar reasons. But you still need a way to learn new things.

Second, you will want to build something somewhat secretively, so that the value you are producing is captured rather than released to the community. This means that you will be cut off from external sources of expert feedback.

As an organization, you might want to build the skills of your staff in similar ways.

The trade in this benefit is code for knowledge. You release the code or changes, and in return you expect your users to provide you good bug reports, and to induce at least some of them to become co-developers.

Outsourcing

As an engineer, you can only do so much on your own. Perhaps you want to have some influence over your infrastructure so you want to write it, but you also want to have a communal place to keep your infrastructure such that you can make a change to something to suit your needs, but you know that even if you walk away, someone else will maintain that change and keep it working across years or even decades of changes to underlying platforms, hardware, etc.

This sort of communal maintenance effort can be shared among all interested participants; if a thousand companies all need the same tool, if even a few dozen can share it, that reduces even their own load massively, let alone everyone else's.

The trade in this benefit is more complex, since there's less symmetry between the main maintainer and peripheral community members who also contribute code. The main maintainer is actually trading a namespace, a central place for people to contribute, coordinate, and release changes, rather than the code. They are a sort of market maker where then all the other contributors trade code for code within that market-ish structure.

In practice, this motivation produces a game theory problem where, when maintenance drops below a critical threshold, it creates a big enough crisis that at least some freeloading stakeholders will be forced to start making contributions.

Ultimately, however, this saves all involved parties a ton on maintenance, more eager volunteers who do not freeload in the first place get all the other benefits mentioned above as well.

A Brief Aside about your Chart of Accounts

Most companies account for open source maintenance work as simple overhead on ongoing projects. Sometimes it's CapEx, sometimes it's OpEx, but it's just "whoever happens to be working on this thing to support whatever random product it's a part of".

This type of accounting creates distorting incentives, because it doesn't recognize all the benefits above. Under such a fiscal regime, ongoing healthy maintenance becomes a ZIRP because when resources are more constrained, this apparent indulgence gets corrected.

The ancillary benefits that open source creates ought to be properly recognized. It shouldn't just be buried as Wages or IT or whatever. If it's helping you hire better engineers, some of that expense should be allocated to Recruitment Costs. If it's materially improving your reputation among your customer base, some of it should go to Goodwill. If it's getting your product in front of developers who are your customers, it should be in Marketing. Most importantly, if maintenance on an open source project is actually helping you maintain your enterprise-wide platform, it should not be squirreled away in some small team who happened to be the first one to adopt it.3

Exactly how these costs should be allocated and cross-charged to different departments depends heavily upon your organization and your specific chart of accounts. But "whatever, it's just part of the software product" or "I guess it's DevRel because the SDK is in there" is guaranteed to have your open source organization destroyed along with all those side-benefits the next time that there's a cash crunch.

The Things that Aren't Supposed To Be Benefits

These categories could be made as explicit, rational trade-offs, even if they are often implicit and subtle in practice. They are transactions where the maintainer gets something and the user gets something.

However, not everything that you are getting as a maintainer is something you are actually supposed to use to your own benefit. Being given trust in service of a responsibility is not a transaction.

"Oops, All Root Shells"

Open source code is code. In our modern world of absolutely pathetic sandboxing, installing code from somebody else gives them control over your system, even if it is somewhat indirect.

There is an unwritten rule that if I create an open source library, and you use it, it probably shouldn't have a backdoor in it that gives me the credentials to your bank account. There is a trust relationship between the user and the maintainer, and here, we see the first obligation that the maintainer has. The maintainer is obligated not to use the user's computer for their own gain.

This rule might seem obvious and straightforward. It might even seem unfair to you that I call the rule "unwritten", because the rule is, in fact, written down in a few places: for example, in the npm Acceptable Content Policy, it says right there:

A few examples of unacceptable content:

…

  1. Content containing malicious computer code, such as computer viruses, computer worms, rootkits, back doors, or spyware. This includes content submitted for research purposes. Tools designed and documented explicitly to assist in security research are acceptable, but exploits and malware that use the npm registry as a deployment or delivery vector are not.

I think we can all agree that a script which steals your bank credentials and sends them to me to buy a totally sick jet ski would qualify as "malware", so clearly that is forbidden.

There is also an enormous gray area here. npm also explicitly allows "Information on how to pay, donate to, and otherwise support Package development", but then goes on to explicitly forbid "Packages that display ads at runtime, on installation, or at other stages of the software development lifecycle, such as via npm scripts."4 How are the lines drawn around these gray areas? "npm will continue to apply its judgment when deciding what content is acceptable."

But also... this is forbidden by npm, not by the transcendental nature of "open source". I could give away code that displays all kinds of ads to its users as a "gift" on my website. The exact structure of this policy is not uncommon, but it also isn't exactly the same as other such sites. PyPI, for example, explicitly bans "cryptocurrency mining", which NPM does not. Is cryptocurrency mining "not open source"? A lot of judgement calls are happening here about what is allowable in these "gifts" that you are giving to your users.

But I digress.

My point is that policy-making around this concept is not clear, there are lots of little disagreements around the edges, but there is a very strong consensus that while the user is giving you their trust here, that is not a trade. The deal is not "you give the user some code, the user gives you unlimited compute and access to all their financial accounts". The user has made themselves vulnerable to your code on the strength of your reputation.

This creates an obligation for you to not do anything evil with that code, either intentionally or through negligence.

Security Updates Are Just Command And Control In A Funny Hat

All of this is just about the initial download of some code, and that is the way that Rich Hickey describes it, as if you just grabbed some code off a web page and put it in a folder that you like on your desktop. But that is not how open source relationships work today, if indeed it ever was.

The way it works today is that you add a dependency to your pyproject.toml or your package.json or your Cargo.toml and now your users are vulnerable not just to whatever you happened to upload in the first place, but to whoever happens to have your package index credentials.

This creates an obligation to maintain an operational security posture that protects your users from malicious updates.

The Roadmap Is Someone's Life

Another kind of trust that the user is placing in you is the trust that you are going to have at least some kind of regard for their usage of your software.

In a perfect world, the user's expectations could be clearly circumscribed. Whatever ongoing maintenance you commit to perform would be encapsulated in clear policies that you'd write up in advance, about exactly what kind of security response policy you have, how you will communicate when you no longer have the resources for maintenance, and so on.

But anyone who has been involved in any project at anything but the most extreme tier of operational maturity knows that 99% of the ecosystem relies on a set of loose conventions around how all that stuff works. We expect that maintainers will generally be around, that they'll use existing tools like an issue tracker for triaging user bugs, GHSA and CVEs for security reporting, that they will mark the project as "archived" and maybe do a final release before abandoning it, that they will maintain a ChangeLog explaining at least a little bit of what is going on.

Users assume that those conventions will be followed when there are any gaps in explicit policy, or indeed if policy is lacking entirely. This assumption is reasonable, because otherwise nobody could ever use any open source without a stack of service contracts that nobody has any time to write.

The strongest such convention is that an actively maintained program will, at least, more or less keep doing what it does as time goes on. A user who has elected to use a bit of open source software has made themselves vulnerable to changes and breakages in that software by the mere fact of using it. In the time that they have used it and invested in it, they have not invested in:

This can, and does, go badly wrong, when those expectations are mismatched.

How It Goes Wrong

Let's say a maintainer creates an open source paint program, OpenPaint.

An artist, known for their unique style of making blended collages, switches from their previous app, ProprietaryPaint, to this new OpenPaint to make these culturally significant works of art. However, the maintainer decides that the 'blend' tool is kind of a pain to maintain, and they remove it in OpenPaint 2.

A few months later, the artist's operating system vendor issues a security update that breaks OpenPaint, because older versions of OpenPaint were unknowingly abusing some platform API.

The maintainer releases a new OpenPaint 2.0.1 that addresses this incompatibility, but doesn't care about version 1.x any more so they don't bother to update that one.

This places the artist in an impossible situation. They can stay on an old version of their operating system, putting all their personal data at risk. Or they can upgrade to the new operating system, effectively either cutting off access to their livelihood, or forcing them to change their art style entirely.

Now, proprietary software can place users in similarly untenable positions (and in fact, it is more often proprietary software that does). But does the openness completely remove any obligation for this consideration? Should the OpenPaint team have to at least communicate the reasons for doing this, to give the artist some recourse?5

The only thing that "open source" does is that it allows the artist to pay a prohibitive amount of money to a new maintenance team to create a fork. This is rarely the kind of thing that individuals can manage.

This creates an obligation to at least consider how your users might be relying on you.

This is the most complex obligation of the bunch. Obviously it does not entitle every single user to infinite work from the maintainer, but it also shouldn't entitle the user to nothing for having trusted these subtle implied claims that the maintainer is making by making their work public.

It is a nuanced and ongoing negotiation and I do not think we have a clear moral intuition about how it should work out. But we do need to figure out a way to work it out.

It also raises a clarifying question.

Why Are We Even Doing This, and Who Are We Doing It For?

People generally like to do things for more than one reason. We live in an economy where people need to make money, but we mostly prefer to make that money doing things that are useful, and that make other people happy.

So, yes, we create open source for self-interested reasons to improve our reputations, to improve our skills, to increase our influence and to share our maintenance burdens. In so doing we take on some level of obligation to not abuse the trust that is placed in us, even if that level of obligation is not clear.

But if we are not doing it to serve those users at least a little bit, then those motivations are going to quickly ring hollow. We will not increase our reputation with a person if we respond to their every request by telling them that we owe them nothing and that their opinions are worthless. We will not gain influence over a community if we ignore their desires.

Many interactions with open source maintainers are unnecessarily adversarial. This is of course partially the fault of those users, who should calibrate their expectations appropriately.

Still: maintainers could do a better job of listening before these interactions become toxic. There's no reason that "open source users" should be an especially toxic group of people. At this point in history, that group is basically just … people with computers.

It's like that old truism. If you meet one person who is a jerk to you, that's their problem. But if everyone you meet, everywhere you go, is constantly abrasive to you and treats you like you're doing something wrong, maybe it's time to look inward.

If all open source users are entitled assholes, maybe it's time to look for a structural problem.

Surprise, It's About AI Again

Sigh.6

Users hate slop.

I know, dear AI-positive reader, your AI outputs are different from everyone else's, you aren't pushing thoughtless slop into your code, just because everyone else is and it is the inevitable terminus of using those tools. You aren't "lazy vibe coding" with Claude, you're doing "responsible agentic engineering", which is different because you're just built different.

Still, humor me, for a moment. Your users don't know that. They know what it looks like when products that they like adopt slop. They know that they will start leaking data. Developers know that it will make them personally less secure. They know that they can expect more outages and that your code will inexorably decline in quality.

In other words, your users are going to assume that this means you are violating that final obligation that the software should keep working.

Your users are going to tell you to stop, and they are probably going to get mad. Maybe you, or a plurality of your team, also want to stop, maybe you disagree with them, but in any case you need some way to have that conversation in a way that does not immediately overflow into every adjacent discussion forum. Users need to feel welcome in some space so they can have the discussion in that space, and not explode out into a thousand different group chats and social media threads.

This post was inspired by yet another prominent open source community discourse where a ton of angry users showed up to yell at developers to stop accepting LLM-generated code. I'm not going to link to any of these, because we don't need any more fuel for the discourse fire. But there is more than one such case and the pattern is becoming familiar.

On social media - usually BlueSky or Mastodon, but sometimes a user group forum - users become aware of some AI-adjacent policy. They show up in a horde to the developer forum or mailing list. They loudly start demanding the project take a hard stand7 against AI. This pressure is simultaneous, but uncoordinated; extremely repetitive, very diverse, often inconsistent, and pretty stressful, especially if you're a burnt-out maintainer with other things to be doing who may not even like AI yourself in the first place.

Believe me, I get it. It can be very unpleasant to deal with.

Like most problems that AI is causing, though, it's not really an "AI" problem as much as it is a pre-existing dumpster fire that "AI" is pouring gasoline onto. In this case, an online mob is the language of the unheard8.

If Users Are Mad It's Probably Already Too Late (But Maybe You Can Get Ready For Next Time)

One day, all of a sudden, you're getting feedback from a bunch of users that are using inappropriate channels to complain. But did they already have appropriate channels to use?

Did you have a place for people to congregate and discuss your project? To make orderly complaints in a way that will be legible to you? Or do you just have a GitHub Issues page, which non-technical users have no idea how to interact with, and a forum for developers, where users don't know the norms and any arriving brigade of pissed-off users will be seen as disruptive and inappropriate?

I don't want to be throwing any stones from within my particular glass house. Setting up such a place has gotten harder over the years. I don't really have one, either.

Could I have one, though? IRC has been slowly dying, mailing lists are unpopular and present increasingly annoying moderation challenges, forum software is expensive to operate and keep maintained, Discord is a confusing mess and the upshot of all of this is every community needs community management and forum moderation. Which means that for my own small solo projects, I couldn't possibly have such infrastructure because such infrastructure requires a dedicated second person to maintain it, and until someone volunteers for that, it's not really feasible. Even for my larger projects you'd be surprised how slim of a skeleton crew we are getting by with, and we definitely don't have a whole spare maintainer to go manage this, especially as we are under attack from the slopocalypse ourselves.

The nature of open source community is that most communities start too small to need such a thing, grow incrementally until one day they are suddenly way too big and needed one yesterday, and then suddenly they are too small again when interest wanes even a little bit. Even as we need it more and more, building and maintaining community infrastructure remains a challenge.

Even so, having a dedicated place for users - not maintainers - to converse amongst themselves, be an actual community, and present feedback to the developers, is fast becoming a necessary component of a successful community and not a nice-to-have.

In Conclusion

As trying as it can be sometimes, we maintainers all do get something out of open source, and it is good to be honest with your users - and with yourself - exactly what you want to get out of it. In order to know whether the juice is worth the squeeze, we must know both what the juice is, and what the squeeze is.

Part of the metaphorical squeeze is a set of obligations, and those are the most poorly defined of all. We should try to be clear about what those are too. Both about exactly what we believe we are signing up for, and also, about how we are willing to let our users hold us to account for them. Codes of conduct are a start here, but only the absolute barest bare minimum; "do not harass your colleagues or your users" is not a standard of excellence to aspire to, it's just basic manners.

I can't tell you exactly what your obligations are, only try to gesture at my idea of the outlines of the fuzzy moral intuition we've all been implicitly sharing up until now.

Drawing this line is not just for the benefit of the users, either. Maintainers already feel pressure, we already feel obligations. We resent that feeling of obligation. While there are a diverse array of reasons for that resentment, one big one is that it's not clear, even to ourselves where the obligations end. Lashing out by saying "I promised nothing and I owe you nothing!" followed by some choice expletives feels cathartic, but it doesn't really solve the problem, because we clearly don't really believe that's where the line is, or we would have already stopped there. We wouldn't feel the need to say it.

It is going to be a very big collective endeavor to figure out exactly where that line is. The best time to have gotten started on that endeavor was 50 years ago.

But the second best time is today.

Acknowledgments

Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!9


  1. Somewhat to everyone's surprise, I, too, do things for money, like writing this post. Please remember to like and subscribe ↩

  2. Whether it was originally yours, or developed by an employee who happened to be on staff at the time, or adopted by an employee who just started contributing to it a lot, in any of these scenarios a company can benefit from increased consistency and increased familiarity. ↩

  3. If the rule is that they must forever endure the searing budgetary pain of gripping the white-hot potato that they unwittingly caught when they first made a good technical choice, this creates a perverse long-term incentive. ↩

  4. I also find it darkly amusing that there is an explicit affordance here made for advertising, specifically, "Packages with code that can be used to display ads are fine. Packages that themselves display ads are not." This distinction rather gives the game away, that this is a website for carnies and not for marks, and that at some level we expect our users to deserve a lower level of respect than ourselves. But a full exploration of that is another blog post, or maybe a book, that I don't have time to write right now. ↩

  5. If you want the turbocharged ultra-dramatic version of this problem, make it open source drivers for an optical prosthesis that lets the users see instead of an art app. That level of immediate physical dependency could be clarifying. It does also start to edge into an area where you could say that biomedical devices ought to be regulated differently, and that's not really a "software" problem but a "healthcare" problem and I'd mostly agree. Except for the fact that this is a very short distance away from breaking everyone's screen-reader with no notice or recourse. ↩

  6. Did you believe I could write a blog post in 2026 which wasn't somehow about AI? I wish I could still believe that. ↩

  7. It doesn't help that many of the most pro-AI voices are starting to have an, ahem, discernible political valence that is very unpopular among users. ↩

  8. My apologies to MLK. ↩

  9. If you read this whole post you can see that I sure need the help with all that. ↩

25 Sep 2026 12:50am GMT

06 Sep 2026

feedPlanet Twisted

Glyph Lefkowitz: ... but what about video games?

I get asked this rhetorical question a lot, in various forms:

Sure, datacenters might use a lot of energy, but you don't have to use a hosted frontier model to do software development. What if I just run a local open-weights model to do some coding, with an open-source coding agent? Video games also use my GPU. Is local model development any worse than playing a video game?

So I want to write down my comprehensive answer to this: Yes, using an LLM to write some code is worse than playing a video game, for a few reasons.

Video Games Are Interactive, LLMs Are Batch Jobs

Video games use compute to respond to human input. You are using your GPU while you are looking at a screen, displaying an image. When you are done playing, you shut off the game, and your computer goes back to idle. It's much less energy. By contrast, agentic loops with evals (the only kind of "AI" that is meaningfully any good at coding) are running hot, for days. To use the most recent example of such a thing, a very rough first sketch of an implementation of a Windows graphics API backend to help port a paint program to other platforms, it took 3 weeks of Claude time, "day and night". Do you play a lot of video games for 500 hours to make it past the tutorial level, while also using other computers for other things, as well as the rest of your carbon footprint?

Video Games Need Development, LLMs Need Training

Video games use compute to respond to human input during development, too. Your game has to be made, but your LLM has to be trained. LLMs use a historically extreme amount of power, probably using more than the entire Internet, but it's kind of hard to say. Still, it seems a reasonable estimate to within several orders of magnitude that even over a multi-year project with hundreds of developers, the power used to develop an individual video game is nowhere close to training even a small LLM.

This is true even for local models. OpenAI has openly claimed that DeepSeek "stole its intellectual property", and I have heard grumblings that none of the open-weights generalist models could realistically exist without the massive lift that the frontier labs are doing with their training, in various other ways too. Secrecy throughout the industry makes this kind of impossible to understand rigorously, but it seems fair to say that you are partially culpable for all that famously energy-intensive frontier lab training if you're using a local model.

And They Keep Needing Training

You also can't dismiss this as a sunk cost, because in order to stay current with industry developments, models need to be updated with new information from the rest of the world, which means that you need to keep training them. Beyond the energy for your own use, if you want a real-life agentic workflow that actually does useful stuff, practically speaking you would still need to update your local models over and over again, at least once every few months, which means you would be incentivizing continued energy consumption by whoever was doing that training for you, including the energy cost of scraping.

Let's Be Real Here, You Aren't Actually Using A Local Model

This question is a hypothetical thought experiment. Despite synthetic benchmarks that keep showing there isn't much difference between open weight and frontier models, nobody's actually using local models for much of anything beyond sharing those talking points. Depending on which benchmark you're looking at, maybe it's good enough or maybe it's worse.

As an inveterate AI hater, all these systems seem pretty bad to me, but it seems that people who find them useful tend to subjectively believe the frontier models are worth the premium, and that's what they're actually using. Once you have accepted that it is OK to use LLMs for coding at all, it seems like a very quick slippery slope on down to "we'll go ahead and use the frontier models for now anyway, but we could be ethically better in the future by switching to an open weights one, that option is always available".

There's A Reason We Have Data Centers

Devolving power usage to local LLMs might be good to make users responsible for their costs and decrease the impacts to communities that are physically next to huge concentrations of power utilization, not to mention generation. However, there's a reason that it makes sense for the providers to build these giant facilities: economies of scale reduce total power consumption, they don't increase it. If you do all the same stuff with a local model that they have to do in hosted environments, it will probably take more power, even though you will be incentivized to do different stuff. This incentive to "do different stuff" is why although local models can hypothetically hold their own against the frontier labs for some tasks, when people or businesses take their inference costs in-house they often find that it's too painful and move back to hosted LLMs.

There Are Problems Other Than Power

These are subjects for a different post, but you have to consider a lot of other externalities: AI psychosis, de-skilling, comprehension debt, cultivating a dependency, introducing security defects, limiting your design space based on what LLMs can understand, context rot, wasting time on invalid solutions, introducing unpredictability into your workflows. You still have to consider the total cost benefit ratio.

To Sum Up

Local LLMs might alleviate some of the harms from using the hosted frontier providers. There are fewer privacy concerns, you can measure your power utilization and be more directly responsible for it, you can build interfaces with affordances that are less oriented towards addiction and dependency than the major frontier labs' harnesses.

But they're not automatically "the same as playing a video game" just because they can use the same GPU.

Acknowledgments

Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!

06 Sep 2026 10:57pm GMT