09 Oct 2026
Planet Twisted
Glyph Lefkowitz: Programming Isn’t Special
Creative Work
Writers went on strike to get protections against "AI". Thousands of artists have signed open letters in protest of "AI". There are so many copyright lawsuits from creative industry groups against AI that there's a whole dedicated website for it. Popular YouTubers absolutely hate it. If they're also musicians, they REALLY hate it. Across all creative industries, there is a concerted push to reject this technology.
Yet, almost unique among creative fields, many experienced programmers remain convinced that it's fine to use "AI" for programming. We do seem to hate it, and it's making us all miserable, and what it's doing to our industry, but we are using it anyway.
A lot of the justification of this resignation seems to be to be because programming is not Art. If the tool can get the job done, and the job is just functional, then why does it matter?
It does matter, though. It matters because we shouldn't be using AI to produce Art, and programming is Art.
Art can be Mundane
Some people will say that programs cannot be art because programs are functional, rather than being expressive. Programs are mundane whereas art is transcendent.
This is based on a distorted understanding of what Art actually is.
In John Berger's "Ways of Seeing", he names this type of distortion "mystification". His example of this process is both amusing and illustrative. I encourage you to read it in its entirety.
In summary, though: Berger critiques the florid prose of an art historian describing a commissioned group portrait, including phrases like "subtle modulations of the deep, glowing blacks" and "harmonious fusion". The portrait is described as sublime, in nearly ecstatic terms.
Berger reveals that the reality of this portrait is that a poor old painter needed some work, and some officials probably thought it might be nice to have an official portrait. So they paid some money to the poor old man, and he painted it, and then they had a painting. It's a well-executed portrait of a group of people. Beautiful, even. But it was work commissioned for a fairly mundane purpose and it suited that purpose just fine. It was not, and is not, a divine relic.
Culturally, we are prone to mystifying painting, and sculpture, and film, and music. We imbue them with "subtle modulations". We ignore their functional aspects - we desire decoration, amusement, and distraction - and focus on their emotional impact.
Don't get me wrong: I love me some good aesthetic philosophy. I think it's great to really examine our reactions to artwork and to try and gain a deeper understanding of our culture and our selves through media analysis. If anything we really need to do more of it.
This does not mean that the creation of such works is mystical or that it should be venerated beyond any other sort of labor.
Not least of which other types of labor that are adjacent to, but not as culturally venerated, as fine art. We tend to mystify the work of a novelist, but to denigrate the work of a journalist. In reality, the functional prose of the journalist is no less important and deserves no less respect.
Although they might be far below the ethereal realm that novelists inhabit in our collective imagination, even journalists receive more respect and thus more mystification than lowly copywriters. Yet, there is no transcendental distinction between "novelist" and "copywriter"; the many of the skills are the same, and the distinction is merely an accident of commerce and opportunity.
In fact, many famous writers have famously inhabited both roles. This is not an accident! Working with words professionally, even (perhaps especially) mundane words, is excellent practice for working with words in a more purely artistic context, because even mundane creativity is still artistic.
Code can be Beautiful
One thousand Internet years ago, when I was in my late teens, I would describe myself as a "code poet". I was relentlessly mocked for this as what the Youth would today call "being cringe", and at the time was referred to as "pretentious".
I succumbed to the peer pressure, removed it from my email signature and my bio. While I still believed strongly in the parallels, I accepted that - socially, at least - comparing code to poetry, or indeed to Art, was a silly thing to do.
However, I never abandoned the idea, in my heart.
The thing that I am most well-known for, the invention of Deferred, was specifically an aesthetic reaction to the tedium of passing callback and errback parameters to every single remote procedure call in an RPC client/server application. Those two callbacks got the job done just fine. But they were ugly, and annoying to work with.
Deferred is an intentional poem about asynchronous task execution, with a deliberate eye to the aesthetics of the problem and the experience of using it. It was influential because of its focus on aesthetics.
I do not want to overstate the beauty or profundity of this minor contribution, or indeed its durability. That a poem exists does not mean it is a great poem, merely that it is a poem.
Our aesthetic culture around programs is more like folk epic poetry than fine art, so the influence of this contribution is less about its specific enduring power than it is about its influence on what came next; from MochiKit.Async to JQuery Deferred to JavaScript Promises and eventually to async/await; a long chain of different artisans each adding something of their own until the original has all but dissolved. (And I wasn't the "original" here, either, as I drew heavily from the E language's Promises, among other things.)
In order to make code into a deliberate artistic expression, one must have spent quite a bit of time contemplating the problem domain. Without having experienced the tedium of manually passing a thousand callback parameters, I would have had neither the skill, nor indeed the motivation, to bother creating such a thing.
Now, most code does not have to be like this. Most code does not get to be like this. Most code is functional, workday code. Most code could not make a lady weep. It's just copy-writing, if you will.
As I explained previously, most writing couldn't do that either. Most writing is just copy-writing, too. Most visual art is advertising. Most live music performance is background music in bars that will go largely ignored.
However, code that is intentionally aesthetically designed tends to be important, both socially and technologically.
We do have some tradition of self-mystification in software. As Abelson memorably put it, "Programs must be written for people to read, and only incidentally for machines to execute.", so we have long had some conception of programs as highly expressive, even if we can't always agree on what they're expressing or to whom. We will occasionally wax poetical about the philosophical implications of a particular piece of software. This is not unique to a single piece of software, either; more than one community has indulged in similar philosophizing.
The expressive and aesthetic qualities of software are not limited to reading source code or interacting with other programmers via APIs, either. For example, every year, Federico Viticci does a review of Apple's new operating system, which is (among other things) an aesthetic critique. Such a project would not be possible if the software did not have an aesthetic impact on its users.
Not to mention that every video game review is also a software review.
A Brief Aside about Software Literacy
It does make me a bit sad that we don't have much of a critical reading tradition in the software community. Literate Programming is often praised, but rarely practiced.
Moreover, it makes me sad that users have a pretty jumbled idea of what goes into making software, that programming literacy is pretty low, and that modern programming practices often deliberately produce a bad mental model of what the software is doing so it's even harder for the user to understand. The aesthetic experience of software is often wildly detached from its internal state.
While all of these problems predate AI by years or indeed decades, that's no reason to enthusiastically make them worse.
Defend The Mundane
If we use AI to erase all the copy-writing, all the graphic design, all the boring mundane art, and yes, all the boring custom WordPress theme development, then we will be removing all the practical opportunities for the vast amounts of practice and contemplation required for people to elevate their craft to eventually achieve great things. Education is great, but the majority of true skill development happens on the job and always has.
This doesn't mean that we can't use abstractions, or automation, to make our work easier. Programming is the art of abstraction, of understanding how to compose smaller ideas into bigger ones, of how to understand the automation of a larger system by understanding the rules that automate smaller ones and understanding how to combine them.
When we use "AI" to eliminate that understanding rather than raise it up to a higher level, to entirely destroy that creative decision-making process, we do a disservice both to ourselves as programmers and to our users. We would be doing a disservice to our users and our downstream fellow developers in the same way that a visual artist would be doing a disservice to their viewers or a musician would be doing a disservice to their listeners if they served them auto-generated filler instead of their own creative output.
Slop is slop, no matter the medium.
Each mundane project has some tiny chance - let's say, something like 0.1% - of achieving greatness. If we do a single project with AI, then sure, whatever, there's almost no chance that that project was going to be the one hit to create that career-defining moment for an engineer working on it. If we make a habit of doing all projects that way, though, we take the total likelihood of those moments of greatness from "definitely sometimes" to "never".
The precisely appropriate ways in which to resist AI encroachment on all software development lie well beyond the margins of this one short post. How much you can resist and which specific uses you should resist are up to you. But it is worth resisting in software just as much as it would be worth resisting in any creative medium.
Programming isn't special. It's just Art, and Art is the most human - and thus, the most universal - thing that there is.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
09 Oct 2026 7:00am GMT
28 Sep 2026
Planet Twisted
Glyph Lefkowitz: What Would A Serious AI Product Look Like?
One of the issues that I have with the current generation of "AI" products is that they do not appear to take their own premises seriously. I look at a plethora of obsequious chatbots claiming to be serious tools for problem solving, and I think, this is not what a problem-solving tool would look like.
Even before we get to the tremendous ethical problems with the frontier labs, it is this impression of their composition as a product that makes me feel, constantly, whenever I am interacting with them, that they are less a software product than that they are a grift, a scam designed to make me feel like I am interacting with a product that has capabilities that it simply does not, to try to lull me into a false sense of security that I can trust it.
The frontier labs are of course the worst offenders, but every criticism here applies just as much to Ollama, which (if anything, due to the obviously poorer quality of the available models themselves) needs these features even more than the frontier labs do.
Here, I will set down a few features that might convince me that an LLM-based product, particularly one focused on research or software development, was actually serious about helping me do useful things with it.
Make "Checking For Mistakes" A First-Class Feature
This is the biggest issue, and the major reason that I was inspired to write this post.
It is a truth universally acknowledged, that AIs cannot reliably provide information.
I could cite a ton of news articles and studies about this fact, but there is no need. Every single chatbot admits this, up front, in a fine-print disclaimer as a core part of their user interface. Gemini says "AI can make mistakes, so double-check responses", Claude says "Claude is AI and can make mistakes. Please double-check responses.1" ChatGPT says "ChatGPT can make mistakes. Check important info.".
Every time I see that last one, I wonder how I'm supposed to know what "info" is supposed to be "important".
All of these warnings are all small, gray text, painfully obviously included as legalese to push responsibility back onto the user rather than to help with anything. This is a core limitation of all these products. Checking their output is a part of the workflow for using them that:
- you absolutely cannot skip or skimp on without creating risks to yourself and whoever you are conveying its output to, and,
- it is very easy to skip or skimp on and you are encouraged at every turn to do so, because "just trust the output" is one of the quickest ways to save time.
A chatbot product that took this weakness seriously, as an actual consideration for using it, would put a checkbox next to every claim in its output. It would be a 2-column worksheet, where you've got the LLM output in the first column, and next to it, human notes in the second column, explaining what work went into checking this claim, and a big checkbox that you would only check off after you believe you'd checked its claims thoroughly enough.
Coding assistants would need to have some version of this as well. Right now, this is pushed off into code review, which means it is a dark pattern which subtly encourages the "author"2 to offload this work to their code reviewer without ever looking. Once again, "it's probably fine, I don't need to check" is the quickest way to save time and churn out those PRs faster.
It might even be useful for coding harnesses to have some affordance for checking code before it even runs tests. As the vendors themselves have admitted, it's not just expensive to burn tokens on your "AI", you also end up burning far more compute on the AI. Being able to check your diffs before sending them over to uselessly exhaust your testing compute cluster would be useful.
If your product tells me that it makes mistakes and I must be the one to check for the mistakes, but then gives me zero tools to check for mistakes, I cannot take it seriously.
More Citations to Check, And More Details
Most chatbots prefer to give an answer, rather than a citation. In my own personal use, I find that when asked to provide a list of citations with clearly marked sources for each one, they will appear to "get bored" halfway through the list and simply stop including citations at some point.
When the bots include citations at all, present them as inline annotations that say nothing but the domain name of the search result, in a font so small that it's barely legible, and an equally indecipherable icon that is fewer than 16 pixels on a side.
This is backwards.
Now, I am aware that these citations do come from somewhere, and in an attempt to reduce hallucinations, all of the major providers support some form of "grounding"3, and that those little barely-readable citation links are referencing actual structures in the RAG pipeline and not just potentially-hallucinated tokens, but I'm not talking about the underlying machinery in the model, I'm talking about the presentation to the user.
Plus, regardless of whether a snippet of text came from a RAG query, we know that LLMs can never provide an authoritative result; it's a fundamental limitation of the technology. They can still garble the results of RAG as much as they can misrepresent any other training data. This means that it must never present its results as authoritative.
If you ask an AI to do research queries, every result should be presented as a list of citations. Moreover, the presentation should display each citation as a large object of in its own right, with clearly identified metadata, including not just the site where it was found but its publication date and, if possible, the name of the author. The literal, unmodified quotation (not from RAG, not a summary: a quotation extracted with a regular program and not an LLM) should be front-and-center, larger than any AI-generated text.
If the AI product wants to editorialize or summarize (which should not always be necessary!), the AI-generated text should be presented as small text underneath the citation that has been found, de-emphasized as much as the disclaimer is right now, at the very least until the user has verified that the summary is accurate. Perhaps, for a research project, a "did you read the citation" checkbox might even be helpful.
If your product openly tells me that it will scramble, misrepresent, or omit its citations in its summaries, and I must read the original human-authored citations to be sure, but then gives me no tools to track my reading of those citations or even any way to find them, I cannot take it seriously.
No First-Person Output, No Apologies
There is no reason for a software development or research tool to use first-person language to describe itself. They should not do so. In fact they should not be allowed to do so.
There is also no reason that they should ever apologize. It is a waste of everyone's time; it's a waste for the chatbot to generate the apology, it's a waste for the user to read the apology, and it's a waste for the user to respond to the apology. Yet they unfailingly do this upon every correction.
The vendors of these tools know that they are routinely causing mental-health crises. In response, they have added non-functional "guard rails" that can still, in 2026, easily be bypassed.4
A product seriously interested in helping with productivity would correct this glaringly obvious flaw, focus on the task at hand, and stop emitting useless verbiage.
In the previous two sections, I tried to focus on ways in which the harness would be constructed differently even if the LLM technology is fundamentally impossible to improve; in this case, I have to assume that the labs have some control over the model itself. But unless they are truly incapable of influencing their output (and all their "benchmarks" and "capabilities" seem to indicate that they can control it very tightly) they ought to be building models that are much less verbose.
More Non-Natural-Language User Interfaces
Although natural language could hypothetically be a powerful interface for interacting with a computer system, the practical upshot of LLM natural language interfaces is that these interfaces are imprecise and repetitive, full of superstitions masquerading as "best practices". The inputs are a mess and the resulting outputs are a mess.
The general way of addressing this unstructured mess is to allow the chatbot to directly take action in response to the user's input; in other words to supply it with "tools" via an MCP server. But again, this is backwards. If we cannot even express our intent clearly in the first place, why are we trusting this system to take potentially destructive and harmful actions on our behalf?
Instead, I would expect a product that was seriously invested in helping me accomplish specific tasks, to have user interfaces specific to those tasks. Is it supposed to be able to be a security scanner that can discover OWASP top 10 bugs in a codebase? Have a button for that. Build that functionality into your harness, train it directly into the model, use smaller models that can satisfy that functionality more effectively than throwing it at the planet-sized brain of Fable or whatever.
I'm aware that there are small software startups that do something like this, but they are bolted on to the side of the main model providers' APIs, not integrated into the core of the product and not using their own models and AI systems to achieve consistent and repeatable results.
Strong Data Provenance Indicators
Chatbots produce data tables pulled from websites, from APIs, from MCP tools or from summarizing and scrambling the user's input. In order to provide the illusion of a seamless interface, this data is presented in-line regardless of where it comes from. But some of these outputs are produced mechanically via regular old API calls, for example, from the result of calling a tool or querying a website, but presented uniformly.
But there is a huge difference between an authoritative data source being inlined as part of a chatbot conversation, being treated as input by the chatbot, and some ad-hoc hallucinated data being treated as output of the chatbot.
If a product is trying to help me make accurate, empirically-grounded, data-driven decisions, the source of the data is critical.
Integrated into the "check for mistakes" and "verify citations" workflow I described above, there's a necessary "verify data programmatically" pass as well; to have tools that will treat portions of the output as a regular spreadsheet, allowing regular-old computer arithmetic to verify things and showing where such arithmetic was used, and how.
Better User Control of Reproducibility
Anyone familiar with the technical specifics of LLMs will know that they have a variable called "temperature" which controls the degree of randomness that the LLM uses to produce its outputs. But most users don't know this, because it isn't exposed as part of the user interface by default.
This leads to a subjective impression that you asked ChatGPT, and you got ChatGPT's authoritative answer.
You can't just set the temperature to zero and still get useful results - I am aware that it does more than just scramble the output at random, and there are perhaps good reasons that simply exposing just a temperature setting would not be that useful to users. But if we followed some more of my earlier recommendations for making more structured UI elements to solve specific problems rather than having long back-and-forth chats where each refinement depends on the previous response, perhaps those elements could also re-play the process so that users can see how reliable the bot is at a particular task and develop a sense of how the stochastic nature of the process actually affects it.
Similarly, if a user is trying to solve the same problem repeatedly with a chatbot, and the chatbot product has numerous computational tools that don't really have anything to do with the LLM, such as deterministic data-processing tools, then having a way to freeze the non-deterministic parts of the transcript but re-populate a particular data frame with updated information and fork / continue the conversation from there would be a way to avoid introducing pointless additional randomness when you already know what tool you're trying to use.
The fact that every conversation is presented as this flat chat prompt that doesn't let me interact with any of the widgets that were previously produced except through more chatting, really makes me feel like the whole product is just doing predatory social-media style "increase time on site" optimization, just trying to lure me into further repetitive and unreliable chats, rather than letting me get in, solve my problem, and get out.
Context Visibility
Managing the LLM context is the ongoing challenge facing organizations that are trying to use "agentic" workflows. Filling up the context with too much information causes well-known problems. In response, advanced LLM users attempting to solve larger problems must break up very long prompts into "skills", give access to lengthy information via "tools", and delegating sub-problems to "sub-agents" rather than simply extending a single prompt indefinitely.
All of these strategies have flaws, because even on the largest models, compared to the breadth and depth of knowledge-work problems, LLM contexts are quite small.
And yet, none of these products will show the context to the user by default. There are third-party addons that can show you a simple progress bar but for addressing the premier engineering difficulty with this technology, that is below the bare minimum.
This lack of visibility means that almost all of the tools for extending the context are flying blind. Rather than responding meaningfully to a full context, everyone just kind of guesses how much state they need by guessing and trying over and over again with progressively more elaborate skill and sub-agent layouts. Even managing context compaction ends up being an advanced API-driven workflow5.
A serious product that was trying to help the user understand would not only show "available context" but explain the impact of context compactions, make it easier to see harness-generated prompts, and so on. This would be a first-class feature, combined with the aforementioned reproducibility / replay tools, would allow users to do real experiments to develop an understanding about how to make good use of the context window.
A Sandbox That Actually Works
I've been focused on the chatbot interface here because it is the most immediately egregious upon looking at the UI. But the "agentic loop" tools used for coding are equally dangerous, if not more so. Coding tools keep destroying everyone's data, over the course of years.
These catastrophic incidents that become front-page news are relatively rare compared to the amount of coding-agent use out there. But they also aren't the only kind of sandbox violation. Coding models will so routinely edit test code instead of the system under test that there are "pro tips" articles all over the web giving you the flawed advice to simply ask the agent not to cheat. News write-ups of the catastrophic incidents themselves will also offer glib and wrong advice, like "use a docker container". That might prevent it from literally deleting your operating system, but it won't prevent it from destroying all the local work you have in your codebase (it needs access to a checkout, after all!)
There is a flurry of activity in the infosec space where people are rushing to plug the gaps left by these coding harnesses. Everyone's got their own version of an MCP approval gateway where you can optionally place a proxy between your agent and your production infrastructure.
In the best case, though, all these mitigations and proxies and prompts simply turn the user into an auto-approval automaton, hitting Y, Y, Y, Y over and over again, until you finally are driven mad and hit "yes to all", turn on full-auto mode and submit yourself to the void. With nothing between your personal vigilance and disaster, there are no workflows left beyond decrementing your own vigilance until there's nothing left and then hoping the disaster never arrives.
The fact that some mitigations exist that can be deployed by extra-cautious users does not change the fact that "agentic coding" is an unsafe-by-default technology deployed without concern or guidance. Every frontier lab has tied a spring-loaded shotgun to a dog; the fact that dog owners can publish thoughtful blog posts explaining how you can teach your dog the basics of gun safety or how you can have your dogs play in a bullet-proof room does not mitigate the fact that the product should not have been allowed in the first place, nor should it continue to exist without VERY strong security controls.
I might believe that a frontier lab were seriously interested in providing developers with a useful tool if they shipped something that had safety built-in.
That means tools in the harness, detached from any LLM, independent of the prompt, that could:
- sandbox all filesystem operations and strictly limit ANY deletions outside of specified scopes, regardless of operating system,
- enforce snapshotting of the entire repo on every operation for easy rollbacks and minimal lost work,
- remove the disaster of "auto mode" (not to mention nonsense like
--dangerously-skip-permissions) entirely, and - carefully consider a structure for presenting plans to the user where, rather than provoking immediate alert fatigue by asking for checks on every action, make structured plans which can be submitted to the user as a group of actions and reviewed and approved as a batch.
In the same way that I suggested above that research-based tasks should have a way of re-issuing prompts to determine how reproducible a result is, or whether other sources might be found, agent-based tasks should have a way of being executed against mock services for popular APIs, so that the verification can match both on the front-end (review the plan for making the API calls before they're executed) and the back end (review the API calls that were issued to the mock service and verify that they matched).
Instead, the frontier labs provide us products that are disasters out of the box, give us "best practices" to build massive and elaborate, as well as incomplete and error-prone, security perimeters of our own design. Then they blame "operator error" when it inevitably goes wrong. I cannot believe that these design choices are intended to help us be productive.
Bonus: Human Processes
Organizations deploying AI also frequently come across as unserious, for similar reasons. In 2023, naive exuberance could perhaps be forgiven. But today, as we near the close of 2026, there are several well-known problems, that have been extremely well-covered in the press. None of these things should be surprising, but most orgs deploying these tools are still just letting them rip and hoping it all works out.
Organizations deploying these tools would need at least three kinds of major modifications to their internal processes, if they wanted to be serious about using them safely:
1. Shift Rotations to Prevent Vigilance Decrement
There have been several high-profile incidents where software developers' gradual acquiescence to accepting LLM output have lead to serious economic consequences for the companies deploying them, perhaps best typified by Amazon's "millions of lost orders" due to a gradual decay of their engineering processes from LLM use.
These outages, and other AI-related failures, are due to the difficulty of maintaining focus on the same problems. In other words, as I described above, vigilance decrement is a constant problem, because AI outputs are most often correct, but continue to be incorrect in surprising and non-intuitive ways. As I have previously written, you cannot trust yourself to catch every bug with code review, and LLM output.
Aviation, for example, has very strict rules around rest requirements. There is also a specific rule that "No certificate holder may operate an aircraft without a second in command if that aircraft has a passenger seating configuration, excluding any pilot seat, of ten seats or more.". Other safety-critical professions have similar rules.
And yet, even in the age of the supposed "AI revolution", most software teams are still assigning every engineer a full feature load, not planning for any rest, and telling people to review code whenever they happen to have some "free time".
Maintenance of vigilance has to be your top priority. Regular, scheduled, inviolable rest periods where people do work without AI assistance, and are not exposed to any AI output for review or otherwise, would be crucial in order to stay mentally sharp enough.
The tools themselves should have this sort of thing built in. The mistake-review process described above should have a periodic spot-check mode where a second reviewer periodically reviews a chatbot log, doing their own independent verification of claims, to see if they spot the same errors. This could provide a feedback loop to determine how much rest is necessary to maintain continuous attention and actually spot hallucinations.
2. Skill Practice To Prevent Skill Loss
It is also well-known that AI use leads to AI reliance, and AI reliance leads to skill loss.
I like to use the analogy to dockworkers at a seaport6 adopting automation.
If you employ dockworkers to load and unload ships all day long, they are going to be getting tons of exercise. They will be able to lift heavy objects on demand, whenever. They might have plenty of health problems and injuries from this type of work, but "lack of exercise" will not be a problem.
With the development of standardized container ships and mechanized cranes, you are going to be changing their job description substantially: now they mostly spend all day sitting in a small cubicle moving a control lever back and forth, not lifting heavy stuff. They will get worse at lifting heavy objects.
In this analogy however, the cranes are not all that reliable. We know they break, and they drop their payloads sometimes, and the stuff needs to be manually moved. But this only happens a few times a week, at most. If you need whoever is driving the crane to be able to jump out at any moment and still move stuff around manually, then you need to make an affordance for that. You need to give them time to go to the gym and do some lifting for practice, or every crane failure is going to be a major emergency.
An organization doing an AI transformation would also need a massive increase to learning & development budget, both in terms of resources and in terms of schedule. If your people are going to lose skills because they've lost regular practice in the incidental course of doing their duties, then they are going to need deliberate, intentional, non-incidental practice of those skills to keep them sharp.
But rather than trying to accommodate new workflows and give time for people to adjust, most AI mandates are simply dropped on workers like a ton of bricks, with no time to adapt and no affordance for maintaining their skills. Operate the crane and stay fit and healthy and ready to switch back to manual lifting at any time and then get back in the crane cockpit right afterwards. Don't mess up.
Then an accident happens and everyone is surprised, as if this process weren't practically designed to produce a terrible result.
3. Mental Health Resources to Deal with Mental Health Risks
AI psychosis often begins with practical problem-solving, and beyond that, it can start specifically at work. Not to mention the more pedestrian condition of "AI brain fry".
If you are mandating your employees to use a hazardous tool that may seriously and directly damage their mental health, you need trainings and resources. You need in-house therapists and you need to be making sure to check in with people actively to make sure that this is not happening.
Again, the tool itself ought to have some way of dealing with this. An occasional "take a break" popup is easily dismissed; they need a user-visible AI personal dosimeter so you can see your cumulative usage over time.
I don't even know if "usage over time" is a sufficient metric to gauge risk. Maybe if your work chatbot start to talk about resonance too much, unless you literally work as an acoustic engineer, that should be flagged for someone.
We are, again, years into dealing with these tools, and we know these risks exist. Yet no serious mitigations are provided. Not even any way of measuring the risk exposure.
And More
There are also many other risks associated with the technology. There are intellectual property risks with the foundation models, due to recklessness with their training data. There are existential financial risks associated with the infrastructure build-out. The extent to which most "open" models are simply derivatives of frontier models is an open question.
What I Think
If any one of these things were regularly overlooked by AI vendors or users, that would be a totally normal product oversight. Room for improvement for the next version, but nothing catastrophic.
Shipping without any of them doesn't seem like lean product management, it seems like a careless attitude towards risk and a product design philosophy oriented entirely towards short-term demos, with no regard for how to realize actual productivity gains.
Furthermore, being available for years without anything like these features, despite hundreds of incidents demonstrating the risks, with hundreds of billions of dollars of funding, makes it seem to me like if they were to add all the features that would make their product actually safe and hypothetically useful, these features would reveal that it is actually not an improvement to productivity.
In the year since I first wrote about measuring the cost/benefit ratio of AI, I have heard from numerous people who have shown this to management to try to illustrate why their AI initiatives - like almost all AI initiatives - were either failing or burning out their engineers.
I've also heard from lots of people that have told me that it's obviously useful and they don't need to measure so carefully, because they are getting lots of work done that they couldn't have otherwise.7
I have yet to hear from a single person who has said "yeah, we measured according to your methodology8, and it turns out that our AI work is going great and that our ratio is 0.75".
Obviously, I cannot say for sure why this is; absence of evidence is not evidence of absence. But at this point I think the null hypothesis is that AI tools provide, in aggregate, zero value. They make mistakes too often, and the externalities they produce are so bad and so difficult to control that even before we get to the places where they are just physically poisoning people, even the negative effects on their direct users end up cancelling out whatever benefit to they provide to their organizations.
If I were wrong, then including tools to measure an AI's effectiveness at the tasks their users are actually trying to accomplish, rather than meaningless benchmarks, would show big productivity gains. The frontier labs would be champing at the bit to add such features, and crowing about their fantastic results.
I think the labs know that if they did that, it would present a grim picture to their users. Such tools would let their users see that it's making mistakes much more often than they realized, that they're spending much more time with it than they want to be, and that it's just generally not fit for purpose.
If they prove me wrong by adding in all of these safety mechanisms, and in the process, they make all of their AI technology less harmful, I'll be thrilled to be debunked.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
-
It is also interesting that for the next section, sometimes it seems that Claude's disclaimer is "Please double-check cited sources." instead. ↩
-
... by which I mean the "prompter", since authorship is not what's happening here. ↩
-
Claude has the "citations API", Google has various different kinds of "grounding" against its own APIs, and I guess Microsoft can check OpenAI's homework if you want. ↩
-
Given the relatively slow speed of the justice system and the mainstream press around the world, we probably will not hear about whether people are managing to incidentally break through these guard rails to self harm right now, but there are no shortage of stories still being reported right now where people were still doing just that, such as in this story where the effect of the much vaunted "guard rails" in 2025 was that if you wanted it to write you a suicide note, it would refuse twice but acquiesce on the third try. I don't see any reason to believe this fundamental issue has been addressed in the meanwhile, since it had been happening for years at that point. ↩
-
In this tutorial we can also see an incredibly rosy scenario presented, where a long-running workflow effortlessly compresses all of the necessary information into the new context, even if it uses a lower-fidelity model to do so, rather than the tangled and gnarly problem of problems which really are too big to fit in the context, which is to say, "most real-world problems". This presents the context limit instead as a minor speedbump to be worked around rather than the fundamental flaw in LLM tooling. ↩
-
A heavily fictionalized seaport. This is not how actual dockworkers work. This is a simplistic metaphor about incidental benefits of instrumental tasks, it is not supposed to delve deeply into the mechanics of maritime shipping. In particular I know that cranes are more reliable than this and this is not actually how you would respond to a crane malfunction anyway. Feel free to share fun facts about maritime shipping if that is your special interest but please do not @ me to correct this metaphor. ↩
-
To my knowledge, none of their publicly-traded employers have posted a measurable improvement to efficiency outside the margin of error. ↩
-
Or any similar methodology. I don't need people to adopt the exact practice that I proposed there. ↩
28 Sep 2026 4:27am GMT
25 Sep 2026
Planet Twisted
Glyph Lefkowitz: Who Is Open Source About?
Open source is, at least in part, about you, where "you" refers to the user.
Open Source Is Not About "Open Source Is Not About You"
In other words: Rich Hickey was wrong when he wrote "Open Source Is Not About You" and I'm tired of pretending otherwise.
Of course he's not completely wrong, or his famous post would not have resonated quite so much in the first place. Obnoxious users who demand their personal use-cases be immediately addressed by volunteer maintainers for free should indeed be viewed as the pariahs that they are. Similarly, corporate users who want free support from the community that supplies their infrastructure to lower their costs. As should those who profit from this type of externalization by their own customers.
But the exchange of "open source" (or even "free software") is not as simple as "I have prepared some software for you, please enjoy it, you have no right to complain", and maintainers ought to have a precise understanding of the costs and benefits - as well as the ethical implications - of that exchange.
Right now we barely even articulate that the exchange exists, let alone that it establishes a long-term, subtle, and implicit relationship between maintainer and user.
Let's fix that.
A Brief Aside about Meta-Ethics
When we talk about "obligations" and "rights", of "shoulds" and "musts", we are constructing an ethical system. The purpose of such a system is to develop social expectations and social consequences. There is not much use in me telling you that you are transcendentally evil for failing to follow some arbitrary recommendation that I have. But I am implying that I believe there should be consequences for your behavior. I am also implying that there probably already are some consequences, and they're just not written down anywhere yet.
Therefore, a post like this, where I say that we should view our social obligations in a certain way, that is the beginning of a broader social conversation. I think there should be some consequences, so I am gesturing towards that possibility. Exactly what consequences?
For now, I'm not sure. Let's figure it out.
What Are We Doing When We Do An Open Source?
Hickey, and his many acolytes in the years since his fateful post, asserts that the process of "open source" goes like this:
- Maintainer makes a thing, and makes it available to users as a gift.
- Maintainer may "love working with the team".
- Maintainer may be "proud of the work we do".
- Users accept the gift, and extract utility from it.
- (Users MUST be grateful for this.)
- A tiny fraction of users reciprocally contribute to the thing.
- (Maintainers may be grateful for this.)
He makes various oblique references to the specific activities of his company, which does things vaguely related to his projects for money1. These activities are exclusively characterized as for "customers", however, a subset of the aforementioned users so tiny ("fewer than 1%") as to nearly be an entirely distinct group.
Breezing past this process in an essay about obnoxious users demanding things they are not entitled to, one might nod along, as this sounds mostly sensible. Giving gifts is nice. I too love working with good teams and taking pride in things.
Examined more closely, however, it starts to logically fall apart. If you have consulting clients and that's where all of your money is coming from, why are you bothering (as he repeatedly insists) "doing [things] for the community"? What was the point of releasing this code in the first place? You could love working with your team and be proud of the work that you do in a lot of different contexts; why bother implicating this horde of entitled and obnoxious people, if that's all you're getting out of it? What's in it for you?
If we've left out something as fundamental as "why is the maintainer doing this", perhaps this story leaves out some other important bits as well.
Why Are You Doing This?
There are many possible motivations for releasing and maintaining open source software. They are often subtle, often overlapping, and rarely clearly stated. Maintainers are not a monolith and not everyone does it for similar reasons. But let's review a few reasons that someone might want to contribute.
Reputation
One reason that you might want to release some open source software is advertising. The most common form of this is self-promotion; if you are a visible, prominent contributor to an open source project, it stands to reason that you will have an easier time finding work in the domain of that project.
If you operate a consultancy, as Rich Hickey did at the time of his famous rant, then this reputational currency translates into advertising for your services. It's a practical demonstration of the skills of your team.
The trade in this benefit is most like the traditional "gift economy" that open source has been compared to. You give the code to your users, which has some value, but the users give you back some reputation, in the form of their attention, their esteem, and possibly even their money if they become customers or employers.
Influence
Infrastructure is the most popular type of open source for a good reason. Programmers working on a problem are often hemmed in by sclerotic architectural choices which prevent them from solving problems in the way that they'd prefer to solve them. Major infrastructural investments are difficult to justify in a planning process, as their benefits are hard to prove. Sometimes the benefits are highly personal; different engineers have different aesthetic preferences about what types of equally-valid solutions they'd prefer to work with.
If you can develop your preferred type of solution and release it as open source, then you can influence how everyone else solves this type of problem. As an individual, such a position of influence can allow you to have some transferable expertise between employers. You know how to use the tool you developed, so you can be very quick and effective with it, and you can shape it to your ongoing taste over time.
If you're an employer, and you can get everyone else to use your open source thing2, this can reduce both your hiring and training costs. Potential employees can read the code, see that it's good, and want to work at a place that produces good code like that. They can also read the code and become familiar with it in advance of coming to work for you, which means that you have a ready supply of developers who already know how your internal systems work.
The trade in this benefit is more like "soft power" than a gift economy. You give the code to your users, which has some value, but the users give you back the ability to dictate their technological agenda. You gain both the ability to influence their initial direction, and, as part of ongoing maintenance, to dictate their behavior over time.
Improvement
As an engineer, you might want to improve your own skills. Writing something proprietary and commercial cuts against this in two ways.
First, you will want to build something that already exists within your skill set, so that it will attract commercial interest and actually be competitive. Within the context of a larger team, you will want to personally be able to be immediately effective for similar reasons. But you still need a way to learn new things.
Second, you will want to build something somewhat secretively, so that the value you are producing is captured rather than released to the community. This means that you will be cut off from external sources of expert feedback.
As an organization, you might want to build the skills of your staff in similar ways.
The trade in this benefit is code for knowledge. You release the code or changes, and in return you expect your users to provide you good bug reports, and to induce at least some of them to become co-developers.
Outsourcing
As an engineer, you can only do so much on your own. Perhaps you want to have some influence over your infrastructure so you want to write it, but you also want to have a communal place to keep your infrastructure such that you can make a change to something to suit your needs, but you know that even if you walk away, someone else will maintain that change and keep it working across years or even decades of changes to underlying platforms, hardware, etc.
This sort of communal maintenance effort can be shared among all interested participants; if a thousand companies all need the same tool, if even a few dozen can share it, that reduces even their own load massively, let alone everyone else's.
The trade in this benefit is more complex, since there's less symmetry between the main maintainer and peripheral community members who also contribute code. The main maintainer is actually trading a namespace, a central place for people to contribute, coordinate, and release changes, rather than the code. They are a sort of market maker where then all the other contributors trade code for code within that market-ish structure.
In practice, this motivation produces a game theory problem where, when maintenance drops below a critical threshold, it creates a big enough crisis that at least some freeloading stakeholders will be forced to start making contributions.
Ultimately, however, this saves all involved parties a ton on maintenance, more eager volunteers who do not freeload in the first place get all the other benefits mentioned above as well.
A Brief Aside about your Chart of Accounts
Most companies account for open source maintenance work as simple overhead on ongoing projects. Sometimes it's CapEx, sometimes it's OpEx, but it's just "whoever happens to be working on this thing to support whatever random product it's a part of".
This type of accounting creates distorting incentives, because it doesn't recognize all the benefits above. Under such a fiscal regime, ongoing healthy maintenance becomes a ZIRP because when resources are more constrained, this apparent indulgence gets corrected.
The ancillary benefits that open source creates ought to be properly recognized. It shouldn't just be buried as Wages or IT or whatever. If it's helping you hire better engineers, some of that expense should be allocated to Recruitment Costs. If it's materially improving your reputation among your customer base, some of it should go to Goodwill. If it's getting your product in front of developers who are your customers, it should be in Marketing. Most importantly, if maintenance on an open source project is actually helping you maintain your enterprise-wide platform, it should not be squirreled away in some small team who happened to be the first one to adopt it.3
Exactly how these costs should be allocated and cross-charged to different departments depends heavily upon your organization and your specific chart of accounts. But "whatever, it's just part of the software product" or "I guess it's DevRel because the SDK is in there" is guaranteed to have your open source organization destroyed along with all those side-benefits the next time that there's a cash crunch.
The Things that Aren't Supposed To Be Benefits
These categories could be made as explicit, rational trade-offs, even if they are often implicit and subtle in practice. They are transactions where the maintainer gets something and the user gets something.
However, not everything that you are getting as a maintainer is something you are actually supposed to use to your own benefit. Being given trust in service of a responsibility is not a transaction.
"Oops, All Root Shells"
Open source code is code. In our modern world of absolutely pathetic sandboxing, installing code from somebody else gives them control over your system, even if it is somewhat indirect.
There is an unwritten rule that if I create an open source library, and you use it, it probably shouldn't have a backdoor in it that gives me the credentials to your bank account. There is a trust relationship between the user and the maintainer, and here, we see the first obligation that the maintainer has. The maintainer is obligated not to use the user's computer for their own gain.
This rule might seem obvious and straightforward. It might even seem unfair to you that I call the rule "unwritten", because the rule is, in fact, written down in a few places: for example, in the npm Acceptable Content Policy, it says right there:
A few examples of unacceptable content:
…
- Content containing malicious computer code, such as computer viruses, computer worms, rootkits, back doors, or spyware. This includes content submitted for research purposes. Tools designed and documented explicitly to assist in security research are acceptable, but exploits and malware that use the npm registry as a deployment or delivery vector are not.
I think we can all agree that a script which steals your bank credentials and sends them to me to buy a totally sick jet ski would qualify as "malware", so clearly that is forbidden.
There is also an enormous gray area here. npm also explicitly allows "Information on how to pay, donate to, and otherwise support Package development", but then goes on to explicitly forbid "Packages that display ads at runtime, on installation, or at other stages of the software development lifecycle, such as via npm scripts."4 How are the lines drawn around these gray areas? "npm will continue to apply its judgment when deciding what content is acceptable."
But also... this is forbidden by npm, not by the transcendental nature of "open source". I could give away code that displays all kinds of ads to its users as a "gift" on my website. The exact structure of this policy is not uncommon, but it also isn't exactly the same as other such sites. PyPI, for example, explicitly bans "cryptocurrency mining", which NPM does not. Is cryptocurrency mining "not open source"? A lot of judgement calls are happening here about what is allowable in these "gifts" that you are giving to your users.
But I digress.
My point is that policy-making around this concept is not clear, there are lots of little disagreements around the edges, but there is a very strong consensus that while the user is giving you their trust here, that is not a trade. The deal is not "you give the user some code, the user gives you unlimited compute and access to all their financial accounts". The user has made themselves vulnerable to your code on the strength of your reputation.
This creates an obligation for you to not do anything evil with that code, either intentionally or through negligence.
Security Updates Are Just Command And Control In A Funny Hat
All of this is just about the initial download of some code, and that is the way that Rich Hickey describes it, as if you just grabbed some code off a web page and put it in a folder that you like on your desktop. But that is not how open source relationships work today, if indeed it ever was.
The way it works today is that you add a dependency to your pyproject.toml or your package.json or your Cargo.toml and now your users are vulnerable not just to whatever you happened to upload in the first place, but to whoever happens to have your package index credentials.
This creates an obligation to maintain an operational security posture that protects your users from malicious updates.
The Roadmap Is Someone's Life
Another kind of trust that the user is placing in you is the trust that you are going to have at least some kind of regard for their usage of your software.
In a perfect world, the user's expectations could be clearly circumscribed. Whatever ongoing maintenance you commit to perform would be encapsulated in clear policies that you'd write up in advance, about exactly what kind of security response policy you have, how you will communicate when you no longer have the resources for maintenance, and so on.
But anyone who has been involved in any project at anything but the most extreme tier of operational maturity knows that 99% of the ecosystem relies on a set of loose conventions around how all that stuff works. We expect that maintainers will generally be around, that they'll use existing tools like an issue tracker for triaging user bugs, GHSA and CVEs for security reporting, that they will mark the project as "archived" and maybe do a final release before abandoning it, that they will maintain a ChangeLog explaining at least a little bit of what is going on.
Users assume that those conventions will be followed when there are any gaps in explicit policy, or indeed if policy is lacking entirely. This assumption is reasonable, because otherwise nobody could ever use any open source without a stack of service contracts that nobody has any time to write.
The strongest such convention is that an actively maintained program will, at least, more or less keep doing what it does as time goes on. A user who has elected to use a bit of open source software has made themselves vulnerable to changes and breakages in that software by the mere fact of using it. In the time that they have used it and invested in it, they have not invested in:
- creating alternative software to meet their needs,
- maintaining data in formats that other software can read, or
- learning how to use existing alternative software.
This can, and does, go badly wrong, when those expectations are mismatched.
How It Goes Wrong
Let's say a maintainer creates an open source paint program, OpenPaint.
An artist, known for their unique style of making blended collages, switches from their previous app, ProprietaryPaint, to this new OpenPaint to make these culturally significant works of art. However, the maintainer decides that the 'blend' tool is kind of a pain to maintain, and they remove it in OpenPaint 2.
A few months later, the artist's operating system vendor issues a security update that breaks OpenPaint, because older versions of OpenPaint were unknowingly abusing some platform API.
The maintainer releases a new OpenPaint 2.0.1 that addresses this incompatibility, but doesn't care about version 1.x any more so they don't bother to update that one.
This places the artist in an impossible situation. They can stay on an old version of their operating system, putting all their personal data at risk. Or they can upgrade to the new operating system, effectively either cutting off access to their livelihood, or forcing them to change their art style entirely.
Now, proprietary software can place users in similarly untenable positions (and in fact, it is more often proprietary software that does). But does the openness completely remove any obligation for this consideration? Should the OpenPaint team have to at least communicate the reasons for doing this, to give the artist some recourse?5
The only thing that "open source" does is that it allows the artist to pay a prohibitive amount of money to a new maintenance team to create a fork. This is rarely the kind of thing that individuals can manage.
This creates an obligation to at least consider how your users might be relying on you.
This is the most complex obligation of the bunch. Obviously it does not entitle every single user to infinite work from the maintainer, but it also shouldn't entitle the user to nothing for having trusted these subtle implied claims that the maintainer is making by making their work public.
It is a nuanced and ongoing negotiation and I do not think we have a clear moral intuition about how it should work out. But we do need to figure out a way to work it out.
It also raises a clarifying question.
Why Are We Even Doing This, and Who Are We Doing It For?
People generally like to do things for more than one reason. We live in an economy where people need to make money, but we mostly prefer to make that money doing things that are useful, and that make other people happy.
So, yes, we create open source for self-interested reasons to improve our reputations, to improve our skills, to increase our influence and to share our maintenance burdens. In so doing we take on some level of obligation to not abuse the trust that is placed in us, even if that level of obligation is not clear.
But if we are not doing it to serve those users at least a little bit, then those motivations are going to quickly ring hollow. We will not increase our reputation with a person if we respond to their every request by telling them that we owe them nothing and that their opinions are worthless. We will not gain influence over a community if we ignore their desires.
Many interactions with open source maintainers are unnecessarily adversarial. This is of course partially the fault of those users, who should calibrate their expectations appropriately.
Still: maintainers could do a better job of listening before these interactions become toxic. There's no reason that "open source users" should be an especially toxic group of people. At this point in history, that group is basically just … people with computers.
It's like that old truism. If you meet one person who is a jerk to you, that's their problem. But if everyone you meet, everywhere you go, is constantly abrasive to you and treats you like you're doing something wrong, maybe it's time to look inward.
If all open source users are entitled assholes, maybe it's time to look for a structural problem.
Surprise, It's About AI Again
Sigh.6
Users hate slop.
I know, dear AI-positive reader, your AI outputs are different from everyone else's, you aren't pushing thoughtless slop into your code, just because everyone else is and it is the inevitable terminus of using those tools. You aren't "lazy vibe coding" with Claude, you're doing "responsible agentic engineering", which is different because you're just built different.
Still, humor me, for a moment. Your users don't know that. They know what it looks like when products that they like adopt slop. They know that they will start leaking data. Developers know that it will make them personally less secure. They know that they can expect more outages and that your code will inexorably decline in quality.
In other words, your users are going to assume that this means you are violating that final obligation that the software should keep working.
Your users are going to tell you to stop, and they are probably going to get mad. Maybe you, or a plurality of your team, also want to stop, maybe you disagree with them, but in any case you need some way to have that conversation in a way that does not immediately overflow into every adjacent discussion forum. Users need to feel welcome in some space so they can have the discussion in that space, and not explode out into a thousand different group chats and social media threads.
This post was inspired by yet another prominent open source community discourse where a ton of angry users showed up to yell at developers to stop accepting LLM-generated code. I'm not going to link to any of these, because we don't need any more fuel for the discourse fire. But there is more than one such case and the pattern is becoming familiar.
On social media - usually BlueSky or Mastodon, but sometimes a user group forum - users become aware of some AI-adjacent policy. They show up in a horde to the developer forum or mailing list. They loudly start demanding the project take a hard stand7 against AI. This pressure is simultaneous, but uncoordinated; extremely repetitive, very diverse, often inconsistent, and pretty stressful, especially if you're a burnt-out maintainer with other things to be doing who may not even like AI yourself in the first place.
Believe me, I get it. It can be very unpleasant to deal with.
Like most problems that AI is causing, though, it's not really an "AI" problem as much as it is a pre-existing dumpster fire that "AI" is pouring gasoline onto. In this case, an online mob is the language of the unheard8.
If Users Are Mad It's Probably Already Too Late (But Maybe You Can Get Ready For Next Time)
One day, all of a sudden, you're getting feedback from a bunch of users that are using inappropriate channels to complain. But did they already have appropriate channels to use?
Did you have a place for people to congregate and discuss your project? To make orderly complaints in a way that will be legible to you? Or do you just have a GitHub Issues page, which non-technical users have no idea how to interact with, and a forum for developers, where users don't know the norms and any arriving brigade of pissed-off users will be seen as disruptive and inappropriate?
I don't want to be throwing any stones from within my particular glass house. Setting up such a place has gotten harder over the years. I don't really have one, either.
Could I have one, though? IRC has been slowly dying, mailing lists are unpopular and present increasingly annoying moderation challenges, forum software is expensive to operate and keep maintained, Discord is a confusing mess and the upshot of all of this is every community needs community management and forum moderation. Which means that for my own small solo projects, I couldn't possibly have such infrastructure because such infrastructure requires a dedicated second person to maintain it, and until someone volunteers for that, it's not really feasible. Even for my larger projects you'd be surprised how slim of a skeleton crew we are getting by with, and we definitely don't have a whole spare maintainer to go manage this, especially as we are under attack from the slopocalypse ourselves.
The nature of open source community is that most communities start too small to need such a thing, grow incrementally until one day they are suddenly way too big and needed one yesterday, and then suddenly they are too small again when interest wanes even a little bit. Even as we need it more and more, building and maintaining community infrastructure remains a challenge.
Even so, having a dedicated place for users - not maintainers - to converse amongst themselves, be an actual community, and present feedback to the developers, is fast becoming a necessary component of a successful community and not a nice-to-have.
In Conclusion
As trying as it can be sometimes, we maintainers all do get something out of open source, and it is good to be honest with your users - and with yourself - exactly what you want to get out of it. In order to know whether the juice is worth the squeeze, we must know both what the juice is, and what the squeeze is.
Part of the metaphorical squeeze is a set of obligations, and those are the most poorly defined of all. We should try to be clear about what those are too. Both about exactly what we believe we are signing up for, and also, about how we are willing to let our users hold us to account for them. Codes of conduct are a start here, but only the absolute barest bare minimum; "do not harass your colleagues or your users" is not a standard of excellence to aspire to, it's just basic manners.
I can't tell you exactly what your obligations are, only try to gesture at my idea of the outlines of the fuzzy moral intuition we've all been implicitly sharing up until now.
Drawing this line is not just for the benefit of the users, either. Maintainers already feel pressure, we already feel obligations. We resent that feeling of obligation. While there are a diverse array of reasons for that resentment, one big one is that it's not clear, even to ourselves where the obligations end. Lashing out by saying "I promised nothing and I owe you nothing!" followed by some choice expletives feels cathartic, but it doesn't really solve the problem, because we clearly don't really believe that's where the line is, or we would have already stopped there. We wouldn't feel the need to say it.
It is going to be a very big collective endeavor to figure out exactly where that line is. The best time to have gotten started on that endeavor was 50 years ago.
But the second best time is today.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!9
-
Somewhat to everyone's surprise, I, too, do things for money, like writing this post. Please remember to like and subscribe ↩
-
Whether it was originally yours, or developed by an employee who happened to be on staff at the time, or adopted by an employee who just started contributing to it a lot, in any of these scenarios a company can benefit from increased consistency and increased familiarity. ↩
-
If the rule is that they must forever endure the searing budgetary pain of gripping the white-hot potato that they unwittingly caught when they first made a good technical choice, this creates a perverse long-term incentive. ↩
-
I also find it darkly amusing that there is an explicit affordance here made for advertising, specifically, "Packages with code that can be used to display ads are fine. Packages that themselves display ads are not." This distinction rather gives the game away, that this is a website for carnies and not for marks, and that at some level we expect our users to deserve a lower level of respect than ourselves. But a full exploration of that is another blog post, or maybe a book, that I don't have time to write right now. ↩
-
If you want the turbocharged ultra-dramatic version of this problem, make it open source drivers for an optical prosthesis that lets the users see instead of an art app. That level of immediate physical dependency could be clarifying. It does also start to edge into an area where you could say that biomedical devices ought to be regulated differently, and that's not really a "software" problem but a "healthcare" problem and I'd mostly agree. Except for the fact that this is a very short distance away from breaking everyone's screen-reader with no notice or recourse. ↩
-
Did you believe I could write a blog post in 2026 which wasn't somehow about AI? I wish I could still believe that. ↩
-
It doesn't help that many of the most pro-AI voices are starting to have an, ahem, discernible political valence that is very unpopular among users. ↩
-
My apologies to MLK. ↩
-
If you read this whole post you can see that I sure need the help with all that. ↩
25 Sep 2026 12:50am GMT
06 Sep 2026
Planet Twisted
Glyph Lefkowitz: ... but what about video games?
I get asked this rhetorical question a lot, in various forms:
Sure, datacenters might use a lot of energy, but you don't have to use a hosted frontier model to do software development. What if I just run a local open-weights model to do some coding, with an open-source coding agent? Video games also use my GPU. Is local model development any worse than playing a video game?
So I want to write down my comprehensive answer to this: Yes, using an LLM to write some code is worse than playing a video game, for a few reasons.
Video Games Are Interactive, LLMs Are Batch Jobs
Video games use compute to respond to human input. You are using your GPU while you are looking at a screen, displaying an image. When you are done playing, you shut off the game, and your computer goes back to idle. It's much less energy. By contrast, agentic loops with evals (the only kind of "AI" that is meaningfully any good at coding) are running hot, for days. To use the most recent example of such a thing, a very rough first sketch of an implementation of a Windows graphics API backend to help port a paint program to other platforms, it took 3 weeks of Claude time, "day and night". Do you play a lot of video games for 500 hours to make it past the tutorial level, while also using other computers for other things, as well as the rest of your carbon footprint?
Video Games Need Development, LLMs Need Training
Video games use compute to respond to human input during development, too. Your game has to be made, but your LLM has to be trained. LLMs use a historically extreme amount of power, probably using more than the entire Internet, but it's kind of hard to say. Still, it seems a reasonable estimate to within several orders of magnitude that even over a multi-year project with hundreds of developers, the power used to develop an individual video game is nowhere close to training even a small LLM.
This is true even for local models. OpenAI has openly claimed that DeepSeek "stole its intellectual property", and I have heard grumblings that none of the open-weights generalist models could realistically exist without the massive lift that the frontier labs are doing with their training, in various other ways too. Secrecy throughout the industry makes this kind of impossible to understand rigorously, but it seems fair to say that you are partially culpable for all that famously energy-intensive frontier lab training if you're using a local model.
And They Keep Needing Training
You also can't dismiss this as a sunk cost, because in order to stay current with industry developments, models need to be updated with new information from the rest of the world, which means that you need to keep training them. Beyond the energy for your own use, if you want a real-life agentic workflow that actually does useful stuff, practically speaking you would still need to update your local models over and over again, at least once every few months, which means you would be incentivizing continued energy consumption by whoever was doing that training for you, including the energy cost of scraping.
Let's Be Real Here, You Aren't Actually Using A Local Model
This question is a hypothetical thought experiment. Despite synthetic benchmarks that keep showing there isn't much difference between open weight and frontier models, nobody's actually using local models for much of anything beyond sharing those talking points. Depending on which benchmark you're looking at, maybe it's good enough or maybe it's worse.
As an inveterate AI hater, all these systems seem pretty bad to me, but it seems that people who find them useful tend to subjectively believe the frontier models are worth the premium, and that's what they're actually using. Once you have accepted that it is OK to use LLMs for coding at all, it seems like a very quick slippery slope on down to "we'll go ahead and use the frontier models for now anyway, but we could be ethically better in the future by switching to an open weights one, that option is always available".
There's A Reason We Have Data Centers
Devolving power usage to local LLMs might be good to make users responsible for their costs and decrease the impacts to communities that are physically next to huge concentrations of power utilization, not to mention generation. However, there's a reason that it makes sense for the providers to build these giant facilities: economies of scale reduce total power consumption, they don't increase it. If you do all the same stuff with a local model that they have to do in hosted environments, it will probably take more power, even though you will be incentivized to do different stuff. This incentive to "do different stuff" is why although local models can hypothetically hold their own against the frontier labs for some tasks, when people or businesses take their inference costs in-house they often find that it's too painful and move back to hosted LLMs.
There Are Problems Other Than Power
These are subjects for a different post, but you have to consider a lot of other externalities: AI psychosis, de-skilling, comprehension debt, cultivating a dependency, introducing security defects, limiting your design space based on what LLMs can understand, context rot, wasting time on invalid solutions, introducing unpredictability into your workflows. You still have to consider the total cost benefit ratio.
To Sum Up
Local LLMs might alleviate some of the harms from using the hosted frontier providers. There are fewer privacy concerns, you can measure your power utilization and be more directly responsible for it, you can build interfaces with affordances that are less oriented towards addiction and dependency than the major frontier labs' harnesses.
But they're not automatically "the same as playing a video game" just because they can use the same GPU.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
06 Sep 2026 10:57pm GMT
06 Aug 2026
Planet Twisted
Hynek Schlawack: Production-ready Python Docker Containers with uv
Starting with 0.3.0, Astral's uv brought many great features, including support for cross-platform lock files uv.lock. Together with subsequent fixes, it has become Python's finest workflow tool for my (non-scientific) use cases. Here's how I build production-ready containers, as fast as possible.
06 Aug 2026 12:00am GMT
23 Jun 2026
Planet Twisted
Glyph Lefkowitz: Adversarial Communication
As I have discussed in previous posts, "AIs" can make mistakes. In fact, they do make mistakes, and their mistake-making patterns are such that where and how they will make mistakes is both uncertain and constantly changing.
Thus, in any scenario where you want to attempt to make "productive" use of "AI", you must have a system in place for checking every result. Not checking some results; checking every result. If each result might have a consequence for you (and if it didn't have a consequence, why bother automating it?) and you cannot predict in advance which kinds of results will need verification, then verification is always required.
The verification often ends up being just as expensive as doing the work in the first place, which means that if you want your usage of "AI" to be personally profitable, you have to find someone else to externalize the cost of verification onto. This person becomes your adversary, and, if you are successful, your "AI's" victim.
The Ladder-Climber And Their Reverse-Centaur Rungs
One way that this constellation of facts can straightforwardly assemble themselves into a dystopian nightmare is the phenomenon, described by Cory Doctorow, of the reverse centaur. This is when your employer non-consensually turns you into the verification system. The "AI" does the fun part of initially performing the work, and then you do the boring part where you check if the robot is right and clean up its messes, even if everyone already knows that it would, in aggregate, be cheaper for you to do the work in the first place.
Reverse centaurs can be made from any automation, not only "AI" automation. I think that there is a reason that this term happens to have emerged in the "age of AI", though, and not with earlier automation technologies (even those which were considerably more viscerally horrific). That reason is: the wrongness of "AI" output is not merely a technical feature that must be compensated for, it is a generalized externality.
As I mentioned above, if you are responsible for the entirety of the work, both extruding the "AI" output and checking it, it's usually cheaper to have humans do the entirety of the work to begin with. When humans do the writing directly, we can check as we go, and thus verification doesn't need to be as comprehensive.
When "AI" coding advocates say "code review is the bottleneck", what they are observing is that the LLM is still rolling the dice for each PR, and a human is still necessary to verify that each of those rolls is a winner. But calling this process "code review" is a bit of a misnomer; it's not really "code review" in the traditional sense, it's human understanding.
Before the advent of "AI", the human understanding was implicit in the process of writing the code in the first place1, and the code review was a way of diffusing and extending that understanding. Now that the code can be authored with no initial understanding taking place, that cost has not gone away, it has moved.
Human understanding was always the bottleneck.
However, this is taking a collaborative view of a software project, where satisfying the needs and solving the problems of your customers are the goals. We can see that "AI" is a bad tool to satisfy those goals, because all it's doing is converting the first half of the work, that of understanding the code as you write it, to understanding the agent's output as you read it.
What if, instead, we were to take the view that every software company is a Hobbesian nightmare, red in tooth and claw? In this view, the only goal of a software project is for the individual developers to make their promo cycles and get their bonuses. Given that there is only a certain amount of money to go around, this is a zero-sum game where each programmer wants to look more productive than their colleagues.
Pretty much every organization finds it easy to reward "productivity" as expressed by lines of code emitted, but the benefits of doing thorough and thoughtful design, analysis, and code review very difficult to reward. In this world, an LLM is an invaluable tool for the sociopathic ladder-climber, particularly if your legacy organization is still structuring their workflows as if the person prompting the bot is "writing" the code, and then they get to foist off the act of "reviewing" the code onto someone else.
Here, the prompter effectively externalizes the cost of the LLM's failures but internalizes any benefits. The prompter will vibe-code a big feature, so large that the assigned reviewer can't possibly comprehend it all effectively. When this happens, the reviewer will, eventually, be pressured to approve it, even if they can try to spot a few problems along the way. The reviewer has their own work to get back to, after all, the obligation to review the prompter's (read: the bot's) code is a drain on their time that they are not going to get rewarded for.
If this feature is a big success, the prompter gets a promotion. If it causes a big issue, well, the reviewer must not have been careful enough.
This is why LLMs are "good for coding", and also why their biggest promoters keep having outages.
The Generative Gish Galloper
Coding is the biggest "success story" of this type of adversarial communication, but it is by far not the only instance of such a thing. LLMs create a new form of leverage that can turn Brandolini's law from a linear advantage into an exponential one. If you are engaged in a political debate where you want to overwhelm the other side in nonsense, an LLM can generate bullshit faster than it is physically possible for a human being to type, let alone respond thoughtfully. There is an asymmetry to the utility of this weapon as well: only one side of the political spectrum wants to flood the zone and destroy trust in institutions and the concept of truth. There's a good reason that the fascists love it.
Straightforward Spam and Fraud
This is kind of obvious, but LLMs can generate lightly-customized, plausible-looking text much more quickly than any human being. This facilitates their use in fraud, spam, and scams. In a spamming or fraudulent interaction, once again, the costs are externalized onto the victim: the recipient of a spam message has to do all the work of "checking" the LLM's output. Spammers already expect very low hit rates from boilerplate, and if the LLM can increase those percentages from 1% to 5% the technology will pay for itself; they don't need anything like reliable accuracy.
Customer "Support"
If you have any kind of commercial relationship with a company, I probably don't even need to mention this: customer "support" bots are a misery. Everybody knows it at this point. But customer support is usually conceptualized by businesses as an adversarial interaction, because it is a cost center. They maintain internal metrics on time-to-resolution and try to optimize them. Implicitly, this creates a dynamic where the goal of the customer service agent's job is not to solve your problem, but to emit noise that will cause you to think your problem is resolved, or to give up, as fast as possible. Unsurprisingly, LLMs can emit this noise faster than humans can, getting those customers off the phone. But those customers will remember those interactions, and the story outside the TTR metrics is horrible.
Similarly to the situation in software development, LLMs can look very good on paper for customer support, but mostly what they are doing is illuminating the problems with the industry's existing metrics, by turning "winning the metrics battle against the customer" into a more obvious and immediate defeat for the company's long term reputation.
"Education"
In 2026 it is sadly a fact of life that students cheat all the time using "AI", and that this cheating is very successful, in that the teachers find it very hard to detect.
LLMs are great for cheating on schoolwork because the student is externalizing the work of the checking onto the teachers, who are often starting at a disadvantage to begin with, at least in the US.
My view is that this is happening because of a divergence in the way that students vs. teachers (or, more accurately, "the broader educational system") view grading.
When a student is asked to write an essay, the teachers see the effort as both intrinsically worthwhile for the student, as well as useful as a pedagogical tool to evaluate and react to the student's progress. The student, by contrast, sees a stumbling block designed to knock them off the path to success and into a permanent underclass. It is no wonder that the student sees "AI" as useful to their own goals and has no compunction about deploying it.
There is a bitter irony that the ability to understand the inherent value of actually writing the essay on their own is the sort of thing that students can really only learn by writing a bunch of essays. There's no way that I can think of which makes the benefit legible as long as a shortcut is available.
The net effect here is a downward spiral, where the already-wobbling educational system is sustaining an attack that it doesn't have the resources to recover from. The individual students' attacks against their teachers and their schools' grading systems might appear to momentarily succeed, but they will win the battle and lose the war.
Spamming "For Good"?
Usually when we talk about someone unilaterally choosing to enter into an adversarial relationship, that's an "attack" and for good reasons we have a negative impression of the attacker. However, I would be remiss if I did not point out that there are some cases where the relationship was already adversarial; just because you're the attacker doesn't mean that you are evil.
For example we might imagine use-cases like automatically filing appeals for prior authorizations against health insurance. It's relatively well-known at this point that the main way for-profit insurers maintain their margins is by denying claims right up to the line of the policies themselves being fraud, so using a spamming tool to fight them might be entirely justifiable2 in that case.
Similarly, using an LLM could be justified in a fight against a company refusing to honor a warranty. One could imagine using an LLM to immediately generate replies and escalations.
However, even in imagined cases like these, the underlying problem is that the insurers and the vendors already have a tremendous amount of structural power, so it is more likely that they will have the advantage in deploying a communications weapon like an LLM, as well as enacting policies to simply ignore any LLM-based communication that you might submit. Worse, if these strategies were to become widespread, they might provide an excuse to reject any communications by feeding them into an unreliable "LLM detector" and issuing an automated "computer says no" even to hand-written correspondence.
It is also worth stressing that these cases are imagined, as compared to the very real coworker-abuse, spam, scam, fraud, and disinformation campaigns being waged in real life today.
Therefore, while legitimate uses might exist, it's hard to imagine that there's anywhere they would be genuinely valuable and sustainable. In the best case "AI" will provide a temporary advantage for underdogs that will provoke an arms race which the resource-advantaged adversaries will win in the long run, in the worst case the arms race itself will cement permanent structural change that will make things worse.
"Search" By Stealing
Most of the adversarial utility of "AI" is on the "write" side, since write-amplification is more obviously aggressive than reading. But the "read" side of LLMs - summarization and question-answering - can be a form of attack as well.
To begin with, the act of reading itself is currently enormously destructive, but that's arguably not a fundamental aspect of this technology. They could set reasonable rate-limits and respect things like robots.txt, as search engines have for decades now. They could also refrain from committing criminal levels of copyright infringement. But, today, using "AI" tools does suborn this sort of out-of-control crawling.
More insidiously, consider the scenario described in this YouTube video. The LTT Bros decided to try Linux again, and in the course of so doing, they had problems. When trying to solve these problems, they were faced with a choice: they could consult Reddit, or they could ask an LLM. Asking an LLM would "gaslight the heck out of" them, but they still found it preferable, because they would at least get an answer without getting yelled at.
Initially this sounds great. But it also means that you want to extract knowledge from a community, while mechanically eliding any values or norms that the community may want to impart as part of offering that knowledge. As someone who spent many years in a community tech support role, this is worrying. Many requests for support are people asking how to do things that will momentarily solve a superficial problem but create a long-term reliability problem or even an immediate security risk, that the question-asker doesn't want to hear about. Consider the question "I'm tired of entering my password so much, how do I make it so my laptop unlocks automatically". An obsequious chatbot will helpfully tell you how to do this without pushback.
But, this is also a sort of ethically murky area. The Linux community is somewhat famously, for many years now, a toxic cesspool of general hostility, misogyny, etc. It is certainly a good thing that people can get access to this knowledge without subjecting themselves to abuse. But it also means that the people with the power and the privilege to change the community for the better can just quietly withdraw, rather than fixing the problems. It also means that the positive elements of culture cannot be transmitted, and people will have no opportunity to learn about unknown unknowns.
In this case, the "adversarial" communication is with society. The thing that using an LLM for search lets you do is withdraw from society and avoid forming any personal connections. There are some personal connections which are painful and annoying, and so that can feel like a momentary balm. But the need to make connections in general is, like, the concept of society itself.
Who Am I Hurting?
LLMs are good at adversarial communication. They are so good at it, relative to their other benefits, that they will tend to make communications adversarial if you are not remaining vigilant about the possibility that it might do so. My request to you, dear reader, if you are going to use such tools, is to always ask yourself, "who might I be hurting, if I use an LLM for this?"
If you're using an "AI", who is its adversary? If you haven't given it one yet, who might the "AI" turn into an adversary? Who might you overwhelm with an asymmetric amount of output, or, if you're receiving information and not sending it, who are you taking that information from without consulting?
Figure out the answers to these questions and conduct yourself accordingly; the answer might be "yourself".
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
-
One of the reasons that software developers tend to prefer greenfield development is that when you are given a blank page, you can project your own specific understanding onto it. You can structure the codebase in a way that works for your brain, down to the variable naming conventions and the module layouts. LLM-assisted development makes everything into instant brownfield work, which makes developers instantly miserable; even those who are excited about the technology will frequently complain about how it feels like their agency has been stolen and their joy in the work has been diminished. But I digress. ↩
-
Modulo the massive amount of other externalities involved in using LLMs, of course, but I don't have the time or energy to get into those here. ↩
23 Jun 2026 8:06pm GMT
09 Jun 2026
Planet Twisted
Hynek Schlawack: How to Ditch Codecov for Python Projects
Codecov's unreliability breaking CI on my open source projects has been a constant source of frustration for me for years. I have found a way to enforce coverage over a whole GitHub Actions build matrix that doesn't rely on third-party services.
09 Jun 2026 12:00am GMT
22 May 2026
Planet Twisted
Glyph Lefkowitz: Opaque Types in Python
Let's say you're writing a Python library.
In this library, you have some collection of state that represents "options" or "configuration" for a bunch of operations. Such a set of options is a bundle of potentially ever-increasing complexity. Thus, you will want it to have an extremely minimal compatibility surface, with a very carefully chosen public interface, that is either small, or perhaps nothing at all. Such an object conveys state and might have some private behavior, but all you want consumers to be able to do is build it in very constrained, specific ways, and then pass it along as a parameter to your own APIs.
By way of example, imagine that you're wrapping a library that handles shipping physical packages.
There are a zillion ways to do it ship a package. There are different carriers who can ship it for you. There's air freight, and ground freight, and sea freight. There's overnight shipping. There's the option to require a signature. There's package tracking and certified mail. Suffice it to say, lots of stuff.
If you are starting out to implement such a library, you might need an object called something like ShippingOptions that encapsulates some of this. At the core of your library you might have a function like this:
1 2 3 4 5 |
|
If you are starting out implementing such a library, you know that you're going to get the initial implementation of ShippingOptions wrong; or, at the very least, if not "wrong", then "incomplete". You should not want to commit to an expansive public API with a ton of different attributes until you really understand the problem domain pretty well.
Yet, ShippingOptions is absolutely vital to the rest of your library. You'll need to construct it and pass it to various methods like estimateShippingCost and shipPackage. So you're not going to want a ton of complexity and churn as you evolve it to be more complex.
Worse yet, this object has to hold a ton of state. It's got attributes, maybe even quite complex internal attributes that relate to different shipping services.
Right now, today, you need to add something so you can have "no rush", "standard" and "expedited" options. You can't just put off implementing that indefinitely until you can come up with the perfect shape. What to do?
The tool you want here is the opaque data type design pattern. C is lousy with such things (FILE, pthread_*_t, fd_set, etc). A typedef in a header file can easily achieve this.
But in Python, if you expose a dataclass - or any class, really - even if you keep all your fields private, the constructor is still, inherently, public. You can make it raise an exception or something, but your type checker still won't help your users; it'll still look like it's a normal class.
Luckily, Python typing provides a tool for this: typing.NewType.
Let's review our requirements:
- We need a type that our client code can use in its type annotations; it needs to be public.
- They need to be able to consruct it somehow, even if they shouldn't be able to see its attributes or its internal constructor arguments.
- To express high-level things (like "ship fast") that should stay supported as we add more nuanced and complex configurations in the future (like "ship with the fastest possible option provided by the lowest-cost carrier that supports signature verification").
In order to solve these problems respectively, we will use:
- a public
NewType, which gives us our public name... - which wraps a private class with entirely private attributes, to give us an actual data structure, while not exposing the constructor,
- a set of public constructor functions, which returns our
NewType.
When we put that all together, it looks like this:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 |
|
As a snapshot in time, this is not all that interesting; we could have just exposed _RealShipOpts as a public class and saved ourselves some time. The fact that this exposes a constructor that takes a string is not a big deal for the present moment. For an initial quick and dirty implementation, we can just do checks like if options._speed == "fast" in our shipping and estimation code.
However, the main thing we are doing here is preserving our flexibility to evolve the related APIs into the future, so let's see how we might do that. For example, let's allow the shipping options to contain a concrete and specific carrier and freight method:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 |
|
As a NewType, our public ShippingOptions type doesn't have a constructor. Since _RealShipOpts is private, and all its attributes are private, we can completely remove the old versions.
Anything within our shipping library can still access the private variables on ShippingOptions; as a NewType, it's the same type as its base at runtime, so it presents minimal1 overhead.
Clients outside our shipping library can still call all of our public constructors: shipFast, shipNormal, and shipSlow all still work with the same (as far as calling code knows) signature and behavior.
If you need to build and convey some state within your public API, while avoiding breakages associated with compatibility churn, hopefully this technique can help you do that!
Acknowledgments
Thanks for reading, and thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor.
-
The overhead is minimal, but it is not completely zero. The suggested idiom for converting to a
NewTypeis to call it like a function, as I've done in these examples, but if you are wanting to use this pattern inside of a hot loop, you can use# type: ignore[return-value]comments to avoid that small cost. ↩
22 May 2026 12:33am GMT
04 Apr 2026
Planet Twisted
Donovan Preston: Using osascript with terminal agents on macOS
Here is a useful trick that is unreasonably effective for simple computer use goals using modern terminal agents. On macOS, there has been a terminal osascript command since the original release of Mac OS X. All you have to do is suggest your agent use it and it can perform any application control action available in any AppleScript dictionary for any Mac app. No MCP set up or tools required at all. Agents are much more adapt at using rod terminal commands, especially ones that haven't changed in 30 years. Having a computer control interface that hasn't changed in 30 years and has extensive examples in the Internet corpus makes modern models understand how to use these tools basically Effortlessly. macOS locks down these permissions pretty heavily nowadays though, so you will have to grant the application control permission to terminal. But once you have done that, the range of possibilities for commanding applications using natural language is quite extensive. Also, for both Safari and chrome on Mac, you are going to want to turn on JavaScript over AppleScript permission. This basically allows claude or another agent to debug your web applications live for you as you are using them.In chrome, go to the view menu, developer submenu, and choose "Allow JavaScript from Apple events". In Safari, it's under the safari menu, settings, developer, "Allow JavaScript from Apple events". Then you can do something like "Hey Claude, would you Please use osascript to navigate the front chrome tab to hacker news". Once you suggest using OSA script in a session it will figure out pretty quickly what it can do with it. Of course you can ask it to do casual things like open your mail app or whatever. Then you can figure out what other things will work like please click around my web app or check the JavaScript Console for errors. Another very important tips for using modern agents is to try to practice using speech to text. I think speaking might be something like five times faster than typing. It takes a lot of time to get used to, especially after a lifetime of programming by typing, but it's a very interesting and a different experience and once you have a lot of practice It starts to to feel effortless.
04 Apr 2026 1:31pm GMT
16 Mar 2026
Planet Twisted
Donovan Preston: "Start Drag" and "Drop" to select text with macOS Voice Control
I have been using macOS voice control for about three years. First it was a way to reduce pain from excessive computer use. It has been a real struggle. Decades of computer use habits with typing and the mouse are hard to overcome! Text selection manipulation commands work quite well on macOS native apps like apps written in swift or safari with an accessibly tagged webpage. However, many webpages and electron apps (Visual Studio Code) have serious problems manipulating the selection, not working at all when using "select foo" where foo is a word in the text box to select, or off by one errors when manipulating the cursor position or extending the selection. I only recently expanded my repertoire with the "start drag" and "drop" commands, previously having used "Click and hold mouse", "move cursor to x", and "release mouse". Well, now I have discovered that using "start drag x" and "drop x" makes a fantastic text selection method! This is really going to improve my speed. In the long run, I believe computer voice control in general is going to end up being faster than WIMP, but for now the awkwardly rigid command phrasing and the amount of times it misses commands or misunderstands commands still really holds it back. I've been learning the macOS Voice Control specific command set for years now and I still reach for the keyboard and mouse way too often.
16 Mar 2026 11:04am GMT
04 Mar 2026
Planet Twisted
Glyph Lefkowitz: What Is Code Review For?
Humans Are Bad At Perceiving
Humans are not particularly good at catching bugs. For one thing, we get tired easily. There is some science on this, indicating that humans can't even maintain enough concentration to review more than about 400 lines of code at a time..
We have existing terms of art, in various fields, for the ways in which the human perceptual system fails to register stimuli. Perception fails when humans are distracted, tired, overloaded, or merely improperly engaged.
Each of these has implications for the fundamental limitations of code review as an engineering practice:
-
Inattentional Blindness: you won't be able to reliably find bugs that you're not looking for.
-
Repetition Blindness: you won't be able to reliably find bugs that you are looking for, if they keep occurring.
-
Vigilance Fatigue: you won't be able to reliably find either kind of bugs, if you have to keep being alert to the presence of bugs all the time.
-
and, of course, the distinct but related Alert Fatigue: you won't even be able to reliably evaluate reports of possible bugs, if there are too many false positives.
Never Send A Human To Do A Machine's Job
When you need to catch a category of error in your code reliably, you will need a deterministic tool to evaluate - and, thanks to our old friend "alert fatigue" above - ideally, to also remedy that type of error. These tools will relieve the need for a human to make the same repetitive checks over and over. None of them are perfect, but:
- to catch logical errors, use automated tests.
- to catch formatting errors, use autoformatters.
- to catch common mistakes, use linters.
- to catch common security problems, use a security scanner.
Don't blame reviewers for missing these things.
Code review should not be how you catch bugs.
What Is Code Review For, Then?
Code review is for three things.
First, code review is for catching process failures. If a reviewer has noticed a few bugs of the same type in code review, that's a sign that that type of bug is probably getting through review more often than it's getting caught. Which means it's time to figure out a way to deploy a tool or a test into CI that will reliably prevent that class of error, without requiring reviewers to be vigilant to it any more.
Second - and this is actually its more important purpose - code review is a tool for acculturation. Even if you already have good tools, good processes, and good documentation, new members of the team won't necessarily know about those things. Code review is an opportunity for older members of the team to introduce newer ones to existing tools, patterns, or areas of responsibility. If you're building an observer pattern, you might not realize that the codebase you're working in already has an existing idiom for doing that, so you wouldn't even think to search for it, but someone else who has worked more with the code might know about it and help you avoid repetition.
You will notice that I carefully avoided saying "junior" or "senior" in that paragraph. Sometimes the newer team member is actually more senior. But also, the acculturation goes both ways. This is the third thing that code review is for: disrupting your team's culture and avoiding stagnation. If you have new talent, a fresh perspective can also be an extremely valuable tool for building a healthy culture. If you're new to a team and trying to build something with an observer pattern, and this codebase has no tools for that, but your last job did, and it used one from an open source library, that is a good thing to point out in a review as well. It's an opportunity to spot areas for improvement to culture, as much as it is to spot areas for improvement to process.
Thus, code review should be as hierarchically flat as possible. If the goal of code review were to spot bugs, it would make sense to reserve the ability to review code to only the most senior, detail-oriented, rigorous engineers in the organization. But most teams already know that that's a recipe for brittleness, stagnation and bottlenecks. Thus, even though we know that not everyone on the team will be equally good at spotting bugs, it is very common in most teams to allow anyone past some fairly low minimum seniority bar to do reviews, often as low as "everyone on the team who has finished onboarding".
Oops, Surprise, This Post Is Actually About LLMs Again
Sigh. I'm as disappointed as you are, but there are no two ways about it: LLM code generators are everywhere now, and we need to talk about how to deal with them. Thus, an important corollary of this understanding that code review is a social activity, is that LLMs are not social actors, thus you cannot rely on code review to inspect their output.
My own personal preference would be to eschew their use entirely, but in the spirit of harm reduction, if you're going to use LLMs to generate code, you need to remember the ways in which LLMs are not like human beings.
When you relate to a human colleague, you will expect that:
- you can make decisions about what to focus on based on their level of experience and areas of expertise to know what problems to focus on; from a late-career colleague you might be looking for bad habits held over from legacy programming languages; from an earlier-career colleague you might be focused more on logical test-coverage gaps,
- and, they will learn from repeated interactions so that you can gradually focus less on a specific type of problem once you have seen that they've learned how to address it,
With an LLM, by contrast, while errors can certainly be biased a bit by the prompt from the engineer and pre-prompts that might exist in the repository, the types of errors that the LLM will make are somewhat more uniformly distributed across the experience range.
You will still find supposedly extremely sophisticated LLMs making extremely common mistakes, specifically because they are common, and thus appear frequently in the training data.
The LLM also can't really learn. An intuitive response to this problem is to simply continue adding more and more instructions to its pre-prompt, treating that text file as its "memory", but that just doesn't work, and probably never will. The problem - "context rot" is somewhat fundamental to the nature of the technology.
Thus, code-generators must be treated more adversarially than you would a human code review partner. When you notice it making errors, you always have to add tests to a mechanical, deterministic harness that will evaluates the code, because the LLM cannot meaningfully learn from its mistakes outside a very small context window in the way that a human would, so giving it feedback is unhelpful. Asking it to just generate the code again still requires you to review it all again, and as we have previously learned, you, a human, cannot review more than 400 lines at once.
To Sum Up
Code review is a social process, and you should treat it as such. When you're reviewing code from humans, share knowledge and encouragement as much as you share bugs or unmet technical requirements.
If you must reviewing code from an LLM, strengthen your automated code-quality verification tooling and make sure that its agentic loop will fail on its own when those quality checks fail immediately next time. Do not fall into the trap of appealing to its feelings, knowledge, or experience, because it doesn't have any of those things.
But for both humans and LLMs, do not fall into the trap of thinking that your code review process is catching your bugs. That's not its job.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
04 Mar 2026 5:24am GMT
19 Feb 2026
Planet Twisted
Donovan Preston: Wello Horld.
Onovanday Restonpay is going to logbay here again. It's time to take back the rss-source-rss-reader web of links
19 Feb 2026 2:36am GMT
05 Jan 2026
Planet Twisted
Glyph Lefkowitz: How To Argue With Me About AI, If You Must
As you already know if you've read any of this blog in the last few years, I am a somewhat reluctant - but nevertheless quite staunch - critic of LLMs. This means that I have enthusiasts of varying degrees sometimes taking issue with my stance.
It seems that I am not going to get away from discussions, and, let's be honest, pretty intense arguments about "AI" any time soon. These arguments are starting to make me quite upset. So it might be time to set some rules of engagement.
I've written about all of these before at greater length, but this is a short post because it's not about the technology or making a broader point, it's about me. These are rules for engaging with me, personally, on this topic. Others are welcome to adopt these rules if they so wish but I am not encouraging anyone to do so.
Thus, I've made this post as short as I can so everyone interested in engaging can read the whole thing. If you can't make it through to the end, then please just follow Rule Zero.
Rule Zero: Maybe Don't
You are welcome to ignore me. You can think my take is stupid and I can think yours is. We don't have to get into an Internet Fight about it; we can even remain friends. You do not need to instigate an argument with me at all, if you think that my analysis is so bad that it doesn't require rebutting.
Rule One: No 'Just'
As I explained in a post with perhaps the least-predictive title I've ever written, "I Think I'm Done Thinking About genAI For Now", I've already heard a bunch of bad arguments. Don't tell me to 'just' use a better model, use an agentic tool, use a more recent version, or use some prompting trick that you personally believe works better. If you skim my work and think that I must not have deeply researched anything or read about it because you don't like my conclusion, that is wrong.
Rule Two: No 'Look At This Cool Thing'
Purely as a productivity tool, I have had a terrible experience with genAI. Perhaps you have had a great one. Neat. That's great for you. As I explained at great length in "The Futzing Fraction", my concern with generative AI is that I believe it is probably a net negative impact on productivity, based on both my experience and plenty of citations. Go check out the copious footnotes if you're interested in more detail.
Therefore, I have already acknowledged that you can get an LLM to do various impressive, cool things, sometimes. If I tell you that you will, on average, lose money betting on a slot machine, a picture of a slot machine hitting a jackpot is not evidence against my position.
Rule Two And A Half: Engage In Metacognition
I specifically didn't title the previous rule "no anecdotes" because data beyond anecdotes may be extremely expensive to produce. I don't want to say you can never talk to me unless you're doing a randomized controlled trial. However, if you are going to tell me an anecdote about the way that you're using an LLM, I am interested in hearing how you are compensating for the well-documented biases that LLM use tends to induce. Try to measure what you can.
Rule Three: Do Not Cite The Deep Magic To Me
As I explained in "A Grand Unified Theory of the AI Hype Cycle", I already know quite a bit of history of the "AI" label. If you are tempted to tell me something about how "AI" is really such a broad field, and it doesn't just mean LLMs, especially if you are trying to launder the reputation of LLMs under the banner of jumbling them together with other things that have been called "AI", I assure you that this will not be convincing to me.
Rule Four: Ethics Are Not Optional
I have made several arguments in my previous writing: there are ethical arguments, efficacy arguments, structuralist arguments, efficiency arguments and aesthetic arguments.
I am happy to, for the purposes of a good-faith discussion, focus on a specific set of concerns or an individual point that you want to make where you think I got something wrong. If you convince me that I am entirely incorrect about the effectiveness or predictability of LLMs in general or as specific LLM product, you don't need to make a comprehensive argument about whether one should use the technology overall. I will even assume that you have your own ethical arguments.
However, if you scoff at the idea that one should have any ethical boundaries at all, and think that there's no reason to care about the overall utilitarian impact of this technology, that it's worth using no matter what else it does as long as it makes you 5% better at your job, that's sociopath behavior.
This includes extreme whataboutism regarding things like the water use of datacenters, other elements of the surveillance technology stack, and so on.
Consequences
These are rules, once again, just for engaging with me. I have no particular power to enact broader sanctions upon you, nor would I be inclined to do so if I could. However, if you can't stay within these basic parameters and you insist upon continuing to direct messages to me about this topic, I will summarily block you with no warning, on mastodon, email, GitHub, IRC, or wherever else you're choosing to do that. This is for your benefit as well: such a discussion will not be a productive use of either of our time.
05 Jan 2026 5:22am GMT
02 Jan 2026
Planet Twisted
Glyph Lefkowitz: The Next Thing Will Not Be Big
The dawning of a new year is an opportune moment to contemplate what has transpired in the old year, and consider what is likely to happen in the new one.
Today, I'd like to contemplate that contemplation itself.
The 20th century was an era characterized by rapidly accelerating change in technology and industry, creating shorter and shorter cultural cycles of changes in lifestyles. Thus far, the 21st century seems to be following that trend, at least in its recently concluded first quarter.
The early half of the twentieth century saw the massive disruption caused by electrification, radio, motion pictures, and then television.
In 1971, Intel poured gasoline on that fire by releasing the 4004, a microchip generally recognized as the first general-purpose microprocessor. Popular innovations rapidly followed: the computerized cash register, the personal computer, credit cards, cellular phones, text messaging, the Internet, the web, online games, mass surveillance, app stores, social media.
These innovations have arrived faster than previous generations, but also, they have crossed a crucial threshold: that of the human lifespan.
While the entire second millennium A.D. has been characterized by a gradually accelerating rate of technological and social change - the printing press and the industrial revolution were no slouches, in terms of changing society, and those predate the 20th century - most of those changes had the benefit of unfolding throughout the course of a generation or so.
Which means that any individual person in any given century up to the 20th might remember one major world-altering social shift within their lifetime, not five to ten of them. The diversity of human experience is vast, but most people would not expect that the defining technology of their lifetime was merely the latest in a progression of predictable civilization-shattering marvels.
Along with each of these successive generations of technology, we minted a new generation of industry titans. Westinghouse, Carnegie, Sarnoff, Edison, Ford, Hughes, Gates, Jobs, Zuckerberg, Musk. Not just individual rich people, but entire new classes of rich people that did not exist before. "Radio DJ", "Movie Star", "Rock Star", "Dot Com Founder", were all new paths to wealth opened (and closed) by specific technologies. While most of these people did come from at least some level of generational wealth, they no longer came from a literal hereditary aristocracy.
To describe this new feeling of constant acceleration, a new phrase was coined: "The Next Big Thing". In addition to denoting that some Thing was coming and that it would be Big (i.e.: that it would change a lot about our lives), this phrase also carries the strong implication that such a Thing would be a product. Not a development in social relationships or a shift in cultural values, but some new and amazing form of conveying salted meat paste or what-have-you, that would make whatever lucky tinkerer who stumbled into it into a billionaire - along with any friends and family lucky enough to believe in their vision and get in on the ground floor with an investment.
In the latter part of the 20th century, our entire model of capital allocation shifted to account for this widespread belief. No longer were mega-businesses built by bank loans, stock issuances, and reinvestment of profit, the new model was "Venture Capital". Venture capital is a model of capital allocation explicitly predicated on the idea that carefully considering each bet on a likely-to-succeed business and reducing one's risk was a waste of time, because the return on the equity from the Next Big Thing would be so disproportionately huge - 10x, 100x, 1000x - that one could afford to make at least 10 bad bets for each good one, and still come out ahead.
The biggest risk was in missing the deal, not in giving a bunch of money to a scam. Thus, value investing and focus on fundamentals have been broadly disregarded in favor of the pursuit of the Next Big Thing.
If Americans of the twentieth century were temporarily embarrassed millionaires, those of the twenty-first are all temporarily embarrassed FAANG CEOs.
The predicament that this tendency leaves us in today is that the world is increasingly run by generations - GenX and Millennials - with the shared experience that the computer industry, either hardware or software, would produce some radical innovation every few years. We assume that to be true.
But all things change, even change itself, and that industry is beginning to slow down. Physically, transistor density is starting to brush up against physical limits. Economically, most people are drowning in more compute power than they know what to do with anyway. Users already have most of what they need from the Internet.
The big new feature in every operating system is a bunch of useless junk nobody really wants and is seeing remarkably little uptake. Social media and smartphones changed the world, true, but… those are both innovations from 2008. They're just not new any more.
So we are all - collectively, culturally - looking for the Next Big Thing, and we keep not finding it.
It wasn't 3D printing. It wasn't crowdfunding. It wasn't smart watches. It wasn't VR. It wasn't the Metaverse, it wasn't Bitcoin, it wasn't NFTs1.
It's also not AI, but this is why so many people assume that it will be AI. Because it's got to be something, right? If it's got to be something then AI is as good a guess as anything else right now.
The fact is, our lifetimes have been an extreme anomaly. Things like the Internet used to come along every thousand years or so, and while we might expect that the pace will stay a bit higher than that, it is not reasonable to expect that something new like "personal computers" or "the Internet"3 will arrive again.
We are not going to get rich by getting in on the ground floor of the next Apple or the next Google because the next Apple and the next Google are Apple and Google. The industry is maturing. Software technology, computer technology, and internet technology are all maturing.
There Will Be Next Things
Research and development is happening in all fields all the time. Amazing new developments quietly and regularly occur in pharmaceuticals and in materials science. But these are not predictable. They do not inhabit the public consciousness until they've already happened, and they are rarely so profound and transformative that they change everybody's life.
There will even be new things in the computer industry, both software and hardware. Foldable phones do address a real problem (I wish the screen were even bigger but I don't want to carry around such a big device), and would probably be more popular if they got the costs under control. One day somebody's going to crack the problem of volumetric displays, probably. Some VR product will probably, eventually, hit a more realistic price/performance ratio where the niche will expand at least a little more.
Maybe there will even be something genuinely useful, which is recognizably adjacent to the current "AI" fad, but if it is, it will be some new development that we haven't seen yet. If current AI technology were sufficient to drive some interesting product, it would already be doing it, not using marketing disguised as science to conceal diminishing returns on current investments.
But They Will Not Be Big
The impulse to find the One Big Thing that will dominate the next five years is a fool's errand. Incremental gains are diminishing across the board. The markets for time and attention2 are largely saturated. There's no need for another streaming service if 100% of your leisure time is already committed to TikTok, YouTube and Netflix; famously, Netflix has already considered sleep its primary competitor for close to a decade - years before the pandemic.
Those rare tech markets which aren't saturated are suffering from pedestrian economic problems like wealth inequality, not technological bottlenecks.
For example, the thing preventing the development of a robot that can do your laundry and your dishes without your input is not necessarily that we couldn't build something like that, but that most households just can't afford it without wage growth catching up to productivity growth. It doesn't make sense for anyone to commit to the substantial R&D investment that such a thing would take, if the market doesn't exist because the average worker isn't paid enough to afford it on top of all the other tech which is already required to exist in society.
The projected income from the tiny, wealthy sliver of the population who could pay for the hardware, cannot justify an investment in the software past a fake version remotely operated by workers in the global south, only made possible by Internet wage arbitrage, i.e. a more palatable, modern version of indentured servitude.
Even if we were to accept the premise of an actually-"AI" version of this, that is still just a wish that ChatGPT could somehow improve enough behind the scenes to replace that worker, not any substantive investment in a novel, proprietary-to-the-chores-robot software system which could reliably perform specific functions.
What, Then?
The expectation for, and lack of, a "big thing" is a big problem. There are others who could describe its economic, political, and financial dimensions better than I can. So then let me speak to my expertise and my audience: open source software developers.
When I began my own involvement with open source, a big part of the draw for me was participating in a low-cost (to the corporate developer) but high-value (to society at large) positive externality. None of my employers would ever have cared about many of the applications for which Twisted forms a core bit of infrastructure; nor would I have been able to predict those applications' existence. Yet, it is nice to have contributed to their development, even a little bit.
However, it's not actually a positive externality if the public at large can't directly benefit from it.
When real world-changing, disruptive developments are occurring, the bean-counters are not watching positive externalities too closely. As we discovered with many of the other benefits that temporarily accrued to labor in the tech economy, Open Source that is usable by individuals and small companies may have been a ZIRP. If you know you're gonna make a billion dollars you're not going to worry about giving away a few hundred thousand here and there.
When gains are smaller and harder to realize, and margins are starting to get squeezed, it's harder to justify the investment in vaguely good vibes.
But this, itself, is not a call to action. I doubt very much that anyone reading this can do anything about the macroeconomic reality of higher interest rates. The technological reality of "development is happening slower" is inherently something that you can't change on purpose.
However, what we can do is to be aware of this trend in our own work.
Fight Scale Creep
It seems to me that more and more open source infrastructure projects are tools for hyper-scale application development, only relevant to massive cloud companies. This is just a subjective assessment on my part - I'm not sure what tools even exist today to measure this empirically - but I remember a big part of the open source community when I was younger being things like Inkscape, Themes.Org and Slashdot, not React, Docker Hub and Hacker News.
This is not to say that the hobbyist world no longer exists. There is of course a ton of stuff going on with Raspberry Pi, Home Assistant, OwnCloud, and so on. If anything there's a bit of a resurgence of self-hosting. But the interests of self-hosters and corporate developers are growing apart; there seems to be far less of a beneficial overflow from corporate infrastructure projects into these enthusiast or prosumer communities.
This is the concrete call to action: if you are employed in any capacity as an open source maintainer, dedicate more energy to medium- or small-scale open source projects.
If your assumption is that you will eventually reach a hyper-scale inflection point, then mimicking Facebook and Netflix is likely to be a good idea. However, if we can all admit to ourselves that we're not going to achieve a trillion-dollar valuation and a hundred thousand engineer headcount, we can begin to consider ways to make our Next Thing a bit smaller, and to accommodate the world as it is rather than as we wish it would be.
Be Prepared to Scale Down
Here are some design guidelines you might consider, for just about any open source project, particularly infrastructure ones:
-
Don't assume that your software can sustain an arbitrarily large fixed overhead because "you just pay that cost once" and you're going to be running a billion instances so it will always amortize; maybe you're only going to be running ten.
-
Remember that such fixed overhead includes not just CPU, RAM, and filesystem storage, but also the learning curve for developers. Front-loading a massive amount of conceptual complexity to accommodate the problems of hyper-scalers is a common mistake. Try to smooth out these complexities and introduce them only when necessary.
-
Test your code on edge devices. This means supporting Windows and macOS, and even Android and iOS. If you want your tool to help empower individual users, you will need to meet them where they are, which is not on an EC2 instance.
-
This includes considering Desktop Linux as a platform, as opposed to Server Linux as a platform, which (while they certainly have plenty in common) they are also distinct in some details. Consider the highly specific example of secret storage: if you are writing something that intends to live in a cloud environment, and you need to configure it with a secret, you will probably want to provide it via a text file or an environment variable. By contrast, if you want this same code to run on a desktop system, your users will expect you to support the Secret Service. This will likely only require a few lines of code to accommodate, but it is a massive difference to the user experience.
-
Don't rely on LLMs remaining cheap or free. If you have LLM-related features4, make sure that they are sufficiently severable from the rest of your offering that if ChatGPT starts costing $1000 a month, your tool doesn't break completely. Similarly, do not require that your users have easy access to half a terabyte of VRAM and a rack full of 5090s in order to run a local model.
Even if you were going to scale up to infinity, the ability to scale down and consider smaller deployments means that you can run more comfortably on, for example, a developer's laptop. So even if you can't convince your employer that this is where the economy and the future of technology in our lifetimes is going, it can be easy enough to justify this sort of design shift, particularly as individual choices. Make your onboarding cheaper, your development feedback loops tighter, and your systems generally more resilient to economic headwinds.
So, please design your open source libraries, applications, and services to run on smaller devices, with less complexity. It will be worth your time as well as your users'.
But if you can fix the whole wealth inequality thing, do that first.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
-
These sorts of lists are pretty funny reads, in retrospect. ↩
-
Which is to say, "distraction". ↩
-
... or even their lesser-but-still-profound aftershocks like "Social Media", "Smartphones", or "On-Demand Streaming Video" ... secondary manifestations of the underlying innovation of a packet-switched global digital network ... ↩
-
My preference would of course be that you just didn't have such features at all, but perhaps even if you agree with me, you are part of an organization with some mandate to implement LLM stuff. Just try not to wrap the chain of this anchor all the way around your code's neck. ↩
02 Jan 2026 1:59am GMT
11 Nov 2025
Planet Twisted
Glyph Lefkowitz: The “Dependency Cutout” Workflow Pattern, Part I
Tell me if you've heard this one before.
You're working on an application. Let's call it "FooApp". FooApp has a dependency on an open source library, let's call it "LibBar". You find a bug in LibBar that affects FooApp.
To envisage the best possible version of this scenario, let's say you actively like LibBar, both technically and socially. You've contributed to it in the past. But this bug is causing production issues in FooApp today, and LibBar's release schedule is quarterly. FooApp is your job; LibBar is (at best) your hobby. Blocking on the full upstream contribution cycle and waiting for a release is an absolute non-starter.
What do you do?
There are a few common reactions to this type of scenario, all of which are bad options.
I will enumerate them specifically here, because I suspect that some of them may resonate with many readers:
-
Find an alternative to LibBar, and switch to it.
This is a bad idea because a transition to a core infrastructure component could be extremely expensive.
-
Vendor LibBar into your codebase and fix your vendored version.
This is a bad idea because carrying this one fix now requires you to maintain all the tooling associated with a monorepo1: you have to be able to start pulling in new versions from LibBar regularly, reconcile your changes even though you now have a separate version history on your imported version, and so on.
-
Monkey-patch LibBar to include your fix.
This is a bad idea because you are now extremely tightly coupled to a specific version of LibBar. By modifying LibBar internally like this, you're inherently violating its compatibility contract, in a way which is going to be extremely difficult to test. You can test this change, of course, but as LibBar changes, you will need to replicate any relevant portions of its test suite (which may be its entire test suite) in FooApp. Lots of potential duplication of effort there.
-
Implement a workaround in your own code, rather than fixing it.
This is a bad idea because you are distorting the responsibility for correct behavior. LibBar is supposed to do LibBar's job, and unless you have a full wrapper for it in your own codebase, other engineers (including "yourself, personally") might later forget to go through the alternate, workaround codepath, and invoke the buggy LibBar behavior again in some new place.
-
Implement the fix upstream in LibBar anyway, because that's the Right Thing To Do, and burn credibility with management while you anxiously wait for a release with the bug in production.
This is a bad idea because you are betraying your users - by allowing the buggy behavior to persist - for the workflow convenience of your dependency providers. Your users are probably giving you money, and trusting you with their data. This means you have both ethical and economic obligations to consider their interests.
As much as it's nice to participate in the open source community and take on an appropriate level of burden to maintain the commons, this cannot sustainably be at the explicit expense of the population you serve directly.
Even if we only care about the open source maintainers here, there's still a problem: as you are likely to come under immediate pressure to ship your changes, you will inevitably relay at least a bit of that stress to the maintainers. Even if you try to be exceedingly polite, the maintainers will know that you are coming under fire for not having shipped the fix yet, and are likely to feel an even greater burden of obligation to ship your code fast.
Much as it's good to contribute the fix, it's not great to put this on the maintainers.
The respective incentive structures of software development - specifically, of corporate application development and open source infrastructure development - make options 1-4 very common.
On the corporate / application side, these issues are:
-
it's difficult for corporate developers to get clearance to spend even small amounts of their work hours on upstream open source projects, but clearance to spend time on the project they actually work on is implicit. If it takes 3 hours of wrangling with Legal2 and 3 hours of implementation work to fix the issue in LibBar, but 0 hours of wrangling with Legal and 40 hours of implementation work in FooApp, a FooApp developer will often perceive it as "easier" to fix the issue downstream.
-
it's difficult for corporate developers to get clearance from management to spend even small amounts of money sponsoring upstream reviewers, so even if they can find the time to contribute the fix, chances are high that it will remain stuck in review unless they are personally well-integrated members of the LibBar development team already.
-
even assuming there's zero pressure whatsoever to avoid open sourcing the upstream changes, there's still the fact inherent to any development team that FooApp's developers will be more familiar with FooApp's codebase and development processes than they are with LibBar's. It's just easier to work there, even if all other things are equal.
-
systems for tracking risk from open source dependencies often lack visibility into vendoring, particularly if you're doing a hybrid approach and only vendoring a few things to address work in progress, rather than a comprehensive and disciplined approach to a monorepo. If you fully absorb a vendored dependency and then modify it, Dependabot isn't going to tell you that a new version is available any more, because it won't be present in your dependency list. Organizationally this is bad of course but from the perspective of an individual developer this manifests mostly as fewer annoying emails.
But there are problems on the open source side as well. Those problems are all derived from one big issue: because we're often working with relatively small sums of money, it's hard for upstream open source developers to consume either money or patches from application developers. It's nice to say that you should contribute money to your dependencies, and you absolutely should, but the cost-benefit function is discontinuous. Before a project reaches the fiscal threshold where it can be at least one person's full-time job to worry about this stuff, there's often no-one responsible in the first place. Developers will therefore gravitate to the issues that are either fun, or relevant to their own job.
These mutually-reinforcing incentive structures are a big reason that users of open source infrastructure, even teams who work at corporate users with zillions of dollars, don't reliably contribute back.
The Answer We Want
All those options are bad. If we had a good option, what would it look like?
It is both practically necessary3 and morally required4 for you to have a way to temporarily rely on a modified version of an open source dependency, without permanently diverging.
Below, I will describe a desirable abstract workflow for achieving this goal.
Step 0: Report the Problem
Before you get started with any of these other steps, write up a clear description of the problem and report it to the project as an issue; specifically, in contrast to writing it up as a pull request. Describe the problem before submitting a solution.
You may not be able to wait for a volunteer-run open source project to respond to your request, but you should at least tell the project what you're planning on doing.
If you don't hear back from them at all, you will have at least made sure to comprehensively describe your issue and strategy beforehand, which will provide some clarity and focus to your changes.
If you do hear back from them, in the worst case scenario, you may discover that a hard fork will be necessary because they don't consider your issue valid, but even that information will save you time, if you know it before you get started. In the best case, you may get a reply from the project telling you that you've misunderstood its functionality and that there is already a configuration parameter or usage pattern that will resolve your problems with no new code. But in all cases, you will benefit from early coordination on what needs fixing before you get to how to fix it.
Step 1: Source Code and CI Setup
Fork the source code for your upstream dependency to a writable location where it can live at least for the duration of this one bug-fix, and possibly for the duration of your application's use of the dependency. After all, you might want to fix more than one bug in LibBar.
You want to have a place where you can put your edits, that will be version controlled and code reviewed according to your normal development process. This probably means you'll need to have your own main branch that diverges from your upstream's main branch.
Remember: you're going to need to deploy this to your production, so testing gates that your upstream only applies to final releases of LibBar will need to be applied to every commit here.
Depending on your LibBar's own development process, this may result in slightly unusual configurations where, for example, your fixes are written against the last LibBar release tag, rather than its current5 main; if the project has a branch-freshness requirement, you might need two branches, one for your upstream PR (based on main) and one for your own use (based on the release branch with your changes).
Ideally for projects with really good CI and a strong "keep main release-ready at all times" policy, you can deploy straight from a development branch, but it's good to take a moment to consider this before you get started. It's usually easier to rebase changes from an older HEAD onto a newer one than it is to go backwards.
Speaking of CI, you will want to have your own CI system. The fact that GitHub Actions has become a de-facto lingua franca of continuous integration means that this step may be quite simple, and your forked repo can just run its own instance.
Optional Bonus Step 1a: Artifact Management
If you have an in-house artifact repository, you should set that up for your dependency too, and upload your own build artifacts to it. You can often treat your modified dependency as an extension of your own source tree and install from a GitHub URL, but if you've already gone to the trouble of having an in-house package repository, you can pretend you've taken over maintenance of the upstream package temporarily (which you kind of have) and leverage those workflows for caching and build-time savings as you would with any other internal repo.
Step 2: Do The Fix
Now that you've got somewhere to edit LibBar's code, you will want to actually fix the bug.
Step 2a: Local Filesystem Setup
Before you have a production version on your own deployed branch, you'll want to test locally, which means having both repositories in a single integrated development environment.
At this point, you will want to have a local filesystem reference to your LibBar dependency, so that you can make real-time edits, without going through a slow cycle of pushing to a branch in your LibBar fork, pushing to a FooApp branch, and waiting for all of CI to run on both.
This is useful in both directions: as you prepare the FooApp branch that makes any necessary updates on that end, you'll want to make sure that FooApp can exercise the LibBar fix in any integration tests. As you work on the LibBar fix itself, you'll also want to be able to use FooApp to exercise the code and see if you've missed anything - and this, you wouldn't get in CI, since LibBar can't depend on FooApp itself.
In short, you want to be able to treat both projects as an integrated development environment, with support from your usual testing and debugging tools, just as much as you want your deployment output to be an integrated artifact.
Step 2b: Branch Setup for PR
However, for continuous integration to work, you will also need to have a remote resource reference of some kind from FooApp's branch to LibBar. You will need 2 pull requests: the first to land your LibBar changes to your internal LibBar fork and make sure it's passing its own tests, and then a second PR to switch your LibBar dependency from the public repository to your internal fork.
At this step it is very important to ensure that there is an issue filed on your own internal backlog to drop your LibBar fork. You do not want to lose track of this work; it is technical debt that must be addressed.
Until it's addressed, automated tools like Dependabot will not be able to apply security updates to LibBar for you; you're going to need to manually integrate every upstream change. This type of work is itself very easy to drop or lose track of, so you might just end up stuck on a vulnerable version.
Step 3: Deploy Internally
Now that you're confident that the fix will work, and that your temporarily-internally-maintained version of LibBar isn't going to break anything on your site, it's time to deploy.
Some deployment heritage should help to provide some evidence that your fix is ready to land in LibBar, but at the next step, please remember that your production environment isn't necessarily emblematic of that of all LibBar users.
Step 4: Propose Externally
You've got the fix, you've tested the fix, you've got the fix in your own production, you've told upstream you want to send them some changes. Now, it's time to make the pull request.
You're likely going to get some feedback on the PR, even if you think it's already ready to go; as I said, despite having been proven in your production environment, you may get feedback about additional concerns from other users that you'll need to address before LibBar's maintainers can land it.
As you process the feedback, make sure that each new iteration of your branch gets re-deployed to your own production. It would be a huge bummer to go through all this trouble, and then end up unable to deploy the next publicly released version of LibBar within FooApp because you forgot to test that your responses to feedback still worked on your own environment.
Step 4a: Hurry Up And Wait
If you're lucky, upstream will land your changes to LibBar. But, there's still no release version available. Here, you'll have to stay in a holding pattern until upstream can finalize the release on their end.
Depending on some particulars, it might make sense at this point to archive your internal LibBar repository and move your pinned release version to a git hash of the LibBar version where your fix landed, in their repository.
Before you do this, check in with the LibBar core team and make sure that they understand that's what you're doing and they don't have any wacky workflows which may involve rebasing or eliding that commit as part of their release process.
Step 5: Unwind Everything
Finally, you eventually want to stop carrying any patches and move back to an official released version that integrates your fix.
You want to do this because this is what the upstream will expect when you are reporting bugs. Part of the benefit of using open source is benefiting from the collective work to do bug-fixes and such, so you don't want to be stuck off on a pinned git hash that the developers do not support for anyone else.
As I said in step 2b6, make sure to maintain a tracking task for doing this work, because leaving this sort of relatively easy-to-clean-up technical debt lying around is something that can potentially create a lot of aggravation for no particular benefit. Make sure to put your internal LibBar repository into an appropriate state at this point as well.
Up Next
This is part 1 of a 2-part series. In part 2, I will explore in depth how to execute this workflow specifically for Python packages, using some popular tools. I'll discuss my own workflow, standards like PEP 517 and pyproject.toml, and of course, by the popular demand that I just know will come, uv.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
-
if you already have all the tooling associated with a monorepo, including the ability to manage divergence and reintegrate patches with upstream, you already have the higher-overhead version of the workflow I am going to propose, so, never mind. but chances are you don't have that, very few companies do. ↩
-
In any business where one must wrangle with Legal, 3 hours is a wildly optimistic estimate. ↩
-
In an ideal world every project would keep its main branch ready to release at all times, no matter what but we do not live in an ideal world. ↩
-
In this case, there is no question. It's 2b only, no not-2b. ↩
11 Nov 2025 1:44am GMT
15 Aug 2025
Planet Twisted
Glyph Lefkowitz: The Futzing Fraction
The most optimistic vision of generative AI1 is that it will relieve us of the tedious, repetitive elements of knowledge work so that we can get to work on the really interesting problems that such tedium stands in the way of. Even if you fully believe in this vision, it's hard to deny that today, some tedium is associated with the process of using generative AI itself.
Generative AI also isn't free, and so, as responsible consumers, we need to ask: is it worth it? What's the ROI of genAI, and how can we tell? In this post, I'd like to explore a logical framework for evaluating genAI expenditures, to determine if your organization is getting its money's worth.
Perpetually Proffering Permuted Prompts
I think most LLM users would agree with me that a typical workflow with an LLM rarely involves prompting it only one time and getting a perfectly useful answer that solves the whole problem.
Generative AI best practices, even from the most optimistic vendors all suggest that you should continuously evaluate everything. ChatGPT, which is really the only genAI product with significantly scaled adoption, still says at the bottom of every interaction:
ChatGPT can make mistakes. Check important info.
If we have to "check important info" on every interaction, it stands to reason that even if we think it's useful, some of those checks will find an error. Again, if we think it's useful, presumably the next thing to do is to perturb our prompt somehow, and issue it again, in the hopes that the next invocation will, by dint of either:
- better luck this time with the stochastic aspect of the inference process,
- enhanced application of our skill to engineer a better prompt based on the deficiencies of the current inference, or
- better performance of the model by populating additional context in subsequent chained prompts.
Unfortunately, given the relative lack of reliable methods to re-generate the prompt and receive a better answer2, checking the output and re-prompting the model can feel like just kinda futzing around with it. You try, you get a wrong answer, you try a few more times, eventually you get the right answer that you wanted in the first place. It's a somewhat unsatisfying process, but if you get the right answer eventually, it does feel like progress, and you didn't need to use up another human's time.
In fact, the hottest buzzword of the last hype cycle is "agentic". While I have my own feelings about this particular word3, its current practical definition is "a generative AI system which automates the process of re-prompting itself, by having a deterministic program evaluate its outputs for correctness".
A better term for an "agentic" system would be a "self-futzing system".
However, the ability to automate some level of checking and re-prompting does not mean that you can fully delegate tasks to an agentic tool, either. It is, plainly put, not safe. If you leave the AI on its own, you will get terrible results that will at best make for a funny story45 and at worst might end up causing serious damage67.
Taken together, this all means that for any consequential task that you want to accomplish with genAI, you need an expert human in the loop. The human must be capable of independently doing the job that the genAI system is being asked to accomplish.
When the genAI guesses correctly and produces usable output, some of the human's time will be saved. When the genAI guesses wrong and produces hallucinatory gibberish or even "correct" output that nevertheless fails to account for some unstated but necessary property such as security or scale, some of the human's time will be wasted evaluating it and re-trying it.
Income from Investment in Inference
Let's evaluate an abstract, hypothetical genAI system that can automate some work for our organization. To avoid implicating any specific vendor, let's call the system "Mallory".
Is Mallory worth the money? How can we know?
Logically, there are only two outcomes that might result from using Mallory to do our work.
- We prompt Mallory to do some work; we check its work, it is correct, and some time is saved.
- We prompt Mallory to do some work; we check its work, it fails, and we futz around with the result; this time is wasted.
As a logical framework, this makes sense, but ROI is an arithmetical concept, not a logical one. So let's translate this into some terms.
In order to evaluate Mallory, let's define the Futzing Fraction, " FF ", in terms of the following variables:
- H
-
the average amount of time a Human worker would take to do a task, unaided by Mallory
- I
-
the amount of time that Mallory takes to run one Inference8
- C
-
the amount of time that a human has to spend Checking Mallory's output for each inference
- P
-
the Probability that Mallory will produce a correct inference for each prompt
- W
-
the average amount of time that it takes for a human to Write one prompt for Mallory
- E
-
since we are normalizing everything to time, rather than money, we do also have to account for the dollar of Mallory as as a product, so we will include the Equivalent amount of human time we could purchase for the marginal cost of one9 inference.
As in last week's example of simple ROI arithmetic, we will put our costs in the numerator, and our benefits in the denominator.
The idea here is that for each prompt, the minimum amount of time-equivalent cost possible is W+I+C+E. The user must, at least once, write a prompt, wait for inference to run, then check the output; and, of course, pay any costs to Mallory's vendor.
If the probability of a correct answer is P=13, then they will do this entire process 3 times10, so we put P in the denominator. Finally, we divide everything by H, because we are trying to determine if we are actually saving any time or money, versus just letting our existing human, who has to be driving this process anyway, do the whole thing.
If the Futzing Fraction evaluates to a number greater than 1, as previously discussed, you are a bozo; you're spending more time futzing with Mallory than getting value out of it.
Figuring out the Fraction is Frustrating
In order to even evaluate the value of the Futzing Fraction though, you have to have a sound method to even get a vague sense of all the terms.
If you are a business leader, a lot of this is relatively easy to measure. You vaguely know what H is, because you know what your payroll costs, and similarly, you can figure out E with some pretty trivial arithmetic based on Mallory's pricing table. There are endless YouTube channels, spec sheets and benchmarks to give you I. W is probably going to be so small compared to H that it hardly merits consideration11.
But, are you measuring C? If your employees are not checking the outputs of the AI, you're on a path to catastrophe that no ROI calculation can capture, so it had better be greater than zero.
Are you measuring P? How often does the AI get it right on the first try?
Challenges to Computing Checking Costs
In the fraction defined above, the term C is going to be large. Larger than you think.
Measuring P and C with a high degree of precision is probably going to be very hard; possibly unreasonably so, or too expensive12 to bother with in practice. So you will undoubtedly need to work with estimates and proxy metrics. But you have to be aware that this is a problem domain where your normal method of estimating is going to be extremely vulnerable to inherent cognitive bias, and find ways to measure.
Margins, Money, and Metacognition
First let's discuss cognitive and metacognitive bias.
My favorite cognitive bias is the availability heuristic and a close second is its cousin salience bias. Humans are empirically predisposed towards noticing and remembering things that are more striking, and to overestimate their frequency.
If you are estimating the variables above based on the vibe that you're getting from the experience of using an LLM, you may be overestimating its utility.
Consider a slot machine.
If you put a dollar in to a slot machine, and you lose that dollar, this is an unremarkable event. Expected, even. It doesn't seem interesting. You can repeat this over and over again, a thousand times, and each time it will seem equally unremarkable. If you do it a thousand times, you will probably get gradually more anxious as your sense of your dwindling bank account becomes slowly more salient, but losing one more dollar still seems unremarkable.
If you put a dollar in a slot machine and it gives you a thousand dollars, that will probably seem pretty cool. Interesting. Memorable. You might tell a story about this happening, but you definitely wouldn't really remember any particular time you lost one dollar.
Luckily, when you arrive at a casino with slot machines, you probably know well enough to set a hard budget in the form of some amount of physical currency you will have available to you. The odds are against you, you'll probably lose it all, but any responsible gambler will have an immediate, physical representation of their balance in front of them, so when they have lost it all, they can see that their hands are empty, and can try to resist the "just one more pull" temptation, after hitting that limit.
Now, consider Mallory.
If you put ten minutes into writing a prompt, and Mallory gives a completely off-the-rails, useless answer, and you lose ten minutes, well, that's just what using a computer is like sometimes. Mallory malfunctioned, or hallucinated, but it does that sometimes, everybody knows that. You only wasted ten minutes. It's fine. Not a big deal. Let's try it a few more times. Just ten more minutes. It'll probably work this time.
If you put ten minutes into writing a prompt, and it completes a task that would have otherwise taken you 4 hours, that feels amazing. Like the computer is magic! An absolute endorphin rush.
Very memorable. When it happens, it feels like P=1.
But... did you have a time budget before you started? Did you have a specified N such that "I will give up on Mallory as soon as I have spent N minutes attempting to solve this problem with it"? When the jackpot finally pays out that 4 hours, did you notice that you put 6 hours worth of 10-minute prompt coins into it in?
If you are attempting to use the same sort of heuristic intuition that probably works pretty well for other business leadership decisions, Mallory's slot-machine chat-prompt user interface is practically designed to subvert those sensibilities. Most business activities do not have nearly such an emotionally variable, intermittent reward schedule. They're not going to trick you with this sort of cognitive illusion.
Thus far we have been talking about cognitive bias, but there is a metacognitive bias at play too: while Dunning-Kruger, everybody's favorite metacognitive bias does have some problems with it, the main underlying metacognitive bias is that we tend to believe our own thoughts and perceptions, and it requires active effort to distance ourselves from them, even if we know they might be wrong.
This means you must assume any intuitive estimate of C is going to be biased low; similarly P is going to be biased high. You will forget the time you spent checking, and you will underestimate the number of times you had to re-check.
To avoid this, you will need to decide on a Ulysses pact to provide some inputs to a calculation for these factors that you will not be able to able to fudge if they seem wrong to you.
Problematically Plausible Presentation
Another nasty little cognitive-bias landmine for you to watch out for is the authority bias, for two reasons:
- People will tend to see Mallory as an unbiased, external authority, and thereby see it as more of an authority than a similarly-situated human13.
- Being an LLM, Mallory will be overconfident in its answers14.
The nature of LLM training is also such that commonly co-occurring tokens in the training corpus produce higher likelihood of co-occurring in the output; they're just going to be closer together in the vector-space of the weights; that's, like, what training a model is, establishing those relationships.
If you've ever used an heuristic to informally evaluate someone's credibility by listening for industry-specific shibboleths or ways of describing a particular issue, that skill is now useless. Having ingested every industry's expert literature, commonly-occurring phrases will always be present in Mallory's output. Mallory will usually sound like an expert, but then make mistakes at random.15.
While you might intuitively estimate C by thinking "well, if I asked a person, how could I check that they were correct, and how long would that take?" that estimate will be extremely optimistic, because the heuristic techniques you would use to quickly evaluate incorrect information from other humans will fail with Mallory. You need to go all the way back to primary sources and actually fully verify the output every time, or you will likely fall into one of these traps.
Mallory Mangling Mentorship
So far, I've been describing the effect Mallory will have in the context of an individual attempting to get some work done. If we are considering organization-wide adoption of Mallory, however, we must also consider the impact on team dynamics. There are a number of possible potential side effects that one might consider when looking at, but here I will focus on just one that I have observed.
I have a cohort of friends in the software industry, most of whom are individual contributors. I'm a programmer who likes programming, so are most of my friends, and we are also (sigh), charitably, pretty solidly middle-aged at this point, so we tend to have a lot of experience.
As such, we are often the folks that the team - or, in my case, the community - goes to when less-experienced folks need answers.
On its own, this is actually pretty great. Answering questions from more junior folks is one of the best parts of a software development job. It's an opportunity to be helpful, mostly just by knowing a thing we already knew. And it's an opportunity to help someone else improve their own agency by giving them knowledge that they can use in the future.
However, generative AI throws a bit of a wrench into the mix.
Let's imagine a scenario where we have 2 developers: Alice, a staff engineer who has a good understanding of the system being built, and Bob, a relatively junior engineer who is still onboarding.
The traditional interaction between Alice and Bob, when Bob has a question, goes like this:
- Bob gets confused about something in the system being developed, because Bob's understanding of the system is incorrect.
- Bob formulates a question based on this confusion.
- Bob asks Alice that question.
- Alice knows the system, so she gives an answer which accurately reflects the state of the system to Bob.
- Bob's understanding of the system improves, and thus he will have fewer and better-informed questions going forward.
You can imagine how repeating this simple 5-step process will eventually transform Bob into a senior developer, and then he can start answering questions on his own. Making sufficient time for regularly iterating this loop is the heart of any good mentorship process.
Now, though, with Mallory in the mix, the process now has a new decision point, changing it from a linear sequence to a flow chart.
We begin the same way, with steps 1 and 2. Bob's confused, Bob formulates a question, but then:
- Bob asks Mallory that question.
Here, our path then diverges into a "happy" path, a "meh" path, and a "sad" path.
The "happy" path proceeds like so:
- Mallory happens to formulate a correct answer.
- Bob's understanding of the system improves, and thus he will have fewer and better-informed questions going forward.
Great. Problem solved. We just saved some of Alice's time. But as we learned earlier,
Mallory can make mistakes. When that happens, we will need to check important info. So let's get checking:
- Mallory happens to formulate an incorrect answer.
- Bob investigates this answer.
- Bob realizes that this answer is incorrect because it is inconsistent with some of his prior, correct knowledge of the system, or his investigation.
- Bob asks Alice the same question; GOTO traditional interaction step 4.
On this path, Bob spent a while futzing around with Mallory, to no particular benefit. This wastes some of Bob's time, but then again, Bob could have ended up on the happy path, so perhaps it was worth the risk; at least Bob wasn't wasting any of Alice's much more valuable time in the process.16
Notice that beginning at the start of step 4, we must begin allocating all of Bob's time to C, so C already starts getting a bit bigger than if it were just Bob checking Mallory's output specifically on tasks that Bob is doing.
That brings us to the "sad" path.
- Mallory happens to formulate an incorrect answer.
- Bob investigates this answer.
- Bob does not realize that this answer is incorrect because he is unable to recognize any inconsistencies with his existing, incomplete knowledge of the system.
- Bob integrates Mallory's incorrect information of the system into his mental model.
- Bob proceeds to make a larger and larger mess of his work, based on an incorrect mental model.
- Eventually, Bob asks Alice a new, worse question, based on this incorrect understanding.
- Sadly we cannot return to the happy path at this point, because now Alice must unravel the complex series of confusing misunderstandings that Mallory has unfortunately conveyed to Bob at this point. In the really sad case, Bob actually doesn't believe Alice for a while, because Mallory seems unbiased17, and Alice has to waste even more time convincing Bob before she can simply explain to him.
Now, we have wasted some of Bob's time, and some of Alice's time. Everything from step 5-10 is C, and as soon as Alice gets involved, we are now adding to C at double real-time. If more team members are pulled in to the investigation, you are now multiplying C by the number of investigators, potentially running at triple or quadruple real time.
But That's Not All
Here I've presented a brief selection reasons why C will be both large, and larger than you expect. To review:
- Gambling-style mechanics of the user interface will interfere with your own self-monitoring and developing a good estimate.
- You can't use human heuristics for quickly spotting bad answers.
- Wrong answers given to junior people who can't evaluate them will waste more time from your more senior employees.
But this is a small selection of ways that Mallory's output can cost you money and time. It's harder to simplistically model second-order effects like this, but there's also a broad range of possibilities for ways that, rather than simply checking and catching errors, an error slips through and starts doing damage. Or ways in which the output isn't exactly wrong, but still sub-optimal in ways which can be difficult to notice in the short term.
For example, you might successfully vibe-code your way to launch a series of applications, successfully "checking" the output along the way, but then discover that the resulting code is unmaintainable garbage that prevents future feature delivery, and needs to be re-written18. But this kind of intellectual debt isn't even specific to technical debt while coding; it can even affect such apparently genAI-amenable fields as LinkedIn content marketing19.
Problems with the Prediction of P
C isn't the only challenging term though. P, is just as, if not more important, and just as hard to measure.
LLM marketing materials love to phrase their accuracy in terms of a percentage. Accuracy claims for LLMs in general tend to hover around 70%20. But these scores vary per field, and when you aggregate them across multiple topic areas, they start to trend down. This is exactly why "agentic" approaches for more immediately-verifiable LLM outputs (with checks like "did the code work") got popular in the first place: you need to try more than once.
Independently measured claims about accuracy tend to be quite a bit lower21. The field of AI benchmarks is exploding, but it probably goes without saying that LLM vendors game those benchmarks22, because of course every incentive would encourage them to do that. Regardless of what their arbitrary scoring on some benchmark might say, all that matters to your business is whether it is accurate for the problems you are solving, for the way that you use it. Which is not necessarily going to correspond to any benchmark. You will need to measure it for yourself.
With that goal in mind, our formulation of P must be a somewhat harsher standard than "accuracy". It's not merely "was the factual information contained in any generated output accurate", but, "is the output good enough that some given real knowledge-work task is done and the human does not need to issue another prompt"?
Surprisingly Small Space for Slip-Ups
The problem with reporting these things as percentages at all, however, is that our actual definition for P is 1attempts, where attempts for any given attempt, at least, must be an integer greater than or equal to 1.
Taken in aggregate, if we succeed on the first prompt more often than not, we could end up with a P>12, but combined with the previous observation that you almost always have to prompt it more than once, the practical reality is that P will start at 50% and go down from there.
If we plug in some numbers, trying to be as extremely optimistic as we can, and say that we have a uniform stream of tasks, every one of which can be addressed by Mallory, every one of which:
- we can measure perfectly, with no overhead
- would take a human 45 minutes
- takes Mallory only a single minute to generate a response
- Mallory will require only 1 re-prompt, so "good enough" half the time
- takes a human only 5 minutes to write a prompt for
- takes a human only 5 minutes to check the result of
- has a per-prompt cost of the equivalent of a single second of a human's time
Thought experiments are a dicey basis for reasoning in the face of disagreements, so I have tried to formulate something here that is absolutely, comically, over-the-top stacked in favor of the AI optimist here.
Would that be a profitable? It sure seems like it, given that we are trading off 45 minutes of human time for 1 minute of Mallory-time and 10 minutes of human time. If we ask Python:
1 2 3 4 5 |
|
We get a futzing fraction of about 0.4893. Not bad! Sounds like, at least under these conditions, it would indeed be cost-effective to deploy Mallory. But… realistically, do you reliably get useful, done-with-the-task quality output on the second prompt? Let's bump up the denominator on P just a little bit there, and see how we fare:
1 2 |
|
Oof. Still cost-effective at 0.734, but not quite as good. Where do we cap out, exactly?
1 2 3 4 5 6 7 8 9 |
|
With this little test, we can see that at our next iteration we are already at 0.9792, and by 5 tries per prompt, even in this absolute fever-dream of an over-optimistic scenario, with a futzing fraction of 1.2240, Mallory is now a net detriment to our bottom line.
Harm to the Humans
We are treating H as functionally constant so far, an average around some hypothetical Gaussian distribution, but the distribution itself can also change over time.
Formally speaking, an increase to H would be good for our fraction. Maybe it would even be a good thing; it could mean we're taking on harder and harder tasks due to the superpowers that Mallory has given us.
But an observed increase to H would probably not be good. An increase could also mean your humans are getting worse at solving problems, because using Mallory has atrophied their skills23 and sabotaged learning opportunities2425. It could also go up because your senior, experienced people now hate their jobs26.
For some more vulnerable folks, Mallory might just take a shortcut to all these complex interactions and drive them completely insane27 directly. Employees experiencing an intense psychotic episode are famously less productive than those who are not.
This could all be very bad, if our futzing fraction eventually does head north of 1 and you need to reconsider introducing human-only workflows, without Mallory.
Abridging the Artificial Arithmetic (Alliteratively)
To reiterate, I have proposed this fraction:
which shows us positive ROI when FF is less than 1, and negative ROI when it is more than 1.
This model is heavily simplified. A comprehensive measurement program that tests the efficacy of any technology, let alone one as complex and rapidly changing as LLMs, is more complex than could be captured in a single blog post.
Real-world work might be insufficiently uniform to fit into a closed-form solution like this. Perhaps an iterated simulation with variables based on the range of values seem from your team's metrics would give better results.
However, in this post, I want to illustrate that if you are going to try to evaluate an LLM-based tool, you need to at least include some representation of each of these terms somewhere. They are all fundamental to the way the technology works, and if you're not measuring them somehow, then you are flying blind into the genAI storm.
I also hope to show that a lot of existing assumptions about how benefits might be demonstrated, for example with user surveys about general impressions, or by evaluating artificial benchmark scores, are deeply flawed.
Even making what I consider to be wildly, unrealistically optimistic assumptions about these measurements, I hope I've shown:
- in the numerator, C might be a lot higher than you expect,
- in the denominator, P might be a lot lower than you expect,
- repeated use of an LLM might make H go up, but despite the fact that it's in the denominator, that will ultimately be quite bad for your business.
Personally, I don't have all that many concerns about E and I. E is still seeing significant loss-leader pricing, and I might not be coming down as fast as vendors would like us to believe, if the other numbers work out I don't think they make a huge difference. However, there might still be surprises lurking in there, and if you want to rationally evaluate the effectiveness of a model, you need to be able to measure them and incorporate them as well.
In particular, I really want to stress the importance of the influence of LLMs on your team dynamic, as that can cause massive, hidden increases to C. LLMs present opportunities for junior employees to generate an endless stream of chaff that will simultaneously:
- wreck your performance review process by making them look much more productive than they are,
- increase stress and load on senior employees who need to clean up unforeseen messes created by their LLM output,
- and ruin their own opportunities for career development by skipping over learning opportunities.
If you've already deployed LLM tooling without measuring these things and without updating your performance management processes to account for the strange distortions that these tools make possible, your Futzing Fraction may be much, much greater than 1, creating hidden costs and technical debt that your organization will not notice until a lot of damage has already been done.
If you got all the way here, particularly if you're someone who is enthusiastic about these technologies, thank you for reading. I appreciate your attention and I am hopeful that if we can start paying attention to these details, perhaps we can all stop futzing around so much with this stuff and get back to doing real work.
Acknowledgments
Thank you to my patrons who are supporting my writing on this blog. If you like what you've read here and you'd like to read more of it, or you'd like to support my various open-source endeavors, you can support my work as a sponsor!
-
I do not share this optimism, but I want to try very hard in this particular piece to take it as a given that genAI is in fact helpful. ↩
-
If we could have a better prompt on demand via some repeatable and automatable process, surely we would have used a prompt that got the answer we wanted in the first place. ↩
-
The software idea of a "user agent" straightforwardly comes from the legal principle of an agent, which has deep roots in common law, jurisprudence, philosophy, and math. When we think of an agent (some software) acting on behalf of a principal (a human user), this historical baggage imputes some important ethical obligations to the developer of the agent software. genAI vendors have been as eager as any software vendor to dodge responsibility for faithfully representing the user's interests even as there are some indications that at least some courts are not persuaded by this dodge, at least by the consumers of genAI attempting to pass on the responsibility all the way to end users. Perhaps it goes without saying, but I'll say it anyway: I don't like this newer interpretation of "agent". ↩
-
"Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents", Axel Backlund, Lukas Petersson, Feb 20, 2025 ↩
-
"random thing are happening, maxed out usage on api keys", @leojr94 on Twitter, Mar 17, 2025 ↩
-
"New study sheds light on ChatGPT's alarming interactions with teens" ↩
-
"Lawyers submitted bogus case law created by ChatGPT. A judge fined them $5,000", by Larry Neumeister for the Associated Press, June 22, 2023 ↩
-
During which a human will be busy-waiting on an answer. ↩
-
Given the fluctuating pricing of these products, and fixed subscription overhead, this will obviously need to be amortized; including all the additional terms to actually convert this from your inputs is left as an exercise for the reader. ↩
-
I feel like I should emphasize explicitly here that everything is an average over repeated interactions. For example, you might observe that a particular LLM has a low probability of outputting acceptable work on the first prompt, but higher probability on subsequent prompts in the same context, such that it usually takes 4 prompts. For the purposes of this extremely simple closed-form model, we'd still consider that a P of 25%, even though a more sophisticated model, or a monte carlo simulation that sets progressive bounds on the probability, might produce more accurate values. ↩
-
No it isn't, actually, but for the sake of argument let's grant that it is. ↩
-
It's worth noting that all this expensive measuring itself must be included in C until you have a solid grounding for all your metrics, but let's optimistically leave all of that out for the sake of simplicity. ↩
-
"AI Company Poll Finds 45% of Workers Trust the Tech More Than Their Peers", by Suzanne Blake for Newsweek, Aug 13, 2025 ↩
-
AI Chatbots Remain Overconfident - Even When They're Wrong by Jason Bittel for the Dietrich College of Humanities and Social Sciences at Carnegie Mellon University, July 22, 2025 ↩
-
AI Mistakes Are Very Different From Human Mistakes by Bruce Schneier and Nathan E. Sanders for IEEE Spectrum, Jan 13, 2025 ↩
-
Foreshadowing is a narrative device in which a storyteller gives an advance hint of an upcoming event later in the story. ↩
-
"People are worried about the misuse of AI, but they trust it more than humans" ↩
-
"Why I stopped using AI (as a Senior Software Engineer)", theSeniorDev YouTube channel, Jun 17, 2025 ↩
-
"I was an AI evangelist. Now I'm an AI vegan. Here's why.", Joe McKay for the greatchatlinkedin YouTube channel, Aug 8, 2025 ↩
-
"Study Finds That 52 Percent Of ChatGPT Answers to Programming Questions are Wrong", by Sharon Adarlo for Futurism, May 23, 2024 ↩
-
"Off the Mark: The Pitfalls of Metrics Gaming in AI Progress Races", by Tabrez Syed on BoxCars AI, Dec 14, 2023 ↩
-
"I tried coding with AI, I became lazy and stupid", by Thomasorus, Aug 8, 2025 ↩
-
"How AI Changes Student Thinking: The Hidden Cognitive Risks" by Timothy Cook for Psychology Today, May 10, 2025 ↩
-
"Increased AI use linked to eroding critical thinking skills" by Justin Jackson for Phys.org, Jan 13, 2025 ↩
-
"AI could end my job - Just not the way I expected" by Manuel Artero Anguita on dev.to, Jan 27, 2025 ↩
-
"The Emerging Problem of "AI Psychosis"" by Gary Drevitch for Psychology Today, July 21, 2025. ↩
15 Aug 2025 7:51am GMT