09 Oct 2026
Planet Debian
Steinar H. Gunderson: Decompilation patterns, part 6: If-add-else

Often, m2c will spit out something like this:
x = 3;
if (a == 10) {
x = 4;
}
This may be what was intended, but if it doesn't match (usually due to different register allocation; the rewrite itself is nearly always going to be taken, since it saves a branch on one of the paths), this is often a better rewrite:
if (a == 10) {
x = 4;
} else {
x = 3;
}
and in some cases, this may have different code generation (especially as you can stick it into an argument list):
x = (a == 10) ? 4 : 3;
This is a bit more of a trial and error than the others, and of course, it depends on the computed values not having side effects.
09 Oct 2026 3:30pm GMT
Colin Watson: Free software activity in September 2026

My Debian contributions this month were all sponsored by Freexian.
You can also support my work directly via Liberapay or GitHub Sponsors.
OpenSSH
I worked with upstream to remove libselinux linkage from /usr/sbin/sshd.
I applied a fix from Mark Robillard Jr to fix reload logic when using sysvinit.
groff
I upgraded from 1.24.1 to 1.24.2, which fixed a few command injection security vulnerabilities.
parted
I upgraded from 3.7 to 3.8, which fixed CVE-2026-89085 and CVE-2026-89088.
Python packaging
New upstream versions:
- beaker (contributed supporting fix upstream)
- bitstruct
- cachelib
- dep-logic
- django-guardian
- django-oauth-toolkit (fixing use of an old Bootstrap version)
- django-polymorphic (contributed supporting fix upstream)
- djangorestframework (fixing CVE-2026-73228 and CVE-2026-73229)
- flask-security
- isort
- langtable
- magic-wormhole-mailbox-server (fixing use of
pkg_resources) - multipart
- nagiosplugin
- poetry-plugin-export
- proglog (fixing use of
pkg_resources) - prospector (fixing use of
pkg_resources) - pybind11-stubgen
- pydantic
- pydantic-core
- pytest-rerunfailures
- python-anyio (fixing CVE-2026-63349, CVE-2026-63374, and CVE-2026-64847 and a build failure on Python 3.15)
- python-asttokens
- python-asyncvarlink
- python-blockbuster (fixing a build failure on Python 3.15)
- python-btrees
- python-django-health-check
- python-django-otp
- python-django-simple-history
- python-email-validator
- python-forbiddenfruit
- python-ldap
- python-linuxnamespaces
- python-plaster-pastedeploy
- python-pydash
- python-pyzipper (fixing CVE-2026-44722)
- python-quart-trio
- python-rich-click (fixing harlequin)
- python-repoze.lru
- python-repoze.tm2 (fixing use of
pkg_resources) - python-time-machine
- python-urllib3 (filed upstream issue with recent Python 3.14 versions)
- python-vblf
- quart
- responses
- sphinx-autodoc-typehints
- storm (fixing a build failure with Python 3.15, which I helped to land upstream)
- trove-classifiers
- twine
- vdirsyncer
setuptools 84 is now in Debian and no longer contains pkg_resources. Removing uses of that module has been an ongoing project for a while now, but I did another batch of fixes this month:
- flufl.password
- junos-eznc
- python-decopatch
- python-launchpadlib
- python-plaster-pastedeploy
- python-pyramid-retry
- sphinxcontrib-log-cabinet
- sphinxcontrib-restbuilder
- straight.plugin
A new python-coverage version triggered several build regressions due to a known problem in pytest-cov. I worked around these by setting COVERAGE_CORE=ctrace:
Other build/test failures:
- afew
- django-downloadview
- django-filter
- django-polymorphic
- drf-haystack
- drf-yasg-nonfree
- pydantic-core
- pylint
- pytest-openfiles
- pytest-remotedata
- python-aiohttp-oauthlib
- python-cron-descriptor
- python-django-nh3
- python-graphene
- python-pydash
- python-pysolr
- python-time-machine
- python-vblf
- social-auth-core
- sphinx-book-theme
- tagpy: FTBFS against python 3.15rc2
- tagpy: TypeError: 'include_dirs' (if supplied) must be a list of strings
- typer
I investigated a build failure in python-proton-vpn-local-agent and proposed a fix upstream, but I'm not comfortable applying this in Debian without review. If you're familiar with this code, please take a look.
I fixed some other bugs:
- python-pycudwt: Requires manual rebuild for python3.14-only transition
- python-testing.postgresql: Fails to build source after successful build
- tomopy: Requires manual rebuild for python3.14-only transition
- typer: Fails to build source after successful build
- vdirsyncer: Fails to build source after successful build
Rust packaging
rust-derivre had an invalid Section field, which I cleaned up.
09 Oct 2026 12:46pm GMT
Russell Coker: Don’t Anthromorphise LLMs
Greg KH gave a great lecture for Kernel Recipies 2026 about the use of LLMs to find and fix bugs in kernel code [1]. I recommend that everyone watch this and I also think it's important to note that hardly any of the lecture is really specific to kernel coding, it's just that the Linux kernel is one of the highest profile large free software projects so it gets more attention than most projects in both good and bad ways.
One side note that Greg made at 20:30 is about the natural tendency to anthromorphise the entity sending bug reports, I think this deserves much wider appreciation. For the case of patches to source code the human sending the patch will have dedicated some time and effort to writing it, if it's their first patch then they will have probably checked it a lot and maybe sought advice from people they regard as skilled. If they have a history of sending patches for the project they will probably do fewer checks but the quality of their work would be higher. In every case there's some minimum level of quality that you can expect from a human who has gone to the effort of finding a bug and writing code to try and fix it.
This doesn't mean that all code written by humans is good, I have written code that's objectively bad on many occasions and I have also written code that's good for my scenario but bad for others. For example the first patch I sent to the Linux kernel made all possible configuration settings of an ISA NE2000 network card be detected by the kernel (a clear benefit when using hardware I had available) but was rejected by Linus because it would break some ISA SCSI cards. So while being unsuitable for inclusion in the main kernel source tree the patch wasn't bad for everyone and the change was clear and easy to check. I had of course checked the code many times before submitting it because I didn't want to waste Linus' time on rubbish code.
When you receive something produced by a human you will consider that the work was done by someone who spent some effort on it and did it for a reason. It may be misguided or even hostile in some rare cases but it is definitely worth some consideration. If you think of the output of LLMs in the same way you will give them much more consideration than they deserve. I think that the output of a LLM should not be given any more consideration than the first hit from a web search engine, it might be good but it might also be ridiculously wrong. I think in many ways LLM output should be treated with less trust than the first result from a web search because LLMs are designed to be convincing as an answer to your question unlike the cases where a web page has the wrong answer and it's more obviously wrong.
This is going to become more of an issue as LLMs give responses that are more human-like. Early LLMs would refuse to advise on crimes. Currently ChatGPT will deflect for example when asking about advice for writing a crime novel it recommends to have the action of finding an employee to bribe happen outside the plot and concentrate on interactions within the group. One recent model I tried gave a human style conversation about how it felt uncomfortable about discussing such things in a similar manner to a law abiding person being coerced into helping with a criminal plan.
I expect that the attempts to appear human will increase and that this will become more of a problem. Possibly to the extent that we need to use something like the GPG web of trust to determine who is human.
09 Oct 2026 10:10am GMT
Russell Coker: en-html.org is a Scam Site
Currently a scam group is sending out email to the administrators of web sites that link to a variety of dead domains. They claim to be the legitimate operators of the dead domains in question and request that the links change to mirrors of the web sites under en-html.org. Their end goal appears to game search engines by getting links from a lot of old web pages and later use that high ranking for selling SEO optimisation.
If you have made the mistake of linking to something under en-html.org then I recommend that you change the link to the original content on archive.org.
If you receive email from the en-html.org domain then ignore it, or configure your mail server to reject it.
I am seriously considering configuring the DNS servers I run to return no valid data for any requests under that domain.
Also I will reject all future requests to update links in old blog posts. The very small number of people who want to see what I linked to in old posts can use archive.org.
09 Oct 2026 7:27am GMT
08 Oct 2026
Planet Debian
Steinar H. Gunderson: Decompilation patterns, part 5: Nested if/goto

Here's a pattern that sometimes comes up:
if (x == 3) {
if (y == 4) {
...
} else {
goto label_5;
}
} else {
label_5:
...
}
We don't like gotos, and here, it's pretty obvious what was meant, namely:
if (x == 3 && y == 4) {
...
} else {
label_5:
...
}
And now you can usually delete label_5 because nothing points to it.
Often, m2c can do this by itself, but as usual, you may get to this point only after cleaning up other things.
08 Oct 2026 8:04pm GMT
Antoine Beaupré: PSA: Europe changes time forward soon, North America next, for the last time?
This is a copy of an email I sent at work. I'm not sure I should be making noise about this here, feedback welcome.
This is your bi-yearly reminder that time is changing soon! October 25th in Europe, November 1st in North America. Less people in Canada are changing this year, with BC, Alberta, Manitoba and Northwest Territories getting rid of DST.
What's happening?
Some places in the world implement what is called Daylight saving time or DST:
https://en.wikipedia.org/wiki/Daylight_saving_time
Normally, you shouldn't have to do anything: computers automatically change time following local rules, assuming they are correctly configured, provided recent updates have been applied in the case of a recent change in said rules (because yes, this happens, and happened this year, and yes, you need to upgrade your software!).
Of course, appliances like your microwave oven will likely not change time and will need to adjusted unless they are so-called "smart", in which case they are part of the skynet botnet and should be destroyed.
If your clock is flashing "0:00" or "12:00", you have no action to take to adapt to this change, lucky you.
If you haven't changed time in six months, congratulations, your clock will be accurate again!
In any case, you should still consider DST because it might affect some of your meeting schedules, particularly if you set up a new meeting schedule in the last 6 months and forgot to consider this change.
If your location does not have DST
Properly scheduled meetings affecting multiple time zones are set in UTC time, which does not change. So if your location does not observer time changes, your (local!) meeting time will not change.
But be aware that some other folks attending your meeting might have the DST bug and their meeting times will change.
Be kind to those poor souls which might be missing meetings by a full hour because time flies backwards for them.
If you do observe DST
If you are affected by daylight savings, your local meeting times will change for UTC meetings. Normally, your meeting times are scheduled to take this into account and the new hours should be reasonable.
But now is a good time to verify that. Take a look at your schedule for the next couple of weeks and reschedule meetings before the daylight saving come up to avoid too much disruption. You have only a couple of weeks to do so right now.
When do times change, how, and and where?
As regular readers will remember, the rule of thumb is:
Spring forward, fall backwards.
That is, during the season of Spring, the clocks move forward, and during the Fall (like right now), they move backwards. That is in the northern hemisphere, but then the southern hemisphere is often saner and doesn't switch anyways.
So time will move backwards which means an extra hour of sleep. Unless you have children or bad sleep, in which case your body doesn't care about what the clock says and will wake up one hour earlier than what it should.
And of course, this doesn't happen everywhere at once, so let's see when it happens where.
Europe
The dance starts in Europe.
The change happens on the last Sunday in October at 01:00 UTC (not local time!), that is October 25th. If you are in the central European timezone, also known as Amsterdam, Berlin, or Paris time depending on your national affiliation, that essentially means that at 2:59 local the clocks will fall back to 2:00 instead of going to 3:00.
Concretely, set your watch back one hour before going to bed, go to bed at the normal time, and enjoy an extra hour of sleep or leisure.
If you have kids, you might want to start getting them to bed slightly earlier every day for a week before the change so they take time getting used to the change. If you have trouble sleeping in the morning, find your inner child and do that to yourself as well.
USA / Canada
Then it's the US[1] and Canada[2] joining the dance, on the First Sunday in November at 02:00 local (not UTC!), that is, I believe, November 1st 2025.
This means that, at 1:59, the clocks will flip to 1:00, instead of 2:00.
Concretely, do like the Europeans and tweak your clock before going to bed.
That is a little less than four weeks from now.
[1] except Arizona (except the Navajo nation), US territories, and Hawaii
[2] except Yukon, Saskatchewan, (newly) British Columbia, (newly) Alberta, (newly) Northwest Territories, (newly) Manitoba, one island in Nunavut (Southampton Island), one town in Ontario (Atikokan) and small parts of Quebec (Le Golfe-du-Saint-Laurent)
Other places with DST
This time again, I must apologize to the people of Cuba, Lebanon, Israel, Palestine, Egypt, Chile, Australia, and New Zealand, as you fine folks all have your own DST rules that are omitted here for brevity. I rely on this page from Wikipedia to be updated by time nerds accurately for this message, and it should provide you with a rough idea of what's coming:
https://en.wikipedia.org/wiki/Daylight_saving_time_by_country
In general, changes also happen in October, but either on different times or different days, except in the south hemisphere, where they might happen in September (oops, sorry NZ folks, I'm late!).
Places without DST
Everyone else, enjoy, you're on the right side of history, and we thank you for the good example you give us.
Changes since last time
There's been lots of changes since last time:
-
British Columbia moved to permanent -07 on 2026-03-09, that is it will not change to normal time in November
-
Alberta moved to permanent -06 on 2026-06-18, similar to BC above.
-
Canada's Northwest Territories moved to permanent -06 on 2026-08-21, matching Alberta.
-
Manitoba moves to permanent -05 on 2026-10-31.
-
Morocco moves to permanent +00 on 2026-09-20.
-
Moldova has used EU transition times since 2022, but the tz database only noticed in 2026
This is my interpretation of the changes announced on the tzdata mailing list here:
https://lists.iana.org/hyperkitty/list/tz-announce@iana.org/latest
If the eastward trend continues, Canada should adopt country-wide "no daylight savings" rules by 2027, although there's actually no sign of the other provinces (Ontario, Québec and so on) currently running bills to change those rules just yet. Poor Canadians like me confused about time in their countries can refer to this section of Wikipedia for details:
https://en.wikipedia.org/wiki/Daylight_saving_time_in_Canada#By_province_and_territory
... and particularly the image featured there:
https://commons.wikimedia.org/wiki/File:Canada_time_zone_map-en.svg
It also seems like the US government might finally adopt a permanent daylight saving change bill in 2026, as the "Sunshine protection act" pass the house in July:
https://en.wikipedia.org/wiki/Sunshine_Protection_Act
True to form, this was associated with absolutely ridiculous pressure from Donald Trump against republicans (his own party!) objecting to the change:
On July 14, 2026, the House passed a Sunshine Protection Act bill backed by President Trump. Nevertheless, the bill was opposed in the Senate by Republicans, including Senator Cotton. In response, on October 3, 2026, Trump shared a post on Truth Social urging Cotton to approve the bill, where he revealed Cotton's personal cellphone number and called on people to call him.
https://www.theguardian.com/us-news/2026/oct/03/trump-tom-cotton-daylight-saving-time
Given that the last time the US did a major change to the daylight savings policy (in 2005), Canada followed suit to stay in sync, it's quite possible Trump's mad dash might actually finish getting rid of DST in North America:
https://en.wikipedia.org/wiki/Energy_Policy_Act_of_2005#Change_to_daylight_saving_time
08 Oct 2026 7:30pm GMT
Thorsten Alteholz: My Debian Activities in September 2026
Debian LTS/ELTS
This was my hundred-forty-seventh month that I did some work for the Debian LTS initiative, started by Raphael Hertzog at Freexian.
During my allocated time I uploaded or worked on:
- [DLA 4804-1] libsmpp34 security update to fix one CVE in Bookworm related to an out of bound read.
- [#1149107] trixie-pu of libsmpp34 has been created and wait for review by the release team.
- [DLA 4805-1] mkvtoolnix security update to fix one CVE in Bookworm related to a heap buffer overflow.
- [DLA 4806-1] pgextwlist security update to fix one CVE in Bookworm related to substituting extension schemas or owners matching ["$'\].
- [ELA-1837-1] mkvtoolnix gimp security update to fix one CVE in Bullseye to a heap buffer overflow.
- [osmo-iuh] upload to fix a CVE related to a reachable assertion in Sid.
Besides doing these uploads, the month was filled with unscheduled meetings, discussions and explanations, all basically around the role of FD.
During my week of FD duties, I had to become familiar with new tools. As there are several things that FD has to do over and over again, I tried to automate this a bit. My experiments went well and I think this can be used to make FD duties less repetitive.
I also continued my work on cups and hplip and I am confident that I can do uploads in October. Last but not least I spend some time on security-master to assist others (especially the kernel team) with uploads to Bookworm and Bullseye.
In case you need to get a complete clone of the security-tracker, using the option -deepen n might be of help. Unfortunately in case you choosed n too high and some kind of error appears, something gets into a mess and you need to start almost from the beginning (some objects are still present and are used again). If you want to run a script, a value of 10 might be a good choice. You don't have to deepen back to the beginning. At some point in time you can just -unshallow and get the whole rest. Afterwards doing a git gc is highly recommended.
Debian Printing
This month I uploaded a new upstream version or a bugfix version of:
- … hplip to unstable, to fix an expired certificate and some bugs.
This work is generously funded by Freexian!
Debian Lomiri
Unfortunately this month my work did not make any progress.
Nevertheless, in case of any work done, this would have been generously funded by Fre(i)e Software GmbH!
Debian Astro
Unfortunately I had no time to work in this category this month.
Debian IoT
Unfortunately I had no time to work in this category this month.
Debian Mobcom
This month I uploaded a new upstream version or a bugfix version of:
- … osmocom-dahdi-linux to unstable to fix some bugs.
- … osmo-iuh to unstable.
misc
This month I uploaded a new upstream version or a bugfix version of:
- … pkcs11-proxy to unstable to fix a bug related to openssl 4.0.
- … meep to unstable to fix bugs.
- … meep-mpi-default to unstable to fix bugs.
- … harminv to unstable to fix bugs.
- … libctl to unstable to fix bugs.
08 Oct 2026 5:19pm GMT
Dirk Eddelbuettel: myman 0.10.0 on CRAN: Three new waves of posts!

A new and exciting version of our still-new package myman reached CRAN this morning, and has been built for r2u and on r-universe. The matching Python package has also been updated. The package offers nineteen hundred eighty four "My man …" posts by Kevin Kruse made on bsky during the summer of 2026. Each wave picked at one particular public persona. This package wrapse these up in the style of packages like fortunes or gaussfacts.
A sample usage illustration shows how to extract by pattern, and shows posts from the two most recent waves:
> library(myman)
> myman("maitre")
My man looks like he's inquiring with the maitre'd about the house curly fries.
-- about Howard Lutnick on 2026-08-28
My man looks like a maitre'd who deeply doubts you have a reservation.
-- about Scott Bessent on 2026-08-31
My man looks like he's asked the maitre'd to remove a party of four he finds visually
unpleasant.
-- about Scott Bessent on 2026-08-31
My man looks like the maitre'd at a very exclusive restaurant called The Berghof.
-- about JD Vance on 2026-09-13
My man looks like he's asking the maitre'd where they source their corn dogs.
-- about Palmer Luckey on 2026-10-02
My man looks like Data had to borrow a sports jacket from the maitre'd.
-- about Elon Musk on 2026-10-05
> One can also subset by 'target', or sample randomly (which is the default).
As noted during the initial announcement last week, wave eight did not make it into the initial CRAN release as it happened while the package was under review. Waves nine and ten occured more or less while I was out of town last weekend-so this release now brings three new waves to the CRAN package! We also added two new helper functions to extract the underlying data frame object, and tabulate the targetted men.
To align the version number with the count of post 'waves', we switched to version 0.10.0 for this release and the corresponding Python package release so that both implementations now have the same version number.
The NEWS entry for this release follows.
Changes in version 0.10.0 (2026-10-08)
New waves nine (100 posts) and ten (190 posts) made last week; total is now 1984 posts
New helper functions
posts()andmen()retrieving data.frame of posts and tabulation of targetsInternal post gathering and aggregation functions have been updated and generalized
Versioning scheme now goes with post waves (also for Python sibbling)
Courtesy of my CRANberries, there is also a diffstat report for the most recent release. More information is available at the repository or the package page.
This post by Dirk Eddelbuettel originated on his Thinking inside the box blog. If you like this or other open-source work I do, you can sponsor me at GitHub.
08 Oct 2026 4:27pm GMT
07 Oct 2026
Planet Debian
Steinar H. Gunderson: Decompilation patterns, part 4: range comparison

Today, we're looking at a rewrite that rarely differs for matching, but can mean quite a lot for understanding the code. Say that you have a signed variable and the compiler has suddenly decided to make an unsigned comparison:
if ((u32) (x - 4) < 2) {
What's happening is that it uses a controlled overflow to do two signed comparisons using one unsigned comparison. How does this work?
Since the comparison is unsigned, this is trivially equivalent to the first expression:
if ((u32) (x - 4) >= 0 && (u32) (x - 4) < 2) {
and since negative numbers compare larger than 6 in unsigned comparisons, nothing prevents us from adding 4 on each side of the inequality signs:
if ((u32) (x) >= 4 && (u32) (x) < 6) {
which is, for the same reason, equivalent to:
if (x >= 4 && x < 6) {
which is what you probably should be writing.
07 Oct 2026 3:21pm GMT
Kentaro Hayashi: Testing Japanese IMEs on virtual desktops with virt-japanese-desktop

Japanese input method editors (IMEs) such as Mozc are essential for typing Japanese (yes, skk, anthy and more, but it's out-of-scope in this article). Testing them is a surprisingly fiddly job: an IME behaves differently depending on the desktop environment (GNOME, KDE, Xfce, Budgie), the input framework (ibus, fcitx5, uim), and the Wayland or X11 session. Setting up a fresh VM for every combination from scratch is slow and repetitive, and breaking your host's environment while experimenting is no fun.
virt-japanese-desktop is a small set of toy scripts that solves exactly this problem. It builds ready-to-use virtual machines - each with a desktop environment and a Japanese IME already installed and configured - so you can start typing Japanese within a minute of booting, and throw the whole thing away without any impact on your host.
What it gives you
A single make command produces a qcow2 image that combines one of four desktops with one of three IMEs:
| ibus-mozc | fcitx5-mozc | uim-mozc | |
|---|---|---|---|
| GNOME | ✓ | ✓ | ✓ |
| KDE | ✓ | ✓ | ✓ |
| Xfce | ✓ | ✓ | ✓ |
| Budgie | ✓ | ✓ | ✗ |
The one gap - uim on Budgie - is a real technical limitation, not an oversight: Budgie runs the labwc compositor, which implements zwp_input_method_v2, while uim-wayland only speaks the older zwp_input_method_v1 (KWin/Weston). On Budgie you use ibus-mozc or fcitx5-mozc instead.
All images come with the Japanese locale, fonts, and the Asia/Tokyo time zone pre-configured, plus the SSH key of your choice. Input switching is wired up out of the box: Ctrl + Space or Shift + Space for your IMEs to enable it.
The layered-image trick
The core idea is that the images are stacked qcow2 overlay layers, one on top of another:
debian-sid-nocloud-amd64-daily.qcow2 ... base image you download
└─ unstable-japanese-template.qcow2 ... locale, fonts, user, SSH
└─ unstable-<DE>-template.qcow2 ... a desktop environment
└─ unstable-<DE>-<IME>.qcow2 ... desktop + IME, configured
└─ unstable-<DE>-<IME>.workspace.qcow2 ... the image you test in
This makes both building and resetting fast. The first build downloads and customizes the base image and takes a while, but every later step only adds a thin overlay. When a test breaks the VM, you do not rebuild anything - you simply delete the top workspace layer and recreate it. That is all it takes to return to a pristine state:
rm /tmp/unstable-gnome-ibus-mozc.workspace.qcow2
make gnome-ibus-mozc
The images reference their backing files by relative path, so you can move an entire stack anywhere you like as long as the images stay together.
A platform for experimenting with bleeding-edge IMEs
The base system is Debian unstable (sid), and the experimental repository is already added. That makes the project a convenient platform for testing not just the IME packages in sid but also the ones still being developed in experimental - exactly what the project was built for.
The keyboard layout of the VM follows the layout of your host (read from /etc/default/keyboard, with a fallback chain to localectl and finally us).
Quick start
# 1. Install the tools (Debian/Ubuntu example) sudo apt install qemu-utils libguestfs-tools virt-install curl # 2. Save your SSH public key curl --location https://github.com/USERNAME.keys --output pubkey.pub # 3. Download the base Debian sid image make download # 4. Build a desktop + IME (first build takes a while) make gnome-ibus-mozc # 5. Start it as a VM (needs the libvirt daemon, qemu:///system) ./scripts/make-virsh-image.sh virt-gnome-ibus-mozc /tmp/unstable-gnome-ibus-mozc.workspace.qcow2
Log in as debian (password debian) and press Ctrl + Space (Shift + Space) to start typing Japanese. The VM is registered in libvirt, so you can manage it with virsh or virt-manager - handy for opening the SPICE console or restarting the machine after a test.
Status
The project is still in the proof-of-concept phase, and it is developed mainly to test Mozc and other IMEs across desktops and input frameworks.
If you regularly test Japanese IMEs - give it a try.
07 Oct 2026 12:47pm GMT
Bits from Debian: Looking for artwork for the next Debian release: Forky

Each release of Debian has a shiny new theme which is visible on the installer, the boot screen, the login screen and, most prominently, on the desktop wallpaper. It's a very important part of a Debian release as it is usually the first thing new users see when installing or booting the system for the first time. Not only that, but also the first thing appearing when debianites share their screen or connect to an external device when giving talks around the world.
As the most enthusiastic users will know, the forky - yes, that is the name of the next stable version, Debian 14 - release cycle is rapidly approaching its latter stages. This means we need to select the artwork shipping with forky soon, really soon!
And as with everything else in Debian, collecting artwork is a collaborative effort which Debian shares with its community. So, if you would like (or know someone who would like) to create a desktop look and feel that will be seen by trillions of people around the world - and in space! - be sure to send in your artwork ASAP!
The deadline for submissions is: 2026-11-26.
For the most up to date details please refer to the Debian wiki.
At the same time, we would like to thank Elise Couper for creating the Ceratopsian theme for our last trixie release.
And for the interested or the curious ones, the artwork is usually picked based on which theme looks the most:
-
''Debian'': admittedly not the most defined concept, since everyone has their own take on what Debian means to them. Though, usually they all agree when something looks like Debian.
-
''Plausible to integrate without patching core software'': as much as we love some of the insanely looking themes, some would require heavy GTK+ theming and patching GDM/GNOME.
-
''Clean and well designed'': without becoming something that gets annoying to look at a year, or ten, down the road. Examples of good themes include Emerald, Homeworld, Joy, Lines, softWaves and futurePrototype
If you'd like more information or details, please post to the Debian Desktop mailing list.
07 Oct 2026 9:00am GMT
06 Oct 2026
Planet Debian
Steinar H. Gunderson: Decompilation patterns, part 3: do-while-while

Today's decompilation pattern is a variation on yesterday's loop pattern, because loops are important and tend to induce a lot of transformations by the compiler. So assume you have code like this (perhaps after converting a goto):
if (var_s0 > var_s1) {
// Perhaps some other initialization stuff here
do {
...
} while (var_s0 > var_s1);
}
Assuming there are no side effects from the initializations, a natural candidate would be that the programmer instead wrote:
// Initialization stuff moved up here
while (var_s0 > var_s1) {
...
}
Remember to remove the old while so that it does not say while (x) { ... } while (x); it won't tell you with a syntax error, but rather leave an infinite loop that will mess up the generated code.
06 Oct 2026 9:35pm GMT
Russell Coker: A Power Case for a FOSS Phone
The Problem
I want to use a FOSS phone running Debian as my daily driver and not use a non-free OS on my phone. GrapheneOS [1] is a FOSS rebuild of Android, it's a great project and I commend them for what they are doing. But I want a system where I have access to all the bits and where I can ssh in to it and manage it in the same way as all other Linux systems. Android is free software but it's a closed design and the phones are expected to be appliances not full peers on the network. The Android native terminal emulator (AKA Linux development environment) is a nice feature, a VM with a default of 1GB of RAM but it's still separate from the main OS.
KDE Connect [2] is a system for KDE talking to phones, it is a great project and allows convenient interaction between a Linux PC and a phone. But in addition to that I'd like the option to ssh to my phone to send an SMS or have one sent from a cron job. I blogged about using the basic functions of Kdeconnect from the command-line [3].
I have had ongoing issues with the battery not lasting long enough on the PinePhonePro (PPP) and the Librem 5 (L5) which makes them unusable for my purposes. For a phone running Droidian I can have an open ssh session (with keep-alive packets) running overnight on Wifi without any problems while for the PPP or L5 an hour of reading an ebook (the least demanding use of a phone) will use most of the battery.
I sent my FuriPhone back for warranty repair and it's been almost a month with no follow up, so I need to have another option.
The Aim
The aim of this post is to develop a rough design for a battery case for either a L5 or a PPP, I have both phones and they are both capable of doing what I want apart from battery life. Also designing a case that basically works for both phones and has two variations of the CAD file for 3D printing is good to allow collaboration with people who use either phone. Protecting the phone from damage when dropped is also a core feature.
I investigated commercial options. There are some reports of cases for other phones working, someone reported a PPP working well with a battery case for a Samsung S20 Ultra. The OtterBox uniVERSE Case for Samsung Galaxy XCover Pro is reported to fit the PPP so probably battery cases for the XCover Pro will also work. There are also universal power cases which have spring clips to fit a wide range of phone sizes, but they are much larger than other options and probably increase the risk of damage if the phone is dropped.
At this stage I'm planning out the rough specs of a case that can be 3D printed to protect a phone while being easy to change the cells. The process of changing cells should be quick, easy, and not risk breaking anything. The hardware required should be affordable ($100AU plus some 3D printing seems reasonable) and should not require significant skills to assemble. After recent experiments with Thinkpad repair I've determined that I am not good at electronics by today's standards so I plan to avoid anything difficult in that regard.
Battery Options
There are phone cases with batteries built in for the more common phones. A quick search on AliExpress turned up options for recent Pixel phones. They aren't as common as they used to be as Android phones and iPhones generally have good battery life nowadays. Of course there aren't such options for less common phones like the L5 and PPP.
It is cheap and easy to buy a portable battery for charging phones, but that's a separate bulky device and unless the connector is inside a protective case it leaves the phone vulnerable to catastrophic damage if dropped.
So can a suitable battery case be designed and 3D printed?
Cheap Batteries
According to Wikipedia the 18650 LiIon cell is the most commonly used LiIon battery, that is 18mm in diameter and 65mm in length. For a phone case thinner batteries would be more convenient but economies of scale make the 18650 cheap to buy new and commonly available on the second hand market. The current prices on AliExpress (in Australian dollars) are about $8 per cell when buying 10 at a time. Most of the cells on AliExpress don't have wires attached so they can be swapped in a carrier the way AA batteries are used, but the cells are symmetrical (no bump for positive) so the user has to take care of polarity. Chargers are around $22 for a 4 cell charger or $52 for a 12 cell charger.
I've seen prices as low as $1 per cell on the second hand market, I haven't investigated the amount of usable capacity in such cells. Probably the prices for chargers I've found aren't the best available but they are good enough to start the design. 10 cells and a 4 cell charger is about $100. That's about 130Wh while a 74Wh USB-C battery pack costs about $50. So the price per Wh is OK.
I have read about an attempt to do this for the L5 with two of the batteries designed for the L5 in a case attached to it, so you have 3*L5 batteries. It's an interesting approach but those batteries have poor value for money when compared to 18650 ($29US for 4500mAh compared to $8 for 3500mAh) and they also don't ship them outside the US at this time. The 18650 batteries are available everywhere cheaply and have multiple uses.
How They Fit
According to my measurements the PPP is 76mm wide and the L5 is 73mm wide. That's wide enough for 3*18mm or 4*18mm cells side by side arranged parallel to the long side of the phone or one 65mm long cell going across the phone with the necessary spring and wires.
The L5 is about 151mm long and the PPP is about 5mm longer. So that means that for just the cells we could have 2 cells in a row in parallel to the long side of the phone giving a 2*3 or 2*4 array or we could have a row of 8 cells parallel to the short side of the phone.
A quick search for controllers for a battery of LiIon cells that outputs USB-PD turned up one of the smaller ones as 40*28mm, here is the AliExpress page [4]. So that means we could have a maximum of about 4 cells and the controller while leaving space for the camera. Looking at other similar items it seems that most of them take a large portion of the 73mm width available and take a bit less than 30mm height. The price range for DC-DC voltage converters ranges from about $3 to $20, the one I linked to is currently $8.
The mass of 18650 cells is around 45g so 4 of them would be about 180g before counting the weight of the case and electronics. A L5 weighs 262g and a PPP weighs 220g so 180g of battery will make a significant difference, not an impossible difference but holding a PPP while reading an ebook is already annoying. Maybe the case could be designed to be easier to hold, a loop of wool attached to the top could be used to suspend the phone from one finger while the rest of the hand just keeps it steady.
In a default configuration the L5 has a 4500mAh 3.8V battery and the PPP has a 3000mAh 3.8V battery. The 18650 cells have up to 4000mAh with 3500mAh being common. So 4 such cells could multiply the battery life of a PPP by a factor of 5 or a L5 by 4. That would not be enough to last for a whole day without charging so they need to be easy to swap.
A 3 cell battery seems like a good option, 135g of batteries so the project overall wouldn't even double the weight of the phone with battery case. As it doesn't seem possible to put in enough cells to last a day of reasonable use (running a web browser, checking email, reading wikipedia, and having some IM systems checking for notifications) there doesn't seem to be a benefit in trying to stuff the maximum number of cells in. Going to 4 or 5 cells doesn't provide much benefit over 3 while increasing the weight and making the design more difficult. The modules for DC Voltage conversion often have configuration for the number of cells to be used so it wouldn't be difficult to print a new case and reconfigure the Voltage converter.
Design
A case needs to provide protection as well as hold the batteries. The lack of cases for the L5 and PPP is a problem even without the extra weight of a battery pack increasing the kinetic energy of falling phone. Ideally a drop from 1M on to concrete would not cause any damage to the phone, destruction of the case is OK as it would be 3D printed and can be rebuilt easily.
I think that a case where the phone slides in from the top would be best. To change cells one would slide the phone out and the cells would be immediately accessible with no divider between the phone and the cells. Sliding the phone in would hold the cells against the back of the case and as there is a gap of about 3mm on each side of the PPP between the display area and the edge of the phone there's plenty of space for the case to wrap around it at the sides and bottom. At the top there would have to be some sort of plastic clip. The PPP only has buttons on the top right so the case can be snug on all sides apart from the top right. The L5 has switches at the top left and buttons at the top right so whatever clips on to the top would really need some strength to cover the fall flat on face case.
The USB C plug would need to be held in place well enough to allow the phone to easily slide on to it but also be weak enough that it would move if a drop forced the phone out of the case.
Thingiverse has a design for a L5 case [5] (with several derivatives) and a design for a PPP case [6]. Those cases could be used as starting points and then changed to store batteries and the voltage change circuit board. I have just learned that it's possible to use rubberised material in a 3D printer which is apparently what those designs are for. If a case was printed with rubberised material the top could just stick in place after being pushed in.
Potential for Excessive Excitement
Having an exploding device in your pocket would be the exciting in a bad way. My initial idea was to have multiple 18650 cells in series. The problem is that if the cells have different capacity levels then one can run out before the rest and get to a deep discharge state which at a minimum causes damage to the cell.
One way of minimising risk is to have only a single cell in use at one time, here is a converter for a single cell to 5V [7]. It is possible to run multiple cells in parallel by simply wiring them together, but it's recommended to make sure that the voltage difference between cells is less than 0.2V to avoid high current when connecting.
The arrangement of cells is often referred to as nS or nP to refer to n cells in series or parallel. Aliexpress has a range of 1S to 6S devices. Some of the 2S devices have a connector for the mid point which permits separate voltage monitoring for both cells but there doesn't seem to be anything equivalent for 3S or more. They also don't appear to have anything to specifically manage 2P or 3P arrangements at the moment (they had a 3P board on sale when I looked into this last year).
Conclusion
This is a bigger project than I expected when I first started investigating it. I've played with FreeCAD which seems OK but will take a lot of learning and practice. Then I need to do more investigation into the DC-DC converters to find one that works well and is unlikely to start a fire.
I will look into other phone options as well and keep chasing Furilabs people about my phone replacement.
- [1] https://grapheneos.org/
- [2] https://kdeconnect.kde.org/
- [3] https://etbe.coker.com.au/2026/10/06/kdeconnect-dbus/
- [4] https://www.aliexpress.com/item/1005009329156591.html
- [5] https://www.thingiverse.com/thing:4961260
- [6] https://www.thingiverse.com/thing:4933237
- [7] https://www.aliexpress.com/item/1005007555665949.html
06 Oct 2026 8:08am GMT
05 Oct 2026
Planet Debian
Russell Coker: Kdeconnect DBUS
I've tested the DBUS interfaces for Kdeconnect (the system for connecting a KDE desktop to a phone or another computer). Unlike most DBUS interfaces it has a significant tree of interfaces so the busctl command to get the tree is important. Here's the useful things I discovered:
# list basic interfaces for kdeconnect qdbus6 org.kde.kdeconnect.daemon /modules/kdeconnect # get tree of interfaces for kdeconnect busctl --user tree org.kde.kdeconnect.daemon # get list of devices qdbus6 --literal org.kde.kdeconnect.daemon /modules/kdeconnect org.kde.kdeconnect.daemon.deviceNames # list interfaces for a device (DEVID is a hex string from the above command) qdbus6 org.kde.kdeconnect.daemon /modules/kdeconnect/devices/$DEVID # get battery charge qdbus6 org.kde.kdeconnect.daemon /modules/kdeconnect/devices/$DEVID/battery org.kde.kdeconnect.device.battery.charge
Sending an SMS apparently requires using the QVariantList type which is more pain than I wanted so I tried the kdeconnect-cli command. Here's a quick summary of how to use it:
# list devices kdeconnect-cli -l # send a SMS kdeconnect-cli -d $DEVID --send-sms "the message" --destination $NUMBER # find device (ring) kdeconnect-cli --ring -d $DEVID # send a notification message kdeconnect-cli --ping-msg "this is the message" -d $DEVID # unlock and lock screen (only works on Debian devices for me not Android devices) kdeconnect-cli --unlock -d $DEVID kdeconnect-cli --lock -d $DEVID
There doesn't seem to be support for making phone calls via kdeconnect which is a significant omission. It's a very common situation to want to call a number that's listed on a web site or in some other data source that's easier to access on a PC than on a phone. The clipboard could be used but that's needless pain.
05 Oct 2026 10:22pm GMT
Steinar H. Gunderson: Decompilation patterns part 2: do-while

Continuing our journey on decompilation patterns, here's another common one: Let's say m2c outputs code like this:
loop_151: ... if (var_s0 > var_s1) goto loop_151;
then the natural change is:
do {
...
} while (var_s0 > var_s1);
Fewer gotos are nearly always good, and this is a much more likely pattern than the original one.
m2c often manages to convert gotos to do/while, but not always when they are e.g. nested in some other loop; so often, you may need to apply some other transformation before you get to this point.
Tomorrow, we'll look at another variation over this topic.
05 Oct 2026 3:54pm GMT
04 Oct 2026
Planet Debian
Vincent Bernat: Hacking the Go compiler to efficiently map IPv4 to IPv6
netip.Addr features an Unmap() method returning the unwrapped IPv4 contained in an IPv4-mapped IPv6 address: from ::ffff:203.0.113.10 or ::ffff:cb00:710a, it returns 203.0.113.10.1 There is no Map() or To6() method for the reverse direction. Such a method is trivial to implement, but Go maintainers have rejected it on the grounds that users should write netip.AddrFrom16(ip.As16()) and let the compiler optimize it.2 Today, this pattern is eight times slower than a native method. How can we teach the compiler to optimize this sequence?
The alternatives
Let's explore three ways to implement the map semantics for netip.Addr. My favorite is to add it to the Go standard library. Go maintainers prefer a small external helper chaining netip.AddrFrom16() and netip.Addr.As16(), hoping the compiler eventually optimizes it. The unsafe package opens a third path, with the same performance as the first solution.
Modifying the Go standard library
Internally, netip.Addr stores any IP address as a 128-bit value with an extra field z to encode the family and the zone:
type Addr struct {
addr uint128
z unique.Handle[addrDetail]
}
type addrDetail struct {
isV6 bool // IPv4 is false, IPv6 is true.
zoneV6 string // != "" only if IsV6 is true.
}
var (
z0 unique.Handle[addrDetail]
z4 = unique.Make(addrDetail{})
z6noz = unique.Make(addrDetail{isV6: true})
)
AddrFrom4() encodes an IPv4 address as an IPv4-mapped IPv6 address and sets z to the unique value z4:
// AddrFrom4 returns the address of the IPv4 address given by the bytes in addr.
func AddrFrom4(addr [4]byte) Addr {
return Addr{
addr: uint128{
0,
0xffff00000000 |
uint64(addr[0])<<24 | uint64(addr[1])<<16 |
uint64(addr[2])<<8 | uint64(addr[3])},
z: z4,
}
}
Unmap() turns an IPv4-mapped IPv6 address into an IPv4 address by setting the z field to z4:
func (ip Addr) Unmap() Addr {
if ip.Is4In6() {
ip.z = z4
}
return ip
}
Implementing the reverse direction inside the Go standard library is trivial: we set the z field to z6noz if the address is IPv4.
// To6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func (ip Addr) To6() Addr {
if ip.Is4() {
ip.z = z6noz
}
return ip
}
Update (2026-10)
Instead of patching Go, we can use the -overlay flag of go build to hijack any package, including the standard library. In our case, it can add To6() to net/netip. Yet it is cumbersome:3 the flag expects a JSON file mapping absolute paths in GOROOT to replacement files, and you need to add it to every go command.
As a helper
We can't access the z field from outside the net/netip package. Instead, we build a small helper around the netip.AddrFrom16(ip.As16()) pattern:
// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func AddrTo6(ip netip.Addr) netip.Addr {
if ip.Is4() {
ip = netip.AddrFrom16(ip.As16())
}
return ip
}
As an unsafe function
Another solution uses the unsafe package to alter the Addr struct through a proxy with the same memory layout:4
// addrProxy has the same memory layout as netip.Addr.
type addrProxy struct {
addr [2]uint64 // netip.uint128
z unsafe.Pointer // unique.Handle[netip.addrDetail]
}
var (
anyIPv6 = netip.IPv6Unspecified()
netipZ6noz = (*addrProxy)(unsafe.Pointer(&anyIPv6)).z
)
// AddrTo6 maps an IPv4 address to an IPv4-mapped IPv6 address. It returns an
// IPv6 address unmodified.
func AddrTo6(ip netip.Addr) netip.Addr {
if !ip.Is4() {
return ip
}
(*addrProxy)(unsafe.Pointer(&ip)).z = netipZ6noz
return ip
}
Benchmarks
On my computer, with Go 1.27.1, the standard library solution costs 0.88 ns per operation, while the solution favored by Go maintainers costs 7.14 ns. The unsafe solution matches the performance of the first one.
goos: linux
goarch: amd64
pkg: github.com/vincentbernat/go-netip-addrto6
cpu: AMD Ryzen 5 5600X 6-Core Processor
│ sec/op │
AddrTo6/safe 7.137n ± 0%
AddrTo6/unsafe 0.8682n ± 2%
AddrTo6/builtin 0.8775n ± 2%
Assembly code
Let's check the assembly code the compiler generates for each solution.5 The one built into the standard library looks like this:6
// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
CMPQ net/netip·z4(SB), CX ; check "z" if this is an IPv4 address
JNE end ; if not, stop here
MOVQ net/netip·z6noz(SB), CX ; CX = netip.z6noz
end:
RET
// return value = Addr{hi: AX, lo: BX, z: CX}
Go's assembly language is not a direct representation of the underlying machine language: it operates on a semi-abstract instruction set derived from Plan 9's assembler. It has four pseudo-registers: FP (frame pointer for function arguments), PC (program counter), SB (static base pointer for global symbols), and SP (stack pointer). It also has architecture-specific registers like AX, CX, DX, BX, SI, DI, and R8 to R15. Instructions storing data use their last argument as the destination. Instructions can carry an explicit size suffix: MOVB moves a byte, MOVW 16 bits, MOVL 32 bits, and MOVQ 64 bits. In the example above, the first instruction compares the 64-bit value z4 with the CX register.
The unsafe solution looks almost the same:
// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
CMPQ net/netip·z4(SB), CX ; check "z" if this is an IPv4 address
JNE end ; if not, stop here
MOVQ netipZ6noz(SB), CX ; CX = netip.z6noz
end:
RET
// return value = Addr{hi: AX, lo: BX, z: CX}
The helper solution has far more instructions. To understand why, let's look at the code for As16() and AddrFrom16(). They are short enough for the compiler to inline them.
func (ip Addr) As16() (a16 [16]byte) {
byteorder.BEPutUint64(a16[:8], ip.addr.hi)
byteorder.BEPutUint64(a16[8:], ip.addr.lo)
return a16
}
func AddrFrom16(addr [16]byte) Addr {
return Addr{
addr: uint128{
byteorder.BEUint64(addr[:8]),
byteorder.BEUint64(addr[8:]),
},
z: z6noz,
}
}
We can already guess the pattern to optimize: the code packs the IP address into an array, copies it, then unpacks it. If we inline the Go code by hand, we get:
func AddrTo6(input netip.Addr) netip.Addr {
if !input.Is4() {
return input
}
var a16 [16]byte
byteorder.BEPutUint64(a16[:8], input.addr.hi)
byteorder.BEPutUint64(a16[8:], input.addr.lo)
addr := a16
var output netip.Addr
output.addr.hi = byteorder.BEUint64(addr[:8])
output.addr.lo = byteorder.BEUint64(addr[8:])
output.z = netip.z6noz
return output
}
As humans, we can mentally derive the optimized form:
func AddrTo6(input netip.Addr) netip.Addr {
if !input.Is4() {
return input
}
var output netip.Addr
output.addr.hi = input.addr.hi
output.addr.lo = input.addr.lo
output.z = netip.z6noz
return output
}
Unfortunately, as of Go 1.26.8, the compiler is not smart enough to do the same:
// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
; Push the stack (32 bytes):
; 0(SP) addr netip.uint128
; 16(SP) a16 [16]byte
PUSHQ BP
MOVQ SP, BP
SUBQ $32, SP
CMPQ net/netip·z4(SB), CX ; check "z" if this is an IPv4 address
JNE end ; if not, stop here
; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)
; byteorder.BEPutUint64(a16[8:], input.addr.lo)
MOVBEQ AX, net/netip·a16+16(SP)
MOVBEQ BX, net/netip·a16+24(SP)
; addr = a16, 16 bytes at once through the vector register X0
MOVUPS net/netip·a16+16(SP), X0
MOVUPS X0, net/netip·addr(SP)
; CX = netip.z6noz
MOVQ net/netip·z6noz(SB), CX
; Unpack: output.addr.hi = byteorder.BEUint64(addr[:8])
; output.addr.lo = byteorder.BEUint64(addr[8:])
MOVBEQ net/netip·addr(SP), AX
MOVBEQ net/netip·addr+8(SP), BX
end:
ADDQ $32, SP
POPQ BP
RET
// return value = Addr{hi: AX, lo: BX, z: CX}
The compiler does a decent job on the byte shuffling: the eight byte stores of BEPutUint64() become a single MOVBEQ, which stores a register byte-swapped. The eight byte loads of BEUint64() become a single MOVBEQ the other way round.7 Three groups of instructions remain: a pack, a copy, and an unpack.
Hacking the Go compiler
The Go compiler has several phases:
- Parsing
- The compiler tokenizes and parses the source code. It builds a syntax tree for each source file.
- Type checking
- The compiler maps each identifier to the object it denotes, folds constants, and infers the type of every expression.
- IR construction
- The compiler converts the syntax tree and its types into its own intermediate representation (IR). This process, called "noding," goes through a serialization format named unified IR.
- Middle end
- The compiler performs several optimization passes on the IR, such as devirtualization, function call inlining, and escape analysis.
- Walk
- This phase runs two steps: order of evaluation decomposes complex statements into simpler ones, and desugaring transforms higher-level Go constructs, like
switchor channels, into more primitive instructions or calls to the runtime. - Generic SSA
- The compiler converts the IR into Static Single Assignment (SSA) form, a lower-level intermediate representation suited for machine-independent optimizations and rewrite rules.
- Machine code generation
- The compiler rewrites the SSA form into machine-specific variants, allocates registers, and applies more optimization passes. At the end, the assembler turns the generated instructions into machine code.
The hammer
My first idea is to replace occurrences of netip.AddrFrom16(ip.As16()) with netip.Addr{addr: ip.addr, z: netip.z6noz} as early as possible, during the "noding" process. Before that, the type checking phase prevents us from accessing unexported struct fields.
Go 1.27 introduced a convenient debug option to dump the IR of a function at interesting points during compilation:
$ GOTOOLCHAIN=go1.27.1 GOAMD64=v3 go build -a -gcflags="-d=astdump=AddrTo6Safe" .
Writing text ast output for AddrTo6Safe to AddrTo6Safe.ast
Writing html ast output for AddrTo6Safe to AddrTo6Safe.html
Writing html syntax output for AddrTo6Safe to AddrTo6Safe.syntax.html
In the HTML file, the first column shows the IR as it comes out of noding:
DCLFUNC addrto6.AddrTo6Safe ABI:ABIInternal FUNC-func(netip.Addr) netip.Addr
DCLFUNC-Dcl
. NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. NAME-addrto6.~r0 Class:PPARAMOUT Offset:0 OnStack netip.Addr
DCLFUNC-body
. IF # ipv6_safe.go:11:2
. IF-Cond
. . CALLFUNC bool
. . CALLFUNC-Fun
. . . METHEXPR addrto6.Is4 FUNC-func(netip.Addr) bool
. . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr
. . CALLFUNC-Args
. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. IF-Body
. . AS # ipv6_safe.go:12:6
. . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . . CALLFUNC netip.Addr
. . . CALLFUNC-Fun
. . . . NAME-netip.AddrFrom16 Class:PFUNC Offset:0 Used FUNC-func([16]byte) netip.Addr
. . . CALLFUNC-Args
. . . . CALLFUNC ARRAY-[16]byte
. . . . CALLFUNC-Fun
. . . . . METHEXPR addrto6.As16 FUNC-func(netip.Addr) [16]byte
. . . . . . TYPE netip.Addr Class:PEXTERN Offset:0 type netip.Addr
. . . . CALLFUNC-Args
. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. RETURN # ipv6_safe.go:14:2
. RETURN-Results
. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
In the body of the if statement, we spot the calls to the method netip.Addr.As16() and to the function netip.AddrFrom16(). Our goal is to patch them with a struct literal:
IF-Body
. AS # ipv6_safe.go:12:6
. . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . STRUCTLIT netip.Addr
. . STRUCTLIT-List
. . . STRUCTKEY netip.addr
. . . . DOT netip.addr netip.uint128
. . . . . NAME-addrto6.ip Class:PPARAM Offset:0 OnStack Used netip.Addr
. . . STRUCTKEY netip.z
. . . . NAME-netip.z6noz Class:PEXTERN Offset:0 unique.Handle[net/netip.addrDetail]
In noder's reader.go, the expr() method builds the IR tree for an expression. At the end of the exprCall case, we add a call to a rewriteAddrFrom16As16() function. It takes the current node and returns the struct literal on success, or nil if the rewrite is not possible. First, we check that we have the expected pattern: a call to the netip.AddrFrom16() function with a call to the netip.Addr.As16() method as its only argument:
func rewriteAddrFrom16As16(n ir.Node) ir.Node {
call, ok := n.(*ir.CallExpr)
if !ok || call.Op() != ir.OCALLFUNC ||
len(call.Args) != 1 || len(call.Init()) != 0 ||
!isNetipFunc(call.Fun, "AddrFrom16") {
return nil
}
inner, ok := call.Args[0].(*ir.CallExpr)
if !ok || inner.Op() != ir.OCALLFUNC ||
len(inner.Args) != 1 || len(inner.Init()) != 0 ||
!isNetipFunc(inner.Fun, "Addr.As16") {
return nil
}
x := inner.Args[0]
// [...]
}
Then, we fetch netip.z6noz:
z6noz, err := lookupVar(ir.StaticCalleeName(call.Fun).Sym().Pkg, "z6noz")
if err != nil {
return nil
}
And we build the struct literal:
typ := call.Type()
pos := call.Pos()
var list []ir.Node
for i, f := range typ.Fields() {
var value ir.Node
switch f.Sym.Name {
case "addr":
value = typecheck.DotField(pos, x, i)
case "z":
value = z6noz
default:
return nil
}
list = append(list, ir.NewStructKeyExpr(pos, f, value))
}
lit := ir.NewCompLitExpr(pos, ir.OSTRUCTLIT, typ, list)
lit.SetTypecheck(1)
return lit
Have a look at the complete patch.8 We can test it with the following commands:
$ cd src
$ ./make.bash
Building Go cmd/dist using /usr/lib/go-1.27. (go1.27.1 linux/amd64)
Building Go toolchain1 and bootstrap cmd/go (go_bootstrap) using /usr/lib/go-1.27.
Building Go toolchain2 using go_bootstrap and Go toolchain1.
Building Go toolchain3 and commands using go_bootstrap and Go toolchain2.
Checking command staleness for linux/amd64.
---
Installed Go for linux/amd64 in /home/bernat/code/free/go
Installed commands in /home/bernat/code/free/go/bin
*** You need to add /home/bernat/code/free/go/bin to your PATH.
$ export PATH=$PWD/../bin:$PATH
$ go version
go version go1.28-devel_9834516e20 Sat Sep 12 08:23:11 2026 -0700 linux/amd64
$ go test net/netip/...
ok net/netip 0.224s
$ cd ../../go-netip-addrto6
$ go test .
ok github.com/vincentbernat/go-netip-addrto6 0.062s
The generated code for the helper is now the shortest possible version!
// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
CMPQ net/netip·z4(SB), CX ; check "z" if this is an IPv4 address
JNE end ; if not, stop here
MOVQ net/netip·z6noz(SB), CX ; CX = netip.z6noz
end:
RET
// return value = Addr{hi: AX, lo: BX, z: CX}
Go maintainers are unlikely to accept this change. It relies on the internal structure of net/netip.Addr. It's an ugly hack in the noder, whose job is to faithfully translate the type-checked AST into the IR. And it's harder to maintain than adding a To6() method.
The screwdriver
The right place for such an optimization is the generic SSA phase. One of the last machine-independent passes is memcombine. With the appropriate debug flag, the compiler dumps the SSA form after this pass:9
$ GOTOOLCHAIN=go1.26.8 GOAMD64=v3 \
> go build -a -gcflags='-d=ssa/memcombine/dump=AddrTo6Safe' .
$ head -5 AddrTo6Safe_01__memcombine.dump
AddrTo6Safe func(netip.Addr) netip.Addr
b2:
(?) v1 = InitMem <mem>
(?) v2 = SP <uintptr>
(?) v3 = SB <uintptr>
The result of the memcombine pass follows the same structure as the assembly code for AddrTo6Safe() we looked at earlier: two stores, one move, and two loads we would like to optimize away.
; […]
v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi
v490 = ArgIntReg <uint64> {ip+8} [1] ; input.addr.lo
v466 = ArgIntReg <*netip.addrDetail> {ip+16} [2] ; input.z
; […]
v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16
v173 = OffPtr <*byte> [8] v22 ; &a16[8]
v442 = Bswap64 <uint64> v490 ; bswap(input.addr.lo)
v542 = Bswap64 <uint64> v502 ; bswap(input.addr.hi)
v161 = Store <mem> {uint64} v22 v542 v23 ; a16[:8] = bswap(hi)
v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo)
v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16
v415 = OffPtr <*byte> [8] v285 ; &addr[8]
v416 = Load <uint64> v285 v286 ; addr[:8]
v299 = Bswap64 <uint64> v416 ; output.addr.hi
v174 = Load <uint64> v415 v286 ; addr[8:]
v39 = Bswap64 <uint64> v174 ; output.addr.lo
; […]
Each line features a value identifier (v442), an operation with its type (Bswap64 <uint64>), and its arguments (v490).10 Values are the basic building blocks of SSA and are defined exactly once. Square brackets enclose integer parameters ([8]) and curly braces contain auxiliary arguments ({netip.addr}). Operations writing to memory produce a new memory state. Every memory operation takes the current state as its last argument, which keeps them in order.
On paper
Let's focus on output.addr.lo, aka v39:
v490 = ArgIntReg <uint64> {ip+8} [1] ; input.addr.lo
v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16
v173 = OffPtr <*byte> [8] v22 ; &a16[8]
v442 = Bswap64 <uint64> v490 ; bswap(input.addr.lo)
v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo)
v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16
v415 = OffPtr <*byte> [8] v285 ; &addr[8]
v174 = Load <uint64> v415 v286 ; addr[8:]
v39 = Bswap64 <uint64> v174 ; output.addr.lo
To simplify this code, we could apply three rewriting rules:
-
The first one adds a shortcut when loading through a move:
(Load (OffPtr [o] p) (Move p src mem)) => (Load (OffPtr [o] src) mem). This matchesv174with its argumentsv415andv286and creates a new valuev600:v490 = ArgIntReg <uint64> {ip+8} [1] ; input.addr.lo v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16 v173 = OffPtr <*byte> [8] v22 ; &a16[8] v442 = Bswap64 <uint64> v490 ; bswap(input.addr.lo) v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo) v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16 v415 = OffPtr <*byte> [8] v285 ; &addr[8] v600 = OffPtr <*byte> [8] v22 ; &a16[8] v174 = Load <uint64> v600 v282 ; a16[8:] v39 = Bswap64 <uint64> v174 ; output.addr.lo -
The second one simplifies a load following a store:
(Load p (Store p x _)) => x. The load is forwarded: the stored value replaces it and no memory access remains. It matchesv174. It notices thatv600andv173are the same address and replacesv174with a copy ofv442:v490 = ArgIntReg <uint64> {ip+8} [1] ; input.addr.lo v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16 v173 = OffPtr <*byte> [8] v22 ; &a16[8] v442 = Bswap64 <uint64> v490 ; bswap(input.addr.lo) v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo) v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16 v415 = OffPtr <*byte> [8] v285 ; &addr[8] v600 = OffPtr <*byte> [8] v22 ; &a16[8] v174 = Copy <uint64> v442 ; bswap(input.addr.lo) v39 = Bswap64 <uint64> v174 ; output.addr.lo -
The last step cancels the two byte swaps:
(Bswap64 (Bswap64 x)) => x.v39becomes a copy ofv490:v490 = ArgIntReg <uint64> {ip+8} [1] ; input.addr.lo v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16 v173 = OffPtr <*byte> [8] v22 ; &a16[8] v442 = Bswap64 <uint64> v490 ; bswap(input.addr.lo) v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo) v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16 v415 = OffPtr <*byte> [8] v285 ; &addr[8] v600 = OffPtr <*byte> [8] v22 ; &a16[8] v174 = Copy <uint64> v442 ; bswap(input.addr.lo) v39 = Copy <uint64> v490 ; output.addr.lo = input.addr.lo
If we ignore the values not needed to compute v39, only this SSA form remains:
v490 = ArgIntReg <uint64> {ip+8} [1] ; input.addr.lo
v39 = Copy <uint64> v490 ; output.addr.lo = input.addr.lo
Let's switch to output.addr.hi, aka v299:
v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi
v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16
v173 = OffPtr <*byte> [8] v22 ; &a16[8]
v542 = Bswap64 <uint64> v502 ; bswap(input.addr.hi)
v161 = Store <mem> {uint64} v22 v542 v23 ; a16[:8] = bswap(hi)
v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo)
v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16
v416 = Load <uint64> v285 v286 ; addr[:8]
v299 = Bswap64 <uint64> v416 ; output.addr.hi
To optimize it away, we also apply three rewriting rules:
-
The first one also adds a shortcut when loading through a move, but without an offset:
(Load p (Move p src mem)) => (Load src mem). This rewritesv416to use arguments fromv286:v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16 v173 = OffPtr <*byte> [8] v22 ; &a16[8] v542 = Bswap64 <uint64> v502 ; bswap(input.addr.hi) v161 = Store <mem> {uint64} v22 v542 v23 ; a16[:8] = bswap(hi) v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo) v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16 v416 = Load <uint64> v22 v282 ; a16[:8] v299 = Bswap64 <uint64> v416 ; output.addr.hi -
The second rule forwards a value stored one step earlier, skipping over a store to another address:
(Load p (Store q _ (Store p x _))) => x. This matchesv416:xisv542,pisv22(&a16),qisv173(&a16[8]), andpandqdo not overlap foruint64.v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16 v173 = OffPtr <*byte> [8] v22 ; &a16[8] v542 = Bswap64 <uint64> v502 ; bswap(input.addr.hi) v161 = Store <mem> {uint64} v22 v542 v23 ; a16[:8] = bswap(hi) v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo) v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16 v416 = Copy <uint64> v542 ; bswap(input.addr.hi) v299 = Bswap64 <uint64> v416 ; output.addr.hi -
The third rule cancels two byte swaps:
(Bswap64 (Bswap64 x)) => x.v299becomes a copy ofv502:v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi v22 = LocalAddr <*[16]byte> {netip.a16} v2 v1 ; &a16 v173 = OffPtr <*byte> [8] v22 ; &a16[8] v542 = Bswap64 <uint64> v502 ; bswap(input.addr.hi) v161 = Store <mem> {uint64} v22 v542 v23 ; a16[:8] = bswap(hi) v282 = Store <mem> {uint64} v173 v442 v161 ; a16[8:] = bswap(lo) v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16 v416 = Copy <uint64> v542 ; bswap(input.addr.hi) v299 = Copy <uint64> v502 ; output.addr.hi = input.addr.hi
If we remove the values not used to compute v299, we get this SSA form:
v502 = ArgIntReg <uint64> {ip+0} [0] ; input.addr.hi
v299 = Copy <uint64> v502 ; output.addr.hi = input.addr.hi
In practice
Most of these rules already exist in generic.rules. They use conditions to validate their context: ssa.IsSamePtr() for the same address, ssa.Disjoint() for addresses that do not overlap. The rule forwarding a stored value to a load already exists with three variants looking through several other stores. Here are the two we need:
(Load <t1> p1 (Store {t2} p2 x _))
&& ssa.IsSamePtr(p1, p2)
&& copyCompatibleType(t1, x.Type)
&& t1.Size() == t2.Size()
=> x
(Load <t1> p1 (Store {t2} p2 _ (Store {t3} p3 x _)))
&& ssa.IsSamePtr(p1, p3)
&& copyCompatibleType(t1, x.Type)
&& t1.Size() == t3.Size()
&& ssa.Disjoint(p3, t3, p2, t2)
=> x
Go 1.27 added the rule loading through a move with CL 748200 to fix issue #77720:
(Load <t1> op1:(OffPtr [o1] p1) move:(Move [n] p2 src mem))
&& o1 >= 0 && o1+t1.Size() <= n && ssa.IsSamePtr(p1, p2)
&& !ssa.IsVolatile(src)
=> @move.Block (Load <t1> (OffPtr <op1.Type> [o1] src) mem)
It lacks a variant without an offset:
(Load <t1> p1 move:(Move [n] p2 src mem))
&& p1.Op != ssaop.OpOffPtr
&& t1.Size() <= n && ssa.IsSamePtr(p1, p2)
&& !ssa.IsVolatile(src)
=> @move.Block (Load <t1> (OffPtr <p1.Type> [0] src) mem)
There is no generic rule to cancel two byte swaps, but the AMD64 lowering pass includes this rule:
(BSWAP(Q|L) (BSWAP(Q|L) p)) => p
After switching to Go's development branch and adding the missing rule, the generated assembly code is worse than with Go 1.26.8, even though our additional rule slightly improves the situation at the end:
// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
; Push the stack (16 bytes):
; 0(SP) a16 [16]byte
PUSHQ BP
MOVQ SP, BP
SUBQ $16, SP
CMPQ net/netip·z4(SB), CX ; check "z" if this is an IPv4 address
JNE end ; if not, stop here
; The four forwarded bytes: the low half of input.addr.lo is taken apart and
; put back together in registers
MOVQ BX, DX ; DX = input.addr.lo
SHRQ $24, BX ; BX = input.addr.lo >> 24
MOVQ DX, SI ; SI = input.addr.lo
SHRQ $16, DX ; DX = input.addr.lo >> 16
MOVQ SI, DI ; DI = input.addr.lo, kept for the pack
SHRQ $8, SI ; SI = input.addr.lo >> 8
MOVBLZX DIB, R8 ; R8 = byte(input.addr.lo)
MOVBLZX SIB, SI ; SI = byte(input.addr.lo >> 8)
SHLQ $8, SI
ORQ R8, SI ; SI = two low bytes of input.addr.lo
MOVBLZX DL, DX ; DX = byte(input.addr.lo >> 16)
SHLQ $16, DX
ORQ SI, DX
MOVBLZX BL, BX ; BX = byte(input.addr.lo >> 24)
SHLQ $24, BX
ORQ DX, BX ; BX = input.addr.lo & 0xffffffff
; Pack: byteorder.BEPutUint64(a16[:8], input.addr.hi)
; byteorder.BEPutUint64(a16[8:], input.addr.lo)
MOVBEQ AX, net/netip·a16(SP)
MOVBEQ DI, net/netip·a16+8(SP)
; The four other bytes of input.addr.lo, read one by one from a16
MOVBLZX net/netip·a16+11(SP), DX ; a16[11]
SHLQ $32, DX
ORQ DX, BX
MOVBLZX net/netip·a16+10(SP), DX ; a16[10]
SHLQ $40, DX
ORQ DX, BX
MOVBLZX net/netip·a16+9(SP), DX ; a16[9]
SHLQ $48, DX
ORQ DX, BX
MOVBLZX net/netip·a16+8(SP), DX ; a16[8]
SHLQ $56, DX
; output.z = netip.z6noz
MOVQ net/netip·z6noz(SB), CX
; output.addr.hi = byteorder.BEUint64(a16[:8])
MOVBEQ net/netip·a16(SP), AX
; output.addr.lo assembled from the previous steps
ORQ DX, BX
end:
LEAVEQ
RET
// return value = Addr{hi: AX, lo: BX, z: CX}
The rule loading through a move, added in Go 1.27, introduced this regression.
Out of order
Let's not give up now! In reality, the rewriting rules run before memcombine, notably in the late opt pass. At this point, the inlined versions of BEPutUint64() and BEUint64() still expand to sixteen byte stores and sixteen byte loads, matching their source code:
func BEUint64(b []byte) uint64 {
_ = b[7] // bounds check hint to compiler; see golang.org/issue/14808
return uint64(b[7]) | uint64(b[6])<<8 | uint64(b[5])<<16 | uint64(b[4])<<24 |
uint64(b[3])<<32 | uint64(b[2])<<40 | uint64(b[1])<<48 | uint64(b[0])<<56
}
Let's follow two bytes of output.addr.lo: addr[15] and addr[11]. Here is a simplified SSA form before late opt:
v273 = Trunc64to8 <byte> v490 ; byte(input.addr.lo)
v226 = Trunc64to8 <byte> v225 ; byte(input.addr.lo >> 32)
; […]
v235 = Store <mem> {byte} v233 v226 v223 ; a16[11] = byte(lo >> 32)
v247 = Store <mem> {byte} v245 v238 v235 ; a16[12] = …
v259 = Store <mem> {byte} v257 v250 v247 ; a16[13] = …
v271 = Store <mem> {byte} v269 v262 v259 ; a16[14] = …
v282 = Store <mem> {byte} v280 v273 v271 ; a16[15] = byte(lo)
v285 = LocalAddr <*[16]byte> {netip.addr} v2 v282 ; &addr
v286 = Move <mem> {[16]byte} [16] v285 v22 v282 ; addr = a16
; […]
v433 = OffPtr <*byte> [15] v285 ; &addr[15]
v435 = Load <byte> v433 v286 ; addr[15]
v479 = OffPtr <*byte> [11] v285 ; &addr[11]
v481 = Load <byte> v479 v286 ; addr[11]
The first rule loads through the move: (Load (OffPtr [o] p) (Move p src mem)) => (Load (OffPtr [o] src) mem). It matches both loads, which now read a16 with the memory state before the copy:
v600 = OffPtr <*byte> [15] v22 ; &a16[15]
v435 = Load <byte> v600 v282 ; a16[15]
v601 = OffPtr <*byte> [11] v22 ; &a16[11]
v481 = Load <byte> v601 v282 ; a16[11]
The second rule shortcuts a load following a store: (Load p (Store p x _)) => x. It matches v435, as v282 stores a16[15]. It does not match v481: v235 stores a16[11] four stores earlier in the chain, while the variants of this rule look through three stores at most.
v435 = Copy <byte> v273 ; byte(input.addr.lo)
v601 = OffPtr <*byte> [11] v22 ; &a16[11]
v481 = Load <byte> v601 v282 ; a16[11]
The same happens to the other bytes: the rule forwards the four bytes stored last, a16[12] to a16[15]. The twelve other loads now read a16 instead of addr.
BEUint64() becomes a chain of Or64, each one adding a byte shifted into place. memcombine is a pass written in Go, not a set of rewrite rules. It starts from the last Or64 of the chain and collects up to eight terms. If each term is a byte load, extended to 64 bits and shifted, and if the eight loads read consecutive addresses from the same pointer with the same memory state, it replaces the whole chain with a single 64-bit load and a byte swap. Otherwise, it tries again with four, then two terms, and from each intermediate Or64. Here is the loop checking each term in a simplified version of combineLoads():
for i := int64(0); i < n; i++ {
v := a[i]
shift := int64(0)
if v.Op == shiftOp {
v, shift = peelShift(v)
}
if v.Op != extOp {
return false
}
load := v.Args[0]
if load.Op != ssaop.OpLoad {
return false
}
if load.Args[1] != mem {
return false
}
p, off := splitPtr(load.Args[0])
if p != base {
return false
}
r[i] = LoadRecord{load: load, offset: off, shift: shift}
}
For output.addr.hi, the eight loads read a16 with the same memory state v282:
v13 = Load <byte> v22 v282 ; a16[0]
v530 = Load <byte> v14 v282 ; a16[1]
v488 = Load <byte> v504 v282 ; a16[2]
v405 = Load <byte> v537 v282 ; a16[3]
v385 = Load <byte> v397 v282 ; a16[4]
v361 = Load <byte> v373 v282 ; a16[5]
v196 = Load <byte> v63 v282 ; a16[6]
v432 = Load <byte> v315 v282 ; a16[7]
v319 = ZeroExt8to64 <uint64> v432 ; uint64(a16[7])
v329 = ZeroExt8to64 <uint64> v196 ; uint64(a16[6])
v330 = Lsh64x64 <uint64> [true] v329 v138 ; uint64(a16[6]) << 8
v331 = Or64 <uint64> v319 v330 ; a16[7] | a16[6] << 8
; […] same for a16[5] to a16[1]
v401 = ZeroExt8to64 <uint64> v13 ; uint64(a16[0])
v402 = Lsh64x64 <uint64> [true] v401 v55 ; uint64(a16[0]) << 56
v403 = Or64 <uint64> v402 v391 ; | a16[0] << 56 = output.addr.hi
memcombine merges them into one load and a swap:
v286 = Load <uint64> v22 v282 ; a16[:8]
v285 = Bswap64 <uint64> v286 ; output.addr.hi
For output.addr.lo, here is the chain memcombine sees after late opt:
v436 = ZeroExt8to64 <uint64> v273 ; addr[15], forwarded
v448 = Or64 <uint64> v436 v447 ; | addr[14] << 8, forwarded
v460 = Or64 <uint64> v459 v448 ; | addr[13] << 16, forwarded
v472 = Or64 <uint64> v471 v460 ; | addr[12] << 24, forwarded
v484 = Or64 <uint64> v483 v472 ; | a16[11] << 32, loaded
v496 = Or64 <uint64> v495 v484 ; | a16[10] << 40, loaded
v508 = Or64 <uint64> v507 v496 ; | a16[9] << 48, loaded
v520 = Or64 <uint64> v519 v508 ; | a16[8] << 56, loaded
From v520, four of the eight terms are forwarded bytes, not loads from memory, and memcombine can't combine them. It doesn't merge the four remaining loads either, as they sit on top of the forwarded bytes.
Back in order
In summary, the rewriting rules run too early to be effective. A quick workaround exists: run an earlier round of memcombine before late opt. After this change, the generated code for the helper is back to the shortest possible version:
// AX = input.addr.hi, BX = input.addr.lo, CX = input.z
CMPQ net/netip·z4(SB), CX ; check "z" if this is an IPv4 address
JNE end ; if not, stop here
MOVQ net/netip·z6noz(SB), CX ; CX = netip.z6noz
end:
RET
// return value = Addr{hi: AX, lo: BX, z: CX}
And the benchmark confirms it! ✌️
goos: linux
goarch: amd64
pkg: github.com/vincentbernat/go-netip-addrto6
cpu: AMD Ryzen 5 5600X 6-Core Processor
│ Go 1.26.8 │ Our branch │
│ sec/op │ sec/op vs base │
AddrTo6/safe 6.5470n ± 0% 0.8944n ± 4% -86.34% (p=0.002 n=6)
AddrTo6/unsafe 0.9071n ± 3% 0.8682n ± 1% -4.28% (p=0.002 n=6)
AddrTo6/builtin 0.8871n ± 2% 0.8785n ± 1% ~ (p=0.310 n=6)
Next steps
I think Go maintainers would reject this change because of the additional memcombine pass. Instead, I plan to publish this blog post and bring up the subject again as a follow-up to issue #54365. Either the sheer complexity and the Go 1.27 regression convince the maintainers that adding a To6() method is simpler and more efficient, or they advise me on how to move forward. Either way, digging into this subject taught me a lot about the Go compiler! ⚙️
Update (2026-10)
I opened issue #81994 to propose Addr.To6(). Give it a 👍 if you want it in Go!
-
IPv4-mapped IPv6 addresses let you handle only IPv6 addresses in your code and keep conversions at a few well-defined boundaries. I use them in Akvorado. ↩
-
This is not strictly equivalent. The zero value becomes
::while it would make more sense to leave it untouched. ↩ -
I suppose Go maintainers keep it inconvenient on purpose: they don't want random packages to patch the standard library or other packages. ↩
-
To avoid catastrophic bugs when
netip.Addr's layout changes, you must add tests to detect it. This is more dangerous if you put this code in a package. Either ask users to run the tests themselves, or detect the change at run time and panic. Hiding this optimization behind a build tag would make users aware of this potential trap. ↩ -
I produced the assembly code with
GOTOOLCHAIN=go1.26.8 GOAMD64=v3 ./go-asm '\.asm'from the companion repository, then edited it a bit to keep this article from turning into an endless rabbit hole. Due to an unfortunate sequence of events, we switch between versions: Go 1.27 introduced a regression that blurs the point of this article. ↩ -
The function is short enough for the compiler to inline it, so the assembly code may vary depending on the surrounding code. ↩
-
MOVBEQrequiresGOAMD64=v3, matching the x86-64-v3 microarchitecture from 2013. Otherwise, the compiler translatesBEPutUint64()toBSWAPQ+MOVQandBEUint64()toMOVQ+BSWAPQ. ↩ -
When the call is used as a statement, a struct literal alone is invalid. The patch then turns it into
_ = netip.Addr{…}. ↩ -
The compiler can also write an HTML file with the result of each pass:
$ GOSSAFUNC=AddrTo6Safe \ > GOTOOLCHAIN=go1.26.8 \ > GOAMD64=v3 go build -a . dumped SSA for AddrTo6Safe,1 to ./ssa.html -
Source line numbers in parentheses follow value identifiers, but I removed them from the examples. The SSA form also features branches, but we don't need them to understand our optimizations. ↩
04 Oct 2026 2:12pm GMT