02 Oct 2026
planet.freedesktop.org
Mike Blumenkrantz: Welcome Robot Overlord
Hm
I'm no stranger to big tree-wide refactors. I've done lots of them. They suck for everyone.
That's why this time I decided to hand over this job to the new high-tech calculator, Claude.
We've had a lot of discussion about AI use in mesa, and while I can appreciate the various factions, I came out sort of vaguely in favor of allowing AI as long as it was used with some discrimination (i.e., a calculator exists to decrease the workload of humans, not increase it). But I had never actually used it other than asking the free models to find me spec references for various things.
Today, that changed.
Process
I don't know if anyone will find this useful/interesting, but I'm going to document the process I used today.
Install Claude Code VSCode plugin
It's just the normal one from Anthropic. I linked it to my Claude subscription, where I have a 7-day free trial for Claude Pro activated.
Prompt
My exact prompt was as follows:
in the mesa repo I have open:
* do not create or alter any git commits
* delete support for big endian
This took a long while to execute, and it printed out lots of in-progress stuff it was doing. Some neat things I observed along the way were how it refused to modify src/gallium/drivers/asahi due to in-tree policy documentation and how it wrote a script to mimic unifdef when it was unable to find my kernel-devel install directory. One of the times it ran the homegrown script resulted in a fuckup related to line-wrapping, which it detected and then fixed. Cool.
When it was finally finished, it started a test build to verify. The build passed.
Organize
I glanced over all the changes, and they seemed reasonable. My next prompt:
organize the big endian removal into logical git commits:
* each commit must be independently buildable
* each commit must "make sense"
* add `Assisted-by: Claude Code (Opus 5.5 Medium)` at the end of every commit log
This also took a while, but it ended up with 15 commits and refused to write commit logs for them because apparently that's against our AI policy. It then tried to kick off builds to verify the bisectability.
Unfortunately, my main workstation is still an ICL laptop from 2019, so this would've taken literal hours. I told Claude it could ssh to my big test machine with a threadripper for this step. It worked.
Document
Drunk with power, I asked Claude:
briefly describe the effects of each commit
Then I confirmed it did actually reflect the contents of each commit, massaged the output a bit, and that became the commit logs.
Fixups
Our AI policy allows individual drivers/components to ban AI-assisted changes. For this reason, pipe_cap::endianness could not be removed by Claude. I did this manually. It didn't take long.
Done?
In total, this took maybe an hour for me, an actual organic human, to get big endian completely removed from mesa. Doing it manually would've taken days, so I think this is a win.
https://gitlab.freedesktop.org/mesa/mesa/-/merge_requests/44863
02 Oct 2026 12:00am GMT
01 Oct 2026
planet.freedesktop.org
Tomeu Vizoso: Ethos-U NPU update 2: Executorch support
Over the last few months I have been working on Torx, an ExecuTorch backend in Mesa3D that delegates to the NPU drivers Mesa already contains. Initial Torx support has been merged already, with YOLOX support being currently under review.
Torx sits next to Teflon, the existing TensorFlow Lite delegate, and both use the same Gallium ML API. Today Torx works with the ethosu driver.
Torx implements the backend API that ExecuTorch provides for running portions of a model's graph on specialized hardware such as NPUs or GPUs. It then uses the Gallium ML API to compile and execute those portions on the available hardware.
Unlike TensorFlow Lite, ExecuTorch compiles models in a step separate from execution. The model can therefore be compiled on a different machine from the target, where it is practical to apply more expensive optimizations than edge-class hardware could run.
I worked on this alongside Rob Herring from Arm, who supported the effort on the kernel side.
Below is a demo of object detection with YOLOX-Nano at 416x416, quantized to 8 bits integers, on an i.MX93 board with an Arm Ethos-U65 NPU. Inference takes 17 milliseconds per frame.
The convolutional layers of YOLOX run entirely on the U65. The final decoding steps need floating point operations, so they run on the CPU.
Once the other NPU drivers in Mesa support the ahead-of-time compilation extension to the Gallium ML API, they will be usable from ExecuTorch as well.
To try it yourself, you can follow the instructions in the documentation. If you prefer not to build ExecuTorch from sources, you can use this merge request from Meta's Anthony Shoumikhin.
01 Oct 2026 12:09pm GMT
28 Sep 2026
planet.freedesktop.org
Hans de Goede: Snapdragon X1 laptop kernel COPR for Fedora 45
I'm happy to announce the availability of a COPR repository which rebuilds Fedora 45 kernels with the patches from the qcom-laptops git branch added. This is a branch where various upstream pending patches with Snapdragon laptop improvements are gathered. For the exact contents of the latest kernel from this COPR see this Fedora kernel git-repo fork.
Installing the kernel from this COPR adds camera support for various X1 laptop models, adds bluetooth support for the ThinkPad T14s as well various other improvements. Most of these are expected to land in the 7.4 kernel.
comments
28 Sep 2026 1:28pm GMT
Hans de Goede: Snapdragon X1 laptops improvements in Fedora 45
After the initial work to make Fedora live iso media boot OOTB on Snapdragon X1 laptops in Fedora 44, I've continued working on improving the Fedora experience on these laptops for Fedora 45. For Fedora 44 a long list of workarounds was necessary, see the Snapdragon laptop install instructions on the wiki.
For Fedora 45 various improvements have been made in the mainline kernel and I've been working on improvements at the initramfs generator / distro level. Together these result in a much smoother experience, see the install instructions for Fedora 45 on the wiki. A few workarounds are unfortunately still necessary for Fedora 45. For Fedora 46 I hope that the Fedora aarch64 live iso media will just work on X1 laptops.
Note that this blog post is about Snapdragon X1 (Elite,Plus) laptops. Support for Snapdragon X2 laptops at the same level as the current X1 laptop support is quickly coming together in the mainline kernel. But the 7.2 kernel used for the Fedora 45 beta and release isos is still missing some bits and Devicetrees, so Fedora 45 will not work OOTB on Snapdragon X2 laptops.
comments
28 Sep 2026 1:06pm GMT
25 Sep 2026
planet.freedesktop.org
Mike Blumenkrantz: That Time I Fought Draw And Won
Tessellate Your Nightmares
Way, way, way back in the day, there was EXT_shader_object, and inside Big Triangle there was a lot of debate about how things should work. By this time, I had long since infiltrated their organization. They viewed me as one of their own. Because of this, I was able to influence their sinister operation.
My plotting was even more underhanded than the Big Triangle fat-cats, and I nudged them to design tessellation shader objects in a manner which matched OpenGL. Specifically, all the spacing and vertex ordering mechanics were specified in the same shader stages as OpenGL. The D3D members of Big Triangle were asleep at the wheel. My influence went unopposed.
I was the butterfly flapping my wings to cause a hurricane.
That hurricane manifested years later. Suddenly, those sleeping giants awoke and discovered that shader objects were utterly incompatible with their chosen API. Their howls of rage reverberated across the world.
I was unprepared for the backlash. They wielded their monstrous power and upended everything. Suddenly, spacing and vertex ordering could be specified in any tessellation shader stage. It was a nightmare of epic proportions.
Draw Hell
Lavapipe, unbeknownst to many, is really just llvmpipe wearing a funny hat. This means it inherits all the llvmpipe-isms, including its deepest flaws. One flaw is that llvmpipe is primarily a driver for OpenGL rendering, and most of its internals are structured around that. Chief among them is the tessellation support, which goes through auxiliary/draw and then even deeper, into auxiliary/tessellator, a mysterious land where few have tread and even fewer have returned.
And all of this code expects tessellation parameters in the shader stages required by OpenGL.
This was fine due to my initial machinations, but it was no longer fine once the more fiendish parts of Big Triangle awoke. The tessellator was broken. Tests were failing. A new crisis had emerged.
Compatibility
Historically, this mismatch was handled by a function called merge_tess_info. It's still present in a number of drivers. And it did work, propagating that info to the right place in lavapipe before being sent to llvmpipe, except for one wrinkle.
Dynamic domain origin.
OpenGL hardcodes this value to lower-left, but in Vulkan it can be either lower-left or upper-left. Lavapipe worked around this by treating upper-left as equivalent to toggling vertex ordering CCW: for the dynamic state, two versions of the shader were compiled, and lavapipe would run the one which corresponded to (shader_ccw ^ dynamic_domain_ccw). This was fine since it was all restricted to the tessellation evaluation shader.
It was no longer fine once vertex ordering could also be specified in tessellation control.
Another Battle
I had two options: add even more hacks into lavapipe to work around this, or go spelunking deep into gallium to make all the weird bits support setting params in either shader stage. Naturally I went spelunking. This essentially meant shoving merge_tess_info into the depths of auxiliary/draw.
I won't claim it was pretty or easy, but the battle was won. The forces of good have once again triumphed over the evils of Big Triangle's tessellation monster.
25 Sep 2026 12:00am GMT
21 Sep 2026
planet.freedesktop.org
Ricardo Garcia: XDC 2025 CTS Talk
We're only one week away from XDC 2026, which will take place in Toronto from the 28th to the 30th of September. I'll be there one more year together with a sizeable group of Igalia colleagues presenting on a number of different topics related to the open source graphics stack. See the full schedule for more details. I'll open the afternoon lightning talks round on Monday with a brief presentation about a memory usage problem we recently solved in Vulkan CTS. Last year I also gave an overview of Vulkan CTS and listed a bunch of tips and tricks about it.
You can find the slides in the link above but I noticed I never actually embedded my talk in a blog post for posterity, and I didn't provide a transcription of the 2025 talk for those more inclined to read than to watch a video. Please excuse me for being so late. You can find both the video and the transcription below. See you in Toronto!
XDC 2025 Recording
Talk slides and transcription
Thanks everyone! Hello, I'm Ricardo from Igalia and I'm going to talk about Vulkan CTS, giving you some general information and some tips and tricks about it.
I've been contributing to the Vulkan Conformance Test Suite since 2019. It's the same project that has OpenGL tests, but I have almost no contributions to the OpenGL part. I'm not the only one at Igalia working on CTS. Our work on this front is being generously funded by Valve. As part of my job, I frequently interact with Mesa contributors and people at Khronos. There's some overlap between both, specially in this room. So what is my job about?
This.
The project is called VK-GL-CTS and it's a Vulkan and OpenGL Conformance Test Suite. Most people writing Vulkan drivers have had to deal with VK-GL-CTS at some point. It's open source, available on Github under the Apache 2 license. You can see it has a lot of commits from many different contributors apart from us, but not much community or activity inside Github, for reasons that I will explain later. This is very common with Khronos projects. Right now it has around 2.8 million Vulkan tests and ~5 million lines of code specific to Vulkan, but when you run a test you also use a lot of code from a common base that's shared with OpenGL.
To give you some specific details about the Vulkan tests, if you take a look at the top-level README file, it points you to a second README that's specific to Vulkan. If you follow the quick instructions in that second README and you configure and build the project, you end up with a binary called deqp-vk in some subdirectory of the build directory. That's the base binary that allows you to run Vulkan tests by passing the -n option and indicating a test name, but you can also use asterisks in the name to form a glob that allows you to run multiple tests that match that pattern. The program will print a summary of the results in the terminal and put a lot more detail about each test in a file called TestResults.qpa. The tests are organized in a tree, with the leaves of that tree being the test cases that can be run, and the nodes leading to that leaf being the parent test groups.
One example is this: a Vulkan test that tries to see if copying between images with some given formats works. dEQP-VK, at the start of the name, is the root node, and there are subgroups to test the core API commands for copying and blitting images, and a subset of those are run for all formats with different image types, like copying from 1D to 2D using the general image layout for both images.
Starting with the first tip: if you dig deeper in the Vulkan CTS README, it mentions a few useful things. One of them is that, when building the project normally, you build many more things than the deqp-vk binary, like stuff for Vulkan Safety Critical, or some additional extra binaries. Of course, if you build calling cmake or ninja, you can indicate that you only want to build deqp-vk. You can always create a shell alias, script or whatever to do that, but if you don't want to forget about it, there's a cmake configure option that you can use when building the project to specify the targets you want to build by default, called SELECTED_BUILD_TARGETS. You pass a list of target names separated by spaces. If you specify deqp-vk only, only that binary will be built by default, which it's much faster than building everything.
Another configuration option that I find particularly useful is DEQP_LOG_NODE_SOURCE. This was added recently to main, so it's present in the most recent branches, and requires a modern compiler to work. If you enable it, the log file will mention, for every test case that is run, where the test case is being added to the tree. This makes it easier to locate the source code of a particular test. We'll talk a bit more about that later. If you use this, there's a small performance penalty when creating the test tree in memory, but it's probably worth it in many cases. I always use this option.
Another thing: the deqp-vk binary is typically around 400MB big in a Debug build, or more depending on some other build options, so it takes a long time to link it. This is specially problematic when you're making changes to a single file and rebuilding. Many Linux distributions still use BFD or Gold as the default linkers, which are very slow for deqp-vk, so my advice is to switch to LLD or Mold. For example, in my previous laptop, the default linker took 30 seconds to link deqp-vk. With mold, it went down to <2 seconds. Sometimes you can make them the default linker with update-alternatives but you can also set a link option from LDFLAGS and cmake will pick that up when configuring the project.
Moving on from build tips, I wanted to briefly explain how changes are reviewed and merged to the project. This is related to the community aspect I mentioned before. Every change that lands in the project needs to pass an internal review process inside Khronos. Who's reviewing and approving changes to CTS? Typically people who have a Vulkan implementation, that is, a Vulkan driver running on some hardware, but software-only implementations also count. Some reviewers work on Mesa. For example, Samuel here reviews a lot of changes to make sure they work on RADV. However, I think it's fair to say most reviewers do not really look at the source code for the changes. They only want to check if the tests pass on their driver and, if they don't pass, they block the change until they fix the driver, if it's a driver bug. A change cannot be blocked forever in those cases, so typically the result is that changes are merged more or less quickly and all drivers work and are fixed at the same time. To help with that process, Khronos has an internal Gitlab instance to track proposals for new tests and issue reports. If you don't have Khronos access, you can report issues and submit PRs on Github, and someone will eventually review them and move them to the internal tracker.
So, again, if you have a Khronos account, go to Gitlab. If you don't, use Github. If days or weeks pass and nobody replies on Github, feel free to mention us, specially if your issue comes from Mesa. No guarantees, of course, but it may help move your issue forward. You have some of our handles in the slide.
Some basic stuff about reporting issues. First, please mention hardware and driver, and if possible one specific case that is giving you trouble. Maybe mention a larger set of tests if you think all of them are affected by the same problem. Things that can be reported: Test mistakes, specially if the problem is reported by the validation layers. Also lack of coverage if you think something needs to be tested. A couple of tips about coverage: if you've implemented something new and all related tests seem to be passing, try to break something on purpose inside the driver and see if the tests still pass. You may be surprised. Also, if you discover a bug in your driver that is quite simple, obvious and it should be easy to reproduce, maybe it can be converted to a CTS test, so think about that.
Expanding on the topic of reporting issues, sometimes I get the feeling that people assume some minor pain points that they find are simply there and go "well, CTS is like that sometimes", so I want to mention more things that are fair game and worth reporting. Configuration step takes too long: at some point a version of the layers was being built by default and the dependencies of the layers were being downloaded during the configuration step, instead of a previous step that fetched external dependencies. Tests take too long to run: maybe a pathological case of the shader being too slow to compile (this is happening to our Rpi team), or the test hammering the PCI-e bus due to the chosen memory types (this happened to NVK and we introduced a mechanism to more easily select better memory types, and did some fixes to some tests) And more stuff: tests that take too long to skip, or fail without giving a hint about where, or explaining why (neither in the terminal nor in the logs), not logging enough information, build or link time worsens a lot, etc. You don't have to assume "CTS is like that".
More stuff: dealing with test failures. First, if the test logs an image to TestResults.qpa, you can view them with the cherry tool mentioned in the README, but you can also use a self-contained viewer in the scripts subdirectory. It runs entirely in your browser. I have it as a local bookmark in Firefox. It has some limitations. Images in the test log are logged as PNGs. Drawbacks: color depth, 3D images, etc. Ideally this tool would be much more sophisticated. Images would be logged in a format that was more flexible and the tool would allow you to examine layers of an image and raw values with precision and easily by hovering the mouse over a failing pixel, but we don't have that. You can also run the test in RenderDoc, and you should pass an option for that. Finally, if the failure message does not mention a line and file in the source code where it's failing, which is trivial, please report it as an issue. In the mean time, the log-node-source option that I mentioned before can help.
If you want to check what a test is doing, here are some hints about navigating the CTS source code. I don't want to talk too much about this, but some details are worth mentioning. If you log the node source and it points you to a line that mentions addFunctionCase, you only need to check the functions that are being run. Otherwise, the source code contains test cases and instances. Test cases are the leaf nodes of the test tree that we mentioned, and have methods to check the requirements to run the test, to generate the shaders that will be used, and another one to create a test instance. The test instance is the object that runs the actual test, in a method called iterate, and internally they can manage all the resources they need: images, buffers, command buffers, etc. Test Cases in the tree are kept alive since the start of the process and until it finishes while a Test Instance is short-lived: they're created to run the test, they allocate the resources they need, and they're destroyed after the test finishes, freeing all resources.
The hard part: finding tests. If I was given a cent every time someone asks me if we have tests that meet certain criteria, I could amass a small fortune. We can always grep the source code searching for extension names, feature names, etc, but things can get really complicated.
Sometimes it's not only about Feature X or Format Y, or a given API call. We want some specific memory types and alignments, or in combination with some other feature, or some other weird and specific requirements that may not be easy to meet.
Sometimes we look at the source code and find that we do have that coverage. Some other times we have similar coverage but not quite, and it may be easy or hard to add the requested coverage. Of course, sometimes we're lacking some general coverage and we need to create new tests.
Unfortunately, in many cases the real answer is that we do need to grep the source code as we mentioned before. However, some of this is being changed nowadays and we're making progress towards making test requirements more explicit and, in a sense, declarative. Ideally, at some point that should let us build a system that could be used to search test cases by features, formats, etc. You can imagine something similar to Sascha Willems' GPU info database. We could combine it with logging the node source and make it easy to find the source code of those tests.
That's it. Many thanks! Let me know if you have any questions, or maybe you just want to rant about CTS. If not now, maybe later in the hallway.
Q&A
Q: We sometimes see that a particular test case runs on a queue like the graphics queue. Is it possible to run that test on another queue like the compute queue?
A: That's being worked on. There's a very old feature request for CTS to basically be able to run every test that is possible on any other queue that's available on the device, in general. That's one thing. It's a hard problem given the state of the source code but people are working on it right now, and I think we will eventually get there, when we can run the tests on, say, the compute queue. Then that begs more questions like, for conformance purposes, do we have to run all the tests on all the queues supported by the device or not? The answer is not simple. The other possible answer is something that we have been doing for some time, as you know, is that we sometimes create specific variants for some tests that we know are tricky to run on the compute queue or on the transfer queue, and we create variants of those for those specific queues.
Q: Question about lists of packages for the linux distributions to be able to build deqp-vk.
A: I misunderstood part of the question as it was asked but later caught up with the person asking and he was genuinely wondering why deqp-vk was failing to build on some systems following the instructions. As a result of that question, we submitted an update to the documentation to be more precise in the list of required packages to build CTS, so that's a win!
Q: Request to parametrize the side of the framebuffer in some tests, which is important for tilers like Turnip, to be able to run the test both using system memory and graphics memory.
A: Ack the request and admit that's a hard one (because it would need changes in a lot of tests).
21 Sep 2026 8:17pm GMT
20 Sep 2026
planet.freedesktop.org
Danylo Piliaiev: Turnip’s evolution over the years, supporting flat-screen and VR games
It's been more than 5 years since I started working on Turnip, the open source Mesa 3D driver for Adreno GPUs, with Igalia's graphics team. Looking back, it's amazing how much we have achieved together, and how much the driver has improved since then. Here's my take on the evolution through those years, some of the challenges we faced and how we overcame them. Before I start, I want to thank all those great people from Igalia, Valve, and the Mesa community I've had the pleasure to work with.
This post is a bit one-sided as it contains my particular view of the evolution of Turnip through those years and it doesn't touch on a lot of important features implemented, issues debugged, and improvements made by others.
Challenges
There are several overarching challenges that the Turnip driver has had and still has to this day:
- There was no hardware documentation, we had to reverse engineer everything, with the most painful part being hardware bugs that necessitate workarounds. I cannot say I became as good as I would have wanted in all that, but I'm grateful that in Turnip we had people with absolutely amazing HW reverse-engineering skills!
- Hardware wasn't really built to run desktop games, at least at first. It became much better in A7XX generation, but there are still some bumpy parts. Running desktop games is specially important for hardware like the Steam Frame and others.
2021
Way back, when I started to contribute to the Turnip driver, we were fixing CTS tests, running "trivial", by today's standards, applications like "Genshin Impact", "TauCeti Vulkan Technology Benchmark", "3DMark". I have a few posts describing some issues debugged back then:
- Turnips in the wild (Part 1) - Fixing "Genshin Impact";
- Turnips in the wild (Part 2) - Fixing "Genshin Impact" and "TauCeti Benchmark".

If you'd asked me back then about running AAA PC games, I would have nervously laughed at best =)
DXVK Time!
But soon, we found a way to at least test how PC games render on Adreno GPUs with Turnip; we weren't able to run games on our development boards, but the driver, after a bit of massaging, had enough features to work with DXVK (Vulkan-based translation layer for Direct3D 8/9/10/11). So the solution at that time was to run the game on the PC, record Vulkan API calls with GFXReconstruct and a Vulkan profile that constrained the desktop GPU's capabilities to those of the Adreno GPU.
A bit later I was able to "play" simple DirectX games on the development board by playing the game on a PC and in real-time replaying translated Vulkan API calls on the board. You can read more on this in "Testing Vulkan drivers with games that cannot run on the target device".
That was a great start; we weren't able to run games directly on our development boards, but we were now able to test them and begin to implement more features necessary for DXVK.
Measuring Performance
In parallel with trying to support DXVK, fixing issues, and reverse engineering hardware features, we found ourselves needing to better understand performance. We chose to support Mesa's u_trace framework and integrate with Perfetto. At that time the performance measurement support in Mesa was rough, with only Freedreno (OpenGL driver for Adreno GPUs) supporting u_trace. Spoiler: as of now, most Mesa drivers either support or are in the process of merging support for u_trace/Perfetto integration.

2022
At the end of 2021 Turnip became Vulkan 1.1 conformant. We also started testing lots of single frame D3D11 captures of games we had, which uncovered plenty of new issues.
With more testing came more issues!
At that point I met our three two main adversaries over the years, in full force:
- Low-Resolution-Z (LRZ) depth optimization;
- GPU hangs, with the worst of them completely shutting down the SoC.
Low Resolution Z (LRZ)
Low-Resolution-Z is an extremely important optimization to get right in order to get reasonable performance out of our GPU (especially for VR games), but I believe it's also one of the most complicated (at least from a software POV) depth-related hardware optimizations across various GPUs.

Conceptually it's relatively simple - create a low resolution depth buffer during primitive binning pre-pass, throw out primitives that completely fail LRZ tests during binning. While during tiling, prevent a lot of overdraw by testing against this already formed low resolution depth buffer. But, in practice, there are lots of implicit restrictions on what can be done without disabling LRZ (to prevent correctness issues), and being too pessimistic when disabling LRZ can lead to unacceptable performance.
GPU Hangs
GPU hangs, on the other hand, are a plague that every graphics driver developer is intimately familiar with. However, at that time there were also unrecoverable hangs, which, due to certain issues in firmware or GPU HW itself, caused the entire SoC to shut down. They were an extreme pain to debug.
I tried many methods over time to debug them:
- At first I tried Vulkan-level breadcrumbs via "Graphics Flight Recorder", it was useful at that time but far from perfect and abandoned by Google;
- Later, I implemented driver-level breadcrumbs that allowed finding the source of hangs at a finer level, and more importantly, synchronously going through breadcrumbs to debug unrecoverable hangs;
- Prototyped editing a captured submission to the GPU.
That made debugging somewhat bearable.
Turnip supports Vulkan 1.3
Meanwhile Turnip gained Vulkan 1.3 support! It was necessary for DXVK and VKD3D-Proton.
Reviewing VK_EXT_fragment_density_map
The third Harbinger of the Apocalypse has arrived: VK_EXT_fragment_density_map, the first stepping stone of the extensions essential for VR, was implemented by Connor Abbott. Since then I have been on the hook reviewing increasingly complicated interactions between VK_EXT_fragment_density_map and the growing host of extensions.
The fragment density map (Foveated Rendering) explanation can be found in the following blog posts from Qualcomm and Meta:
- Eye Tracked Foveated Rendering
- Improving Foveated Rendering with the Fragment Density Map Offset Extension for Vulkan
2023
Turnip began to support Adreno 7XX GPUs; previously we only supported the single 6XX generation, albeit with several sub-generations. The proprietary driver supported Vulkan since the 4XX generation, but it was only Vulkan 1.0, and 4XX/5XX generations were not powerful enough and didn't have enough features to support anything with Vulkan.

New generation means more issues to fix!
To help with that I implemented:
rddecompilerwhich makes it possible to decompile captured raw submissions to the GPU into editable C code, coupled with the ability to replay them, print from shaders, and print from the command stream - resulted in a much faster debug loop;- A debug option that helps find where we use stale register values
TU_DEBUG_STALE_REGS_RANGE.
/* pkt4: GRAS_SC_SCREEN_SCISSOR[0].TL = { X = 0 | Y = 0 } */
pkt4(cs, REG_A6XX_GRAS_SC_SCREEN_SCISSOR_TL(0), (2), 0);
/* pkt4: GRAS_SC_SCREEN_SCISSOR[0].BR = { X = 32767 | Y = 32767 } */
pkt(cs, 2147450879);
/* pkt4: VFD_INDEX_OFFSET = 0 */
pkt4(cs, REG_A6XX_VFD_INDEX_OFFSET, (2), 0);
/* pkt4: VFD_INSTANCE_START_OFFSET = 0 */
pkt(cs, 0);
/* pkt4: SP_FS_OUTPUT[0].REG = { REGID = r0.x } */
pkt4(cs, REG_A6XX_SP_FS_OUTPUT_REG(0), (1), 0);
After a lot of command stream and shader reverse-engineering - at the end of the year the Adreno 7XX generation was in decent shape in Turnip.
2024
Work continued with Adreno 750 now being the main target. We had a lot of issues to debug and fix:
A Hat In Time
Farming Simulator
Inspecting Dark Souls 3And many, many more games.
Android (Waste)Lands
While Turnip, as far as I can remember, was officially only used on some Google Chromebooks, there apparently was a dedicated community of people who tried and actually ran desktop games on Android via Termux. The community has grown since then, but it was great to have a real user already using the driver to play games and having better outcomes than with the proprietary driver, which often doesn't get updated after a phone's release.
While fascinating, those Android setups were hard to debug, so we used them only a few times.
Turnip Is Vulkan 1.4 Conformant and Preemption Support
Not my achievements at all, but two major milestones for Turnip were:
- We caught up with the Vulkan releases and were day 1 conformant to Vulkan 1.4 on A7XX.
- Preemption support - this is a crucial feature for VR and needed by the Steam Frame. The VR compositor has to be able to preempt games that are being rendered at that moment, otherwise we can miss a frame or several, which feels extremely bad for a user moving their head in the VR environment. An Adreno GPU can be preempted at two boundaries: at the drawcall/dispatch boundary in direct (sysmem) rendering, and at the tile boundary in tiling (gmem) rendering.
2025
More VR Extensions
Even more extensions came in from Connor: VK_QCOM_multiview_per_view_viewports, VK_QCOM_multiview_per_view_render_areas, VK_VALVE_fragment_density_map_layered, VK_QCOM_fragment_density_map_offset, VK_QCOM_subpass_shader_resolve, VK_EXT_custom_resolve. Those are essential for pushing VR performance to its limit. The biggest issue with FDM related extensions is that they don't have good CTS tests; the complexity comes from the fact that a conforming driver implementation may simply choose not to reduce quality in regions specified by the density map. Even the size of those regions is not known to a game using those extensions!
As a result I wrote lots of tests that exercise various combinations of those extensions and require visual inspection for them to pass 🫠🫠🫠.
Result of one of FDM tests (Notice lower resolution at the edges of both eyes)Those tests also measured performance, which helped us fix disparities with the proprietary driver.
Half-Life: Alyx
What is going to use the above extensions to push Steam Frame to its limits? Of course, Half-Life: Alyx.
Even with previously mentioned tests in place, Half-Life: Alyx found plenty of both rendering and performance issues.
That's how I felt when a new issue was foundAre We Performant Yet?
We already had feedback from Android users running games that Turnip performance is sometimes better than the proprietary driver, but sometimes noticeably worse. But how to compare them? Turnip has Perfetto support, the proprietary driver has Snapdragon Profiler, but it's hard to use and even then - we'd just see that some particular renderpass is faster or not. There could be hundreds or thousands of draw calls in a renderpass!
Staring at the command stream from the proprietary driver stopped yielding any insights, so the next step was to take a game trace that can run both on Turnip and Qualcomm's driver and compare them draw by draw, measuring all performance counters along the way. It was done by lots of command stream patching, but the end result was an ability to compare key registers at every draw and every performance counter in existence.
Turnip VS Qualcomm's driver draw by draw comparisonThe resulting table was huge:
This helped us to close several gaps in performance, however the most useful part of the comparison was not the counters, but the register comparison. The counters, aside from execution time and LRZ stats, were surprisingly hard to convert into any useful insight.
Preventing Regressions
With the driver gaining capabilities, but not gaining many more users, we needed a way to prevent regressions while introducing more and more complex features. And while Vulkan Conformance Test Suite is being tirelessly improved upon year after year, it's still far from enough.
It was time to introduce a CI system that would be able to detect visual regressions in rendering and performance regressions. We already had a number of d3d11 and d3d9 captures to start working with, and so it came to be:
Above, you can see one of the results where the rendering regressed. In most cases we run only a single frame capture instead of longer multi-frame traces. This was an explicit choice due to the observation that it's better to have a wider selection of captures, than trying to cover any single game better (VR games are an exception here). In many cases only one or two captures out of hundreds regressed due to some issue; the regression was caused by some very specific pattern the game had, which wouldn't appear in others.
Every night we test several driver configurations:
- Fixed DXVK/VKD3D versions + forced direct (sysmem) rendering;
- Fixed DXVK/VKD3D versions + forced tiling (gmem) rendering;
- Upstream DXVK/VKD3D versions;
- Turnip compiled with
ubsan(undefined behaviour sanitizer).
We also run important MRs through that CI to find regressions early on.

Every nightly run generates a performance datapoint, so we can see the line going down day by day. Don't mind the bump where we had concurrent binning "optimization" enabled 🫠 (that's a sad tale of an incredibly complicated HW feature which failed to deliver performance gains).
At the moment we are testing more than 600 different game frames per driver configuration; the APIs span D3D8-D3D12, Vulkan, and OpenGL. A single configuration runs in under an hour on just two Adreno 750 devices.
With CI in place we became increasingly confident in making changes to the driver.
See more in my XDC 2025 talk:
2026
Steam Frame was announced at the end of 2025; now we still had plenty of things to polish.
Performance
We were now working much more on performance, and the main culprit in bad performance is often Low-Resolution-Z being fully or partially disabled, especially in VR. We improved Perfetto tracepoints a lot: added new ones, fixed tracepoints with complex renderpass suspend-resume setups, and added performance warnings:

Now it was much nicer to work with, and something external developers could reasonably use.
As a result, while working on the performance of Half-Life: Alyx and some other VR games, we improved LRZ support a lot.
We've also compared Turnip against Qualcomm's driver on a new GPU Performance Microbenchmark (gpu-ratemeter) to squash the rest of the performance differences.
There was also a lot of compiler work, done by great compiler engineers working on Turnip.
D3D12 Woes
As we have been testing more D3D12 games running through VKD3D-Proton, we started to find interesting issues. Those generally are: implicit D3D12 features/requirements, or some behaviour out of the D3D12 spec but which all desktop GPUs do, or in one case Turnip having a higher limit than desktop drivers. We saw at least:
- UE5 not working correctly with wave128 (only Adreno has such wide waves) up until a few months ago: vkd3d-proton PR #3265
- D3D12 not having a proper query/limit for number of elements in buffers, implicitly supporting more than Turnip advertises: mesa MR !41477
- One game relying on "fair" execution of dispatches when implementing its own spin locks in compute shaders: mesa MR !41562
- UE5 relying on higher memory allocation alignment than Turnip has: vkd3d-proton PR #3231
- Games not checking for
D3D12_FEATURE_DATA_D3D12_OPTIONS21::ExecuteIndirectTierbefore usingD3D12_EXECUTE_INDIRECT_TIER_1_1commands
Alyx ☆ Rare GPU Hangs
It's "good" when the GPU hangs at a predictable place, it's bad when the GPU hangs randomly, and it's even worse when the GPU stops hanging when you try to isolate the issue in any way to debug it.
One of such hangs happened in Half-Life: Alyx, when moving through a specific location the GPU hung, sometimes, and sometimes it didn't for a long while. I've tried:
- At first it seemed to happen only on the stable Turnip branch, so I've tried to bisect the issue;
- After a while I found out that it happened in any branch;
- I tried to stare at GPU coredumps - no luck;
- Tried to get a gfxreconstruct trace to reproduce the hang; it might hang once or twice out of many replays.
What is almost impossible to do in such cases is test whether disabling a certain driver feature helps; you always have doubts - it didn't hang this time because of the feature I disabled, or I'm just unlucky. This happened several times during the investigation, the hang would disappear for an hour, and then reproduce several times in a row.
Previously I wrote that we have a mechanism to capture and replay raw submissions that are sent to the GPU. Yes, I tried that too, we can capture the submission that hanged the GPU. Guess what? The captured submission executed absolutely normally and didn't hang, not on the first execution, nor on thousands….
At that point I still didn't have a single clue what's going wrong, aside from some kind of hardware errata being involved. It was time to improve the debug tooling even further. Before, the captured .rd submission could be replayed once with one replay invocation, but the issue at hand demanded lots of iterations, and more ergonomic editing of the command stream than we previously had. So I've made improvements (they still are work-in-progress) to:
- Loop specific submission any number of times;
- Quickly disable specific draw call ranges in the submission;
- Automatically bisect which draw/dispatch causes hang/fault.

Only with each bisection step doing 50000 iterations was I able to narrow things down to a few draw calls. Looking at them closely still yielded only more head scratches though. However, with things narrowed down that far, I had something to poke other, more knowledgeable people with.
After some back and forth, it appeared that I hadn't thoroughly checked all shader debug options we had, or rather, I checked the option, but due to the rarity of the hang I misidentified it as not helpful!
In the end it appeared that there were two hardware errata that needed to be implemented in our shader compiler. And we got "lucky" that they were revealed by one of the most important games to run on the headset.
Present
Driver work never ends, there are still games to debug, features to implement, and VR performance to improve. The fact that Steam Frame runs Linux, is based on open-source software, and isn't locked down means that Steam Frame would be used in many ways unforeseen by us. I hope that Steam Frame release would bring improvements to the VR ecosystem and spearhead PC gaming on Linux running on ARM platforms.
20 Sep 2026 10:00pm GMT
18 Sep 2026
planet.freedesktop.org
Mike Blumenkrantz: Out Of Jail
Q3: Big Updates
Hi.
Long time no see.
I've been in Big Triangle jail for the past several months, but now I'm out and "free" once again. I can feel the news sites trembling. I can hear the RSS readers dinging.
That's right. SGC is back.
Future Posts
I had a lot of plans going into 2026. I promised cool stuff. I promised a new level of insanity.
I'll give you a couple weeks to prepare yourselves, but it's coming, and you are not prepared.
Here's a preview of some topics I'll be covering before 2027:
- Zink technical updates
- Zink non-technical updates
- That time I fell into
auxiliary/tessellatorand barely survived Top-secret project(s) that Big Triangle doesn't want you to know about
The Teaser
I'm still easing back into blogging, so today will just be a short warmup post. Nothing major. We're basically done already.
One more small detail:
On both Android and native, it's zink all the way down.
18 Sep 2026 12:00am GMT
11 Sep 2026
planet.freedesktop.org
Erik Faye-Lund: Change of employment
At the end of July, I left Collabora after 8 great years there. This makes it my longest employment to date, which is… something. I'm super grateful for the opportunities Collabora has given me, and I'm leaving a great team filled with lots of talented and passionate people. It's been an honor.
Similarly, at the start of August I started working at Arm. The same company I left over 17 years ago.
Why I left
The main reason I left was that after many years of working at Collabora it had become more and more clear that the top leadership and I had some pretty fundamental disagreements on how the company should be run. I spent a lot of time and energy trying to nudge things in what I think would have been the right direction, but that didn't lead to anything but burnout on my part. In the end, it seemed best to just part ways.
But also, one thing that I always disliked while working at Collabora, was working as a contractor. It's just a lot of paperwork and makes a lot of things that should be simple much harder. This was mostly things between me and the Norwegian government, like tax filing, sick leave and parental issues. This wasn't the straw that broke the camel's back, but it certainly contributed.
I was also kinda done working from home. We had moved closer to downtown, and our new place didn't really have space for a dedicated office. For the last year, I worked out of a co-working space.
An easy choice
One of the major reasons why I left Arm back in 2009 was that I wanted to move back to Oslo, the city I am from. Trondheim, while a very nice city had started to feel a bit small. It's very much a student city, and that's fun when you're in your early twenties. But when you're getting close to 30, you've found that a lot of your non-work friends have finished their studies and moved elsewhere.
For the last 3 or 4 years at Collabora, my work mostly revolved around Panfrost, which was partially funded by Arm. I got to work with some old colleagues again, and it was fun.
So when Arm opened an Oslo office in 2023, that piqued my interest. Several of my friends (like Jake "ferris" Taylor, among others) ended up taking jobs at that office as well.
This checked a few boxes for me, as it would let me continue a lot of my work pretty much uninterrupted, would provide regular employment, and would let me connect with old and new friends. So when it became clear to me that Collabora wasn't the right place for me any more, I reached out to an old friend at Arm, and the rest is history.
So what's the job?
I'm working in the Mesa team at Arm, which means I pretty much work on the same things that I did before, just with a different employer.
It's been less than 1.5 months so far, and Arm has a pretty extensive onboarding process, so a lot of the details haven't materialized yet. But it's clear that I'll still be working on Panfrost and PanVK. I'll also still be serving on the X.Org BoD.
In the short term, I'm preparing for XDC. I have a lightning talk and a workshop to prepare. Plus a bunch of code to write to support the lightning talk, phew.
My longer term goal is to make sure Arm is as good as we can reasonably be at working with upstream. There are certainly some challenges in this area, but I hope that I can work with the stakeholders to make sure we get there.
Closing
So yeah, I'm back at Arm after 17 years away. I'm happy with my choice so far, and it's nice to be back working with a lot of new and familiar faces. I'm looking forward to finding out how I can be of most use, and…
See you at XDC?
11 Sep 2026 10:30am GMT
10 Sep 2026
planet.freedesktop.org
Matthias Klumpp: JPEG-XL as default in AppStream, and better media processing
Two weeks ago, I released AppStream 1.2.0. This release contains a lot of great changes, but one of the most important ones concerns how media are being handled, and AppStream's default image export format.
AppStream is a Freedesktop metadata standard to describe software components. That can be anything from system services over fonts to console and graphical applications. AppStream metadata is supposed to give users enough information to decide whether they want to install a piece of software, to represent that piece of software, and to give the operating system enough information to decide whether a software component should be installed automatically and (to some extent) what capabilities and relations it has, to provide the user with sensible options.
Especially for the first two goals, and especially for GUI applications, AppStream supports icons and screenshots, which are used to showcase applications. Today, AppStream is used by all kinds of services, from Linux distributions over firmware updates to Flatpak and desktops directly. AppStream's original design however comes from the perspective of Linux distributions in 2011, where you may want to browse the software catalog offline, without delay, and without pinging an external server (which could be a privacy concern).
Therefore, a common way to deploy an AppStream-enabled software repository is to ship all icons of all applications in the repository to the user as part of the repository metadata download. AppStream does support remote icon downloads nowadays, and for a while I thought that this would become the default eventually. However, especially in today's world, having a bandwidth-saving, instantly responsive, privacy-protecting application browsing experience seems more important that ever.
PNG images are great!
The only format that AppStream supports for icons and screenshots (which are downloaded on-demand from your distributor's CDN) has always been exclusively PNG. PNG images are perfect for icons, because they compress well (especially for common icon shapes), are fast and simple to load, and can be loaded anywhere, by any toolkit or webbrowser. They also ensure we deliver faithful screenshot images, even though we may have scaled or re-rendered them. Still though, PNG images are less great for screenshots, as they are not very efficient, which puts strain on any CDN that has to deliver them, as well as on people's internet connections when browsing screenshots. Having smaller thumbnails alleviates that problem a little, but does not fully solve it.
But even for icons, PNG could be improved upon: In many cases, icons are re-downloaded with the repository metadata again and again, so having a large icon tarball adds up to the data transferred during metadata refreshes. AppStream also now supports large 128x128px icons, which nobody in 2012 expected we would need, adding even more data that will be re-downloaded. Saving some space here translates directly to lower bandwidth costs as well as faster downloads for users.
To improve PNG file sizes, the AppStream Compose library, which handles all image processing and metadata catalog composition, was running optipng on all generated PNG images. That does create smaller PNG images, but they were still relatively large compared to other image formats.
For a long time though, there was no alternative to PNG images for icons: There was no lossless image compression format that could give us the same quality as PNG images and that was also widely supported.
JPEG-XL vs PNG in AppStream
Since 2021 we have JPEG-XL (JXL), which offers a true lossless mode with often better compression than PNG. The issue was that JPEG-XL wasn't widely supported. Then, in 2025, the PDF Association selected JPEG-XL as the preferred image format for HDR images in PDFs, and now we are finally getting browser support and more ubiquitous availability of the format (you can try it right now in Firefox!).
For screenshots, using JXL's lossy mode, it has obvious and extreme size advantages over PNG, so supporting JXL or WebP for screenshot images was an obvious choice. If JXL would support the lossless case very well as well though, we could serve many use cases with the same exported image format, which is very attractive to me.
So, the obvious next question was whether it was worth the pain of switching the icon format, so I did some measurements on real icons. For that I used the AppStream component icon pool that Debian Unstable ships, which is almost 5000 application icons of various sizes, and converted them to PNG:
| Icon size | Icons | PNG total | JXL total | Pool saved | PNG avg | JXL avg | Median saved | Mean saved | Worst | Best | Larger as JXL |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 48×48 | 1544 | 3.7 MiB | 3.0 MiB | 17.8% | 2.4 KiB | 2.0 KiB | 17.9% | 16.7% | -118.7% | 60.0% | 206 |
| 64×64 | 2018 | 7.0 MiB | 5.8 MiB | 17.8% | 3.6 KiB | 2.9 KiB | 18.0% | 15.8% | -112.7% | 70.0% | 279 |
| 128×128 | 1411 | 11.2 MiB | 8.7 MiB | 22.0% | 8.1 KiB | 6.3 KiB | 20.1% | 17.5% | -89.7% | 61.0% | 209 |
| TOTAL | 4973 | 21.9 MiB | 17.5 MiB | 19.9% | 4.5 KiB | 3.6 KiB | 18.6% | 16.6% | -118.7% | 70.0% | 694 |
PNG images saved with libpng at effort=4, compression=9, then optimized using optipng -o2, JXL images encoded using vips jxlsave lossless=1 effort=7 strip=1 via VIPS/libjxl.
As the table shows, using lossless JXL images over size-optimized PNG images (using optipng's default settings) provides a roughly 20% gain. This does not look like much, until you consider how often these files are downloaded: A 20% file size reduction may only save 1-2 MiB of disk space, but if they are downloaded over and over again by many clients, it will save a lot of bandwidth.
Interesting JXL encoding findings
As a sidequest, I was curious why some images were larger than their PNG counterparts when encoded with JXL, and what the ones that were significantly smaller were.
In short, the biggest size reductions for JXL existed on images that were already small as PNG, and contained large, flat color surfaces with hard edges and simple shapes. They were not very interesting, and much of JXL's wins come from accumulating smaller gains across all files, which compound the bigger icons get (especially at 128x128px, where JXL truly shines).
The events were JXL loses to PNG are more interesting: For example, it does quite poorly with pixel-art images that have a lot of repeating patterns. Those are encoded well by PNG, but less efficiently by JXL. Take for example Vonsh:
Icon of Vonsh, an SDL-based snake game, which PNG compresses better than JXLMy guess is that while PNG can exploit the repeating pixel patterns for compression, JXL's predicts surrounding pixels from its neighbours, which fails too often and makes it pay almost full entropy per pixel. In this single rare case, the PNG is at 5.4 KiB, while the JXL is almost 8 KiB in size.
Other cases I looked at were arguably buggy input data, where color channels were hidden under the alpha channel of the input image. PNG could probably again exploit repeats, while we were forcing JXL to encode pixels that were invisible in the final image. This is arguably a problem with the original input data. Currently, AppStream does not make any changes to icons at all, but in future we might add a filter that removes invisible colors from images to solve this pathological case (it was only two icons out of 5000 though, so it is not a high priority).
The third case I found where JXL loses to PNG were icons with checkerboard-like patterns:
Icon of x3270, an IBM 3270 Terminal EmulatorFor those, PNG can likely again exploit the repeating patterns, while a checkerboard layout is pretty bad for left/top predictors like JXL's. However, in this case the size difference (and loss for JXL) is only 450 bytes, so even though JXL loses to PNG, it does so not by much.
JXL in AppStream
Given these findings, JPEG-XL is the default image format starting with AppStream 1.2.0. AppStream Compose will encode all images losslessly as JXL, while screenshots are encoded in lossy mode at Q=90 effort=7. Since the optipng step does not happen for JXL images, this comes at no speed penalty and is even a bit faster on modern x86_64 CPUs (where libjxl can use SIMD). PNG is still available, and Compose can be told to switch between the two formats.
Upsides of JXL in AppStream right now
If you use JXL in Compose or the recent release of appstream-generator, you will get much smaller images and, for screenshots, will benefit from other JPEG-XL features such as progressive decoding, providing a far nicer user experience. libAppStream has supported JXL icons since version 1.1.3, so your clients will need that version or a newer one, and all software centers will have to support loading JXL images (which all of them do, provided the right plugins are installed).
Downsides of switching to JXL too quickly
JXL is a very new format, so web browsers might not yet display it if you are serving webpages. Your clients may also have bugs in processing JXL images, as the format is still "new". For example, switching on JXL in Debian sent KDE Discover into an infinite loop on startup while trying to load the icons (an issue which has been fixed, but clients will need that patch first before JXL is switched on).
This currently makes JXL enablement only possible when you know that your clients can support it. This is the case for me in Debian Unstable and Debian 14, which are using JXL images for a few weeks now, but not for any older releases. Platforms like Flatpak have it even harder, because they do know even less about their clients. So, even though it has big advantages, you may want to hold off on using JXL right away, and force PNG by setting the ImageFormat key to png in appstream-generator's configuration, or passing --image-format=png to appstreamcli compose.
It is also worth mentioning that JPEG-XL is much, much slower on systems that do not have SIMD instructions or for which the libjxl/jxl-rs library does not have them (such as apparently riscv64 right now). If this is a concern, you might not want to switch to JXL right away.
Media pipeline improvements
Besides the JXL default change, AppStream 1.2.0 also comes with a complete overhaul of its media processing pipeline. While libappstream, AppStream's main library, does not do any media processing and comes with very minimal dependencies to be embedded in client applications and used on servers, the same can not be said about libappstream-compose, AppStream's library to build metadata generating applications (the server-side part, usually).
The compose library has to render fonts into font specimen cards, inspect translation files, render SVG images, decode all kinds of raster images, inspect video files, etc. Especially the fonts, and the fact that fonts can appear in SVG images, has caused issues in the past, as libappstream-compose is a heavily threaded library and most font libraries can only work from a single thread. This forced the library to essentially go into single-thread mode anytime anything that could touch a font was being processed.
AppStream also originally was created for a "safe world" where applications were vetted by the distributors before their metadata was processed. This is increasingly not the case, so it made sense to put at least a few guardrails on the most complex part of the pipeline: The media processing. As part of the change, media processing was split out into a separate worker process. This solved two problems at once: Font handling was isolated in a single-threaded binary - if we wanted to handle fonts in parallel, we could simply spawn more workers. And, being in a separate process, the media processing could now be sandboxed.
As part of the multiprocess changes, Compose also switched from using GdkPixbuf to VIPS for image processing. The latter allows for much more fine-grained control over the image output and encoding, and comes with a lot of well-maintained filters and operations, which made it possible to eliminate a fair chunk of AppStream's hand-rolled image processing operations. As part of this transition, we unfortunately lost the ability to read XPM images, which dropped about 20-30 applications from the pool at Debian. But in the name of security, this is a sensible choice, especially since most XPM icons were very small and low-resolution, and applications using them could benefit from adding a high-quality PNG icon anyway. With VIPS, we also now restrict the amount of image formats we can load to a sensible set, so extremely niche or unexpected formats will be outright rejected (this includes sane-but-unusual formats for screenshots and icons, such as TIFF images).
The Compose library, with all of these changes, will now just request high-level operations (e.g. "render a font card for this font to a JXL image") from the worker, and provide it with input data in sealed memfds and output locations as FDs as well. On Linux systems, the worker will use Landlock if available, to block all write access to the filesystem, deny device access and deny TCP and UDP as well. The sandbox can certainly be tightened a fair bit in future, but this was a good and safe start to gain some experience with it without having things break too easily, given the many places Compose is used in (also, Landlock's API is surprisingly nice to use, so it was easier than I thought to add in this early version).
With all of these changes, the libappstream-compose library is now also officially marked API-stable, so you should be able to rely on it in future to build new things (its API has barely changed in the past, and now with the new media API and defaults change in place, it was time to declare it stable).
I want to see / try this!
Currently, the easiest way to have a look at the new data is to check out Debian Unstable. If you have a JXL-enabled browser, you can also see the icons in AppStream Generator's HTML pages for Debian Sid. If you are using appstream-generator for your distribution, you will also get much more pleasant statistics and HTML pages, as well as fully deterministic media output and a whole bunch of security updates, so, update to its recent 1.0 release.
Please keep in mind that if you switch to JXL, the client tools receiving the image data have to support it. Support varies depending on the Linux distribution, so, test it first and switch the default back to PNG in case you encounter any issues.
What's next?
With so many features and changes landed, the next changes in AppStream will focus on improving what already exists and fixing any issues (there will be more blogposts about the other features 1.2.x delivers!). Testing with the entire Debian archive as data source makes me fairly confident though that there will not be many problems. In the longer term, tightening the media processing sandbox will also be something we might want to do, e.g. by hiding parts of the filesystem tree or filtering syscalls.
For JPEG-XL, one obvious question is "Will you add support for it to the Freedesktop icon-theme specification as supported format alongside PNG, SVG(Z), and XPM?". For on-disk icon repositories, JXL's space-savings are less compelling, and it being HDR-capable is also not necessarily a killer feature (PNG can go a long way!). However, JPEG-XL's ability to immediately decode larger images at reduced resolution without resampling could legitimately be very powerful here, as applications could ship a single large image and quickly decode it at 1/2, 1/4 or 1/8 the size for different purposes in their UI. JPEG-XL also supports spot-color extra channels, which applications could use as masks to recolor raster icons at render time. This could be incredibly nice to color symbolic icons on-the-fly without any SVG and CSS. JXL also provides richer metadata, which might be neat for (license/author) documentation. So, the answer here is: Maybe it makes sense to allow another format, but this will have to be discussed first, as it would force JXL into every toolkit and desktop, which is a much bigger ask than supporting it only in AppStream.
As always, let me know what you think and please report any issues or bugs directly against AppStream or AppStream Generator if you encounter problems that are with the tools, and not with a project's metadata.
10 Sep 2026 5:48pm GMT
09 Sep 2026
planet.freedesktop.org
Dave Airlie (blogspot): nouveau on nvidia spark GB10 - it's alive!
After much back and forth and hoops jumping, I can finally reveal nouveau/nvk running on a NVIDIA Spark box.
This is running on a version of nouveau[1] that has
a) ported to the 610 NVIDIA firmware
b) a bunch of display rework from Moham
c) a bunch of 0 VRAM and L2 cache handling fixes
d) spark boot support
e) spark display support
and NVK[2] with patches to handle gb10 depth/stencil differences and 0 VRAM support.
I'm not sure how best to upstream it all, it's 100 patches and a new firmware which might mean it's a wait for nova type situation, but I just wanted to see it work.
[1] https://gitlab.freedesktop.org/nouvelles/kernel/-/commits/nouveau-610-wip-spark
[2] https://gitlab.freedesktop.org/airlied/mesa/-/commits/nvk-spark-wip
09 Sep 2026 7:56pm GMT
27 Aug 2026
planet.freedesktop.org
Sebastian Wick: Announcing Sovereign Tech Agency Investment in Flatpak
Together with Modal, I'm happy to announce that the Sovereign Tech Agency is investing nearly €510k into Flatpak development. The focus is on closing gaps in Flatpak's sandboxing story: new portals for audio, networking, VPNs, and spell checking, plus infrastructure work on entitlements and intents.
I'll be leading the technical side alongside Adrian, with organizational support from Kateryna and Cade. We've brought on a great team and the project will ramp up over the coming months through the end of 2027.
Read the full announcement on the Modal blog.
27 Aug 2026 10:30pm GMT
Timur Kristóf: The GCN outcast - Radeon HD 7870 XT
AMD is known to be Linux friendly and has had an open source driver stack for their GPUs for more than a decade. However, there was one GPU that has always been an outcast on Linux because it has never worked: the Radeon HD 7870 XT. This post is about how I fixed that so that Linux gamers can enjoy this GPU being fully functional now.
The Radeon HD 7870 XT is built on the GCN 1 architecture (also known as Southern Islands or SI, or GFX6). It has a chip called "Tahiti LE" which is a variant of the Tahiti chip that was quite high-end by 2012 standards. Tahiti and other GCN 1 chips have been supported on Linux for more than a decade, first by the radeon kernel driver and more recently by the amdgpu kernel driver.
We have every reason to assume that the 7870 XT should "just work", so why doesn't it?
Motivation
Why care whether a 15 years old GPU works today or not?
- If the driver stack claims to support GCN 1, we should indeed support all GCN 1 chips without exceptions.
- If there is a serious bug that prevents a GPU from working, there is a chance the bug also affects other GPUs. It is worth an investigation.
- Most importantly, it's a good challenge to see if I can figure out a problem like this.
Story time: what is a harvested (cut down) chip?
If you look at any GPU manufacturer, they have a lot of different products every generation, and those products use different variants of the same few different chips. How is that possible?
Due to economies of scale, there is a common practice among chip makers: they prefer to manufacture massive quantities of the same few chip instead of a small quantity of many different kind of chips. However, in order to have many different products, they can configure the same chip in different ways and sell those variants under different product names. Today we are focusing on AMD's old Tahiti chip, so let's use that as an example. How many different products did they launch that use the Tahiti chip and its different variants or refreshes? We can use Wikipedia to check that:
- Radeon HD 7870 XT, 7950, 7970, 7990, 8950, 8970, 8990
- Radeon R9 280, 280X
- FirePro W8000, W9000, D500, D700, S9000, S9050, S10000
All of those use the Tahiti chip. But what are the differences between those products?
- Different target audience, eg. workstation vs. consumer
- Different memory configuration (memory size and bus width)
- Different shader and memory clock speed
- Different amount of compute units, render output units (aka. render backend or RB), etc.
How is this achieved?
At the factory, each chip is examined automatically. Due to variance in the chip manufacturing process, not all chips end up the same, even if we do everything to make them the same. In practice that means there may be defects on the chip, or maybe not all units perform up to spec. This is the so-called "silicon lottery". The chips are then sorted according to how well they ended up performing and that's when the manufacturer decides what to do with them and what product they can be sold as.
There really are no bad GPUs, just incorrectly priced GPUs. Those that are still usable but ended up less than ideal will still be put to use: some parts (eg. compute units) are fused off and disabled, and the GPU is overall still functional as a weaker, cheaper GPU. That's how we end up with products like the Radeon HD 7870 XT.
Starting point for investigating the problem
Initial testing
To start with this work, open source enthusiast Leonardo Frassetto helped me to buy a used 7870 XT in good condition from an Italian used hardware site. After plugging in this GPU and booting my system, I noticed the following:
- Firmware (BIOS) can recognize the GPU, show a logo and boot grub
- When amdgpu loads, the picture disappears and becomes a colorful mess
- Looking at the logs, it seems the GFX block immediately hangs and goes into a GPU reset loop, as the kernel attempts to reset the GPU to fix the hang, which is expected
Information from Wikipedia
From Wikipedia, we got a list of AMD GPUs and an article about GCN to give us some basic info.
We can see that Tahiti LE is a harvested (cut down) version of the full Tahiti chip as the 7870 XT has lower specs than the 7970 (or R9 280X). There are plenty of GPUs with disabled CUs supported already, so I went with the assumption that the disabled CUs aren't the issue.
| Tahiti (top spec) | Tahiti LE (cut down) |
|---|---|
| 32 CU (compute units) | 24 CU |
| 32 RB (render output units) | 32 RB |
| 384-bit memory bus | 256-bit memory bus |
Information from Freedesktop Bugzilla
Someone opened a bug report on the old Freedesktop Bugzilla in 2013 complaining that the 7870 XT didn't work. Although the issue has never been solved, we can still glean some interesting information from that bug report:
- The display (ie. modesetting) should work, and the issue is "only" with 3D acceleration.
- One commenter was able to get the 7870 XT to work with basic compute shaders but not much else (definitely not a full desktop).
- The register dumps attached to the bug only contain the DCE (display engine) registers, so are not really conclusive.
- There were suggestions to change the
CGTS_TCC_DISABLEregister in the kernel driver, which didn't help. - The developers already corrected the register programming for harvested (cut down) RBs, and the 7870 XT doesn't have those anyway, so that isn't the issue.
Booting a working system
How do we even begin to diagnose what the problem is if the GPU hangs immediately at boot?
Booting in runlevel 3
I started by booting to runlevel 3 (basically just a terminal and nothing else). In this mode, the amdgpu kernel driver can initialize all blocks in the GPU correctly and the display works. We are in a terminal only environment so there is nothing submitting jobs to the GPU.
Booting a desktop with software rendering
I configured the system to use software rendering for both OpenGL and Vulkan by setting the following environment variables temporarily:
# Ask OpenGL loader for software rendering
LIBGL_ALWAYS_SOFTWARE=1
# Force using lavapipe as a Vulkan driver
VK_ICD_FILENAMES=/usr/share/vulkan/icd.d/lvp_icd.x86_64.json
# Force using swrast as a Gallium driver (for OpenGL)
MESA_LOADER_DRIVER_OVERRIDE=swrast
With that, the system can indeed boot with the 7870 XT, although the graphical user interface is slow because we are not using any hardware acceleration (obviously). I would not recommend anyone to regularly use their system this way.
Testing simple shaders
Even though all the system uses software rendering, we can still set different environment variables for specific apps and have just one test application run with "real" drivers for the hardware. I decided to use the vkrunner suite and wrote a very simple test case with a very simple compute shader. The test case allocates two buffers (SSBO). The shader has only one invocation that reads a small piece of data from the values_in buffer and writes it to the values_out buffer.
[compute shader]
#version 450
layout(local_size_x = 1, local_size_y = 1, local_size_z = 1) in;
layout(binding = 1) buffer block_out { uint values_out[]; };
layout(binding = 2) buffer block_in { uint values_in[]; };
void main()
{
values_out[gl_LocalInvocationIndex] = gl_WorkGroupID.x;
}
[test]
ssbo 1 1048576
ssbo 2 1048576
compute 1 1 1
That indeed works correctly, confirming what we found in the Freedesktop bugzilla, that simple compute shaders indeed work. What happens if we slightly complicate it though? The first thing we can do is to increase the number of invocations so that more than one invocation reads and writes memory.
layout(local_size_x = 96, local_size_y = 1, local_size_z = 1) in;
Whoops! This test case now hangs the GPU. I guess the way we changed the memory access must have caused the GPU to hang. With some trial and error, we can find that local_size_x = 64 still works but local_size_x = 65 and higher hangs. After some further trial and error I noticed that I can still use a larger workgroup if I skip memory accesses to certain address ranges. This definitely confirms that the issue is with memory access somehow.
Let's look at registers with umr
One of the developers who commented on the old bugzilla thread were mentioning the CGTS_TCC_DISABLE register. Although the suggested patch didn't help, I thought okay, let's see what is actually the value of this register using umr:
$ sudo umr -O bits -r tahiti.gfx600.mmCGTS_TCC_DISABLE
gfx600.mmCGTS_TCC_DISABLE => 0x09240001
.TCC_DISABLE[16:31] == 2340 (0x00000924)
Each bit represents a disabled TCC. We can see that some TCC units are disabled by default: 2, 5, 8, 11. Why are they disabled and what does that mean exactly? We saw from Wikipedia that the 7870 XT only has a 256-bit memory bus, although other Tahiti based GPUs have a 384-bit bus. This means that only two-thirds of the memory channels on this GPU are enabled and the rest are disabled.
I started to think, is it possible the GPU is trying to read from those disabled memory channels or use some disabled TCC units, or something like that? Technically, the disabled parts should be fused off and not accessible, but maybe it still tries to interact with them?
I went to take a look at the registers and register fields of GFX 6 to see if there is anything that stands out. I basically searched for "TCC" and looked at the results hoping for a clue. I found two interesting things:
- The
TCP_ADDR_CONFIGregister has a field calledNUM_TCC_BANKS - There are
TCP_CHAN_STEER_LOandTCP_CHAN_STEER_HIregisters which have fields for channels
Let's see the value of these registers:
$ sudo umr -O bits -r tahiti.gfx600.mmTCP_ADDR_CONFIG
gfx600.mmTCP_ADDR_CONFIG => 0x000002fb
.COLHI_WIDTH[6:8] == 3 (0x00000003)
.NUM_BANKS[4:5] == 3 (0x00000003)
.NUM_TCC_BANKS[0:3] == 11 (0x0000000b)
.RB_SPLIT_COLHI[9:9] == 1 (0x00000001)
$ sudo umr -O bits -r tahiti.gfx600.mmTCP_CHAN_STEER_LO
gfx600.mmTCP_CHAN_STEER_LO => 0xa9210876
.CHAN0[0:3] == 6 (0x00000006)
.CHAN1[4:7] == 7 (0x00000007)
.CHAN2[8:11] == 8 (0x00000008)
.CHAN3[12:15] == 0 (0x00000000)
.CHAN4[16:19] == 1 (0x00000001)
.CHAN5[20:23] == 2 (0x00000002)
.CHAN6[24:27] == 9 (0x00000009)
.CHAN7[28:31] == 10 (0x0000000a)
$ sudo umr -O bits -r tahiti.gfx600.mmTCP_CHAN_STEER_HI
gfx600.mmTCP_CHAN_STEER_HI => 0x0000543b
.CHAN8[0:3] == 11 (0x0000000b)
.CHAN9[4:7] == 3 (0x00000003)
.CHANA[8:11] == 4 (0x00000004)
.CHANB[12:15] == 5 (0x00000005)
.CHANC[16:19] == 0 (0x00000000)
.CHAND[20:23] == 0 (0x00000000)
.CHANE[24:27] == 0 (0x00000000)
.CHANF[28:31] == 0 (0x00000000)
Whoah!!! It seems like the registers indeed refer to the disabled memory channels:
TCP_ADDR_CONFIG.NUM_TCC_BANKSindicates 12 channels (value of 11) even though the GPU only has 8 active memory channels.TCP_CHAN_STEER_LO/HIrefer to channels 2, 5, 8, 11 even though those memory channels are disabled.
Let's try our luck and see what happens if we remove those. We simply need to edit the bit pattern for these registers, I typically use Gnome Calculator in programming mode to toggle the bits by hand.
# Change NUM_TCC_BANKS to 7
$ sudo umr -w tahiti.gfx600.mmTCP_ADDR_CONFIG 0x000002f7
# Use channels 0, 1, 3, 4, 6, 7, 9, 10
$ sudo umr -w tahiti.gfx600.mmTCP_CHAN_STEER_LO 0x43a91076
# We don't have any more channels so clear this to zeroes
$ sudo umr -w tahiti.gfx600.mmTCP_CHAN_STEER_HI 0x00000000
And… Yesss!!! with that, the test doesn't hang anymore, and the GPU starts working well.
What's actually happening on this GPU?
I had a chat with the lead developer of amdgpu Alex Deucher and shared my findings. He helped me understand what's what:
- TCC (texture cache per channel) is the L2 cache that is attached to each memory channel. For disabled memory channels, the corresponding TCC should be also disabled and should not be used (obviously).
- TCP (texture cache per pipe) is the L1 cache that is part of each CU.
- The
TCP_ADDR_CONFIGandTCP_CHAN_STEERregisters can tell the TCPs which TCCs they are allowed to use. These registers were programmed from the "golden" registers without regard to the fact that some channels may be disabled.
We can now see that the GPU was trying to use the L2 cache that is attached to the disabled memory channels, which (obviously) didn't work and caused the GPU to hang. With that understanding, I wrote a patch. With that patch, the system can boot normally (without a need to force software rendering) and the 7870 XT works fine now.
The patch has been backported to stable kernels too. If you are using a 7870 XT, you can now enjoy it fully working on Linux!
Perf testing
I also did some perf testing to see if it makes any difference how it is configured. I used Rise of the Tomb Raider at 1080p with lowest settings for these tests. This shows that the values mentioned above are indeed optimal for this GPU, so we only need to skip the disabled TCCs but otherwise can keep the order they are in.
TCP_CHAN_STEER_HI |
TCP_CHAN_STEER_LO |
NUM_TCC_BANKS |
frame rate |
|---|---|---|---|
| 0 | 0x43a91076 | 8 | 78 fps |
| 0 | 0 | 12 | 27 fps |
| 0xa147 | 0x43a91076 | 12 | 69.96 fps |
| 0x4310 | 0x43a91076 | 12 | 61 fps |
| 0 | 0xa9764310 | 8 | 66 fps |
What have we learned from this?
It turns out this is one of those rare bugs which only affect a specific variant of a chip and nothing else. The issue with harvested TCCs is a problem unique to Tahiti LE, which was only ever used in the 7870 XT (and the FirePro D500 according to Wikipedia), so the fix doesn't affect other GPUs. There have been no more GCN GPUs with this kind of harvesting (although RDNA1 has variants with harvested TCC as well).
On the bright side, it wasn't a useless exercise for me.
- I learned a lot about how the cache hierarchy works
- It was interesting to see how the cache configuration affects performance
- I feel proud that I was able to solve the mystery
27 Aug 2026 12:00pm GMT
26 Aug 2026
planet.freedesktop.org
Timur Kristóf: How does GPU recovery work with AMD GPUs on Linux?
An inconvenient part of GPU driver development is the fact that GPUs can crash and hang for various reasons, just like CPUs. Really, this can happen on any operating system to GPUs from any manufacturer. We've seen many such issues on the open source Linux driver stack for AMD GPUs. We've been tackling many bugs that lead to GPU hangs. We need to deal with the fact that hangs can happen and when they happen, we need to make sure that the system can recover and the user can keep using their Linux desktop. The purpose of this blog post is to give an overview of what the problem is and what steps we are taking to improve it.
Why does a GPU hang or freeze in the first place?
Many users are referring to this topic as "the bug", as if it was just one problem. If only we lazy driver devs fixed just that one problem, life would be better, wouldn't it?
In reality however, there are many different reasons why a GPU can hang, in fact there are so many possible reasons for hangs that it would be impossible to list all of them. We have fixed so many of these kinds of bugs over the years that it's really impossible to even give a list of those bug fixes.
I'll try to come up with some examples:
- Issues with shader instructions, eg. incorrect opcodes, unmitigated hazards, infinite loops, etc.
- On some GPUs, accessing unmapped memory (ie. a page fault) will also lead to a GPU hang
- Invalid commands submitted to the GPU
- Deadlock, ie. waiting for something that never happens (sometimes also caused by coherency issues)
- Missed interrupt which leads to a deadlock
- And many others…
Setting expectations
When talking about this problem space, it's often difficult to judge what we are talking about, what is possible and what isn't, so first let's start with setting some expectations about what we can do.
Firstly, we need to acknowledge that every GPU has different IP blocks (parts such as graphics, compute, DMA engines, display, video decoder, etc.) and each IP block may offer a very different uAPI and programming model. Therefore, each of them can fail in different ways and need to be recovered in different ways, too. It is impossible to handle every conceiveable issue with every IP block the same way. For example:
- In order to use graphics and compute, we have a command submission uAPI which userspace applications can use to submit jobs. Under the hood, user-mode drivers (UMD) use this uAPI to execute Vulkan, OpenGL commands.
- The display engine offers the modesetting uAPI, which is largely shared between all GPU vendors and works on very different principles.
- Some IP blocks like MC (memory controller) and others are not directly exposed to userspace, but are essential for correct functionality. When something goes wrong with one of these, that requires separate consideration.
This blog post focuses on GPU hangs caused by jobs submitted by userspace applications.
The other topic are out of scope for this post.
So, what can we reasonably expect about job submissions?
- When an app (or game) submits commands that crash or hang the GPU, the graphics context of that app is considered "guilty". We need to accept that it won't be able to continue and will crash (unless it was specifically designed to handle GPU failure). In some cases (such as VM faults, known hazards etc.) we can make an effort to try to stop it from crashing, but not always.
- We should do our best to make sure we don't freeze or crash other "non-guilty" applications or the whole desktop. However, sadly, trade-offs need to be made for the sake of performance.
Why is it difficult to deal with GPU hangs?
Let's start with a quick recap of how graphics drivers work. The way the graphics stack handles GPU jobs is the following:
- A userspace driver generates commands for the GPU (and compiles shaders, etc.)
- A job with those commands is submitted to the kernel driver through a uAPI (userspace API)
- The kernel resolves job dependencies, BO (buffer object) state etc. and writes the job to a hardware ring buffer (which is shared accross many processes)
- The GPU executes the commands from the job and when they are completed, signals a fence
The above model has barely any room for detecting and handling errors. The only thing we can detect is that a GPU job didn't complete within a timeout (eg. 2 seconds). We don't even know if that's because the commands are taking too long (and would complete in a longer timeout) or the GPU is stuck somewhere and isn't making any progress. Furthermore, job execution may overlap, so when it hangs, we can't always know for sure which job was responsible.
Actually, it's even worse than that. Modern GPUs have multiple job queues that can execute in parallel and all share the same resources (eg. compute units for running shaders). That means it is not even possible to be sure which queue is really responsible for a timeout: If two different jobs are executing on two queues in parallel, it can happen that the "guilty" job hogs all compute units, causing the other queue to "starve" and time out.
So, where does that leave us?
- We can detect when a job times out
- We don't know which job is really responsible
- We don't know which queue is really responsible
- We do know which jobs were in flight when a timeout happened
- We can't know why the timeout happened
How do we deal with GPU hangs, anyway?
Thanks to the excellent work of Alex Deucher and his team at AMD on the amdgpu kernel driver, various strategies were developed over the years to recover GPUs from a hung state and to make the problems less severe.
Enforcing isolation
Enforcing isolation means that we try to reduce how much an application using the GPU can affect another, effectively disabling parallel execution from different contexts. This greatly limits the scope of what jobs are potentially affected when there is a crash or hang, so makes it less likely that a "guilty" context can crash others.
- Currently,
amdgpudoes not allow jobs from different contexts to overlap on the same queue. This means when only one queue is active, we can 100% identify the "guilty" context correctly. - There is a kernel parameter
amdgpu.enforce_isolationwhich will additionally isolate different contexts that are running on different queues. This is currently disabled by default for performance reasons, you can enable it for better stability.
ASIC reset
ASIC reset is the simplest, and most destructive type of reset:
- All pending or in-flight jobs are killed
- The GPU is completely reinitialized
- On dedicated GPUs, VRAM is erased
That means that all processes that used the GPU will lose all resources they may have had in VRAM (except on APUs). The user sees that the screen turns black briefly, then every application and the desktop just crash (unless the compositor and apps were robust). This reset strategy means that just one misbehaving application can cause all other applications and the entire desktop to crash. It is better than watching a completely frozen screen, but not by much.
Soft recovery
Soft recovery was the first attempt at making resets less destructive. What it does is it kills all currently running shaders on a specific queue and hopes that the GPU can then move on and complete the job.
This can solve quite a few issues such as infinite loops in shaders, but it is somewhat dangerous because the kernel doesn't really have any knowledge if it worked, and can only judge by seeing whether the job now completes within a certain time or not. There is no way to know if the current job or subsequent jobs will actually work. So, soft recovery should be avoided when better recovery methods are available.
Queue reset
Queue reset (as its name suggests) attempts to reset just one specific queue without affecting others. There are mainly two ways this can be implemented for different hardware blocks:
- Graphics and compute queues on newer GPUs have a firmware-assisted queue reset, which means it's a feature of the CP (command processor) firmware. It basically terminates all operations (including shaders) that are executing or pending on the specific queue, and then moves on.
- SDMA, VCN and other queues can be reset by simply resetting the whole hardware IP block. These blocks usually only have one single queue, so the reset doesn't perturb anything else.
In my opinion, queue reset is the best way to handle GPU recovery. The only issue with queue reset was that in itself, it would still kill all currently executing or pending jobs from the given queue, even those jobs that haven't started yet. From a user perspective, that means it could still crash other apps and the desktop (unless you enabled enforcing isolation too).
Starting from Linux 6.18, an important improvement was made to queue resets: the kernel can now re-emit pending jobs from other contexts after the queue reset is complete, so in practice it is very likely that only the "guilty" misbehaved app is killed and everything else can continue. This is not 100% guaranteed though, because it is still possible that a job from a different queue can starve other jobs. For the safest user experience, you should turn on enforcing isolation too.
IP block soft reset
I introduced IP block soft reset as a GPU recovery method very recently. This method is more blunt than the queue reset, because IP block soft reset will reset an entire IP block including every queue it has. For the graphics/compute block this means it will practically reset all graphics and compute queues at the same time. The reason why I added this is because it works on GPUs that don't have firmware support for queue reset, or in situations where the firmware failed to do the reset. This method uses the same re-emit code that was added for queue reset so it is very likely to be able to keep your system running after a hang.
The main benefit of this reset method is that it's markedly better than a full ASIC reset because it doesn't erase the contents of VRAM, and it can be used on old APUs where ASIC reset is not available at all.
Starting from Linux 7.3 you can benefit from this new recovery method on GCN 1-4.
But… what about the page flip timeout and other issues?
Page flip timeouts are a different beast entirely, because they are caused by issues with the display engine which has a different programming model and userspace interacts with it using a different uAPI. So this is out of scope for the current post.
However, I need to mention that Leo Li has done some excellent work tracking down the root cause for many page flip timeout related issues.
Recommendations
As you can see, a lot of improvements have been made to GPU recovery recently, so if you experience a lot of GPU hangs, my recommendation is to upgrade your kernel if possible. If that's not possible, consider enabling enforcing isolation.
Here is some simple advice in a nutshell:
- For good GPU recovery
- Queue resets with re-emit on RDNA ― use Linux 6.18 or newer
- Queue resets with re-emit on Vega ― use Linux 7.0 or newer
- IP block soft reset with re-emit on GCN 1-4 ― use Linux 7.3 or newer
- For the safest experience, use
amdgpu.enforce_isolation=1
- If you need to use older kernels
- On RDNA, use
amdgpu.enforce_isolation=1which will make queue resets work much better - You need to accept that Vega and older don't have any decent recovery options on those kernels
- On RDNA, use
- If you use newer kernels but still experience problems
- Try
amdgpu.enforce_isolation=1 - If that didn't help, open an issue here and please don't forget to mention your system specs, upload a
dmesglog and write down the steps to reproduce
- Try
Hope this helps!
26 Aug 2026 5:19pm GMT
24 Aug 2026
planet.freedesktop.org
Matthias Klumpp: Sovereign Tech Fellowship for Freedesktop Tasks
In 2025 I was honored to be selected for the first cohort of Sovereign Tech Fellows, a program by Germany's Sovereign Tech Agency to improve the resilience of the open source ecosystem by supporting maintainers directly (complementing their existing support for larger FOSS organizations). Back in 2025, I was only working very limited hours - however, this has changed in 2026.
For the second half of 2026, I am working again as a Sovereign Tech Fellow, but this time with significantly increased hours. After finishing my PhD, I do have time now for new tasks (and new jobs!), and the fellowship presents an amazing opportunity to really advance projects that I maintain or am part of. This also has a very nice effect on contributors and bug reporters, as their feedback gets addressed a lot faster. With some luck, this ultimately will help finding new (co)maintainers for projects as well (although in the age of AI, a lot of how open source used to work is much more uncertain, but that is a matter for a different blog post).
The fellowship is time-limited, so I am intending to make the time I currently have count!
So, what's planned?
I am involved in many projects, but three of them will be getting attention as part of the fellowship. I know I am notoriously slow at blogging, but expect more details on each of them very soon. Here's an overview:
Freedesktop.org, Specifications and Organization
I maintain the Freedesktop Specifications, which is an area of Freedesktop that has traditionally been a bit chaotic. This "worked" in the past, because Freedesktop was never intended to be a formal standards body, but more a shared space where people could throw a lot of code and ideas over the wall and see what sticks and what people can collaborate on.
While I very much love the spirit of this and want to keep it in some form, we definitely would benefit not just from more formalization and better procedures, but also from better organization of the specifications in general. A lot of conflicts can be avoided by that. I will work on improving procedures, crunching through the (lots!) of pending bug reports and MRs, and to make the specifications site better searchable and accessible (similar to how Mozilla's MDN presents information, but I am not sure if we will get quite that far). I also intent to add a compatibility matrix for specifications, so if a desktop opts out of any one of them (or does not implement them yet) that fact is documented and authors of applications know what they can expect. This will allow us to move a lot faster and avoid a lot of conflict, because there is no implicit assumption that "everybody will implement everything" anymore (which has never been quite true anyway).
Hopefully, this will ultimately result in a Freedesktop that is both a lot more useful for application authors who want to bring their project to Linux, as well as developers of desktop environments who need to see which specifications are available and which ones are current.
In addition to that, I have also worked on a Freedesktop.org website refresh, which is pretty much done in its first iteration (pending sysadmin action). The aim there is to have a more official website, separate from user-contributed wiki content, that showcases what Freedesktop is and which projects are using it for hosting. Once the new website is live, I will also review every page again, archive dead projects in their own section and reorganize the software and specifications directory. Those sections are severely outdated and are missing recent efforts from the community, while still containing long-dead old projects (remember HAL?
).
AppStream
A lot of extra maintenance work will be (has been!) done on it. This includes things such as JPEG-XL support (blog post soon), sandboxed media processing, support for newer specification additions, better OARS integration (and potentially migrating it to fd.o infrastructure), improvements and API stabilization for libappstream-compose and a lot of bugfixing and resolution of issues found by AI code review.
AppStream was originally designed to parse only trusted data from vetted Linux distribution sources - this is no longer the case in today's world and in the way Flatpak uses it, so we need to increase resilience of the project.
I am also exploring a project that could vastly improve search accuracy for AppStream. Stay tuned for that.
PackageKit & System Upgrades
Many years ago, people thought we would all migrate to atomic Linux distributions and slowly not need PackageKit anymore. This has not turned out to be the case, and there are still plenty of reasons to use a package-based OS, especially in development environments. At the same time, PackageKit has been basically the same for years, and its older architecture is beginning to show. It being a daemon who's literal job it is to modify the entire system also makes it one of the most security-sensitive components that a Linux system can have, while simultaneously making it near-impossible to sandbox.
My plan is to create PackageKit 2.0 by building on the great foundation of PackageKit 1.0, but modernizing it. This will include simplifying its code and removing a bunch of features that have no more use in modern desktops, while also adding some features that PackageKit never had but that would be useful to expose to frontends (still no to interactivity an terminal-progress forwarding though!). PK 2.0 will also allow me to solve a few design issues that have been worked around in the past, by replacing them with better solutions. This will be a painful transition, as PackageKit 2.0 will break all interfaces PackageKit has - and those interfaces have been frozen for more than a decade. However, I do fully expect this change to be worth the effort.
In addition to that, I intend to look into the offline-update procedure again and improve it. The current multi-reboot operation comes with downsides, that newer systemd features such as soft-reboot can alleviate. The end result should be a much smoother, less annoying offline-update experience for users (I especially want to get rid of updates running on system startup, which I consider quite bad from a usability perspective). The new behavior is in the early drafting stages and may need direct support from systemd. I will share more about it once I can.
That's a lot of tasks!
Yes! I will see how far I get. I am moving project-by-project though, to allow me to focus on one project at a time, rather than scattering my attention continuously. Amazingly, this means that the major tasks for AppStream are already almost done, and we are nearing the 1.2.0 release. AppStream got priority, because the new Freedesktop Flatpak runtime will be released soon, and because I want FlatHub/Flatpak to have access to the new AppStream release sooner. Freedesktop and PackageKit are next on the task list.
Either way, a lot of progress is coming - if you have any feedback or want to help out, please don't hesitate to reach out! All work is happening fully in the open, so you can also chime in on the respective GitHub/GitLab tasks
.
You can also expect blog posts about key features or interesting changes, so stay tuned! 
24 Aug 2026 9:00pm GMT
22 Aug 2026
planet.freedesktop.org
Tomeu Vizoso: Etnaviv NPU update 22: YOLOX support
We have expanded our open-source NPU support for object detection: YOLOX is now running on the Etnaviv driver.

This brings high-performance, open-source AI acceleration to the Vivante VIP line of NPUs, including those found in the NXP i.MX 8M Plus and the Amlogic A311D. Adding YOLOX gives users more options when balancing detection accuracy against available hardware resources.
YOLOX is an open-source object detection model developed by Megvii (Apache 2.0). It is particularly well-suited for edge NPUs, but it is significantly more complex than SSDLite MobileDet (our previously supported model). Getting it running required implementing several new operations in the driver.
As part of this work, we landed support for:
- FullyConnected: A new operation that runs directly on the NN (convolution) cores.
- Reshape, Split, and Concatenate: Handled via metadata changes; these do not execute on the hardware, saving cycles.
- Fused ReLU: Enables activation function hardware on the output.
- Absolute and Logistic: Implemented as lookup table operations on the TP (tensor processing) cores.
- Subtract: Lowered to a convolution, similar to our approach for Add.
- Transpose: Either fused into the next operation or handled as a TP operation.
Additionally, we added support for feature maps in signed 8-bit integers. For certain models, this provides increased accuracy at the exact same computational cost.
This work was performed in partnership with Ideas On Board.
YOLOX is an open-source project developed by Megvii and licensed under the Apache License 2.0.
22 Aug 2026 8:40am GMT
.png)






















