Mesa version missmatch cause GPU driver crash

Hi everyone,

I got this bug for the second time and I believe to be what happen:

If mesa libs and the 32 bit version slightly differ, I get driver crash/page fault like this one

amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:169 vmid:0 pasid:0) 
amdgpu 0000:0b:00.0: in page starting at address 0x0000000000000000 from client 10 
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00040D52 
amdgpu 0000:0b:00.0: Faulty UTCL2 client ID: CPG (0x6) 
amdgpu 0000:0b:00.0: MORE_FAULTS: 0x0 
amdgpu 0000:0b:00.0: WALKER_ERROR: 0x1 
amdgpu 0000:0b:00.0: PERMISSION_FAULTS: 0x5 
amdgpu 0000:0b:00.0: MAPPING_ERROR: 0x1 
amdgpu 0000:0b:00.0: RW: 0x1 
amdgpu 0000:0b:00.0: Dumping IP State 
amdgpu 0000:0b:00.0: Dumping IP State Completed 
amdgpu 0000:0b:00.0: [drm] AMDGPU device coredump file has been created 
amdgpu 0000:0b:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data 
amdgpu 0000:0b:00.0: ring gfx_0.0.0 timeout, signaled seq=7663915, emitted seq=7663917 
amdgpu 0000:0b:00.0: Process Main pid 10780 thread vkd3d_queue pid 10990 
amdgpu 0000:0b:00.0: Starting gfx_0.0.0 ring reset 
amdgpu 0000:0b:00.0: MES failed to respond to msg=RESET 
amdgpu 0000:0b:00.0: failed to reset legacy queue 
amdgpu 0000:0b:00.0: reset via MES failed and try pipe reset -110 
amdgpu 0000:0b:00.0: Ring gfx_0.0.0 reset failed 
amdgpu 0000:0b:00.0: GPU reset begin!. Source: 1 
amdgpu 0000:0b:00.0: MES failed to respond to msg=REMOVE_QUEUE 
amdgpu 0000:0b:00.0: failed to unmap legacy queue 

Locally I can see that

> sudo pacman -Syyu
cachyos/lib32-mesa                         1:26.0.5-1   2:26.0.5-2     2,63 MiB
cachyos/lib32-opencl-mesa                  1:26.0.5-1   2:26.0.5-2     0,66 MiB
cachyos/lib32-vulkan-mesa-implicit-layers  1:26.0.5-1   2:26.0.5-2     0,00 MiB
cachyos/lib32-vulkan-radeon                1:26.0.5-1   2:26.0.5-2     0,49 MiB
cachyos/mesa                               1:26.0.5-1   2:26.0.5-4     0,09 MiB
cachyos/opencl-mesa                        1:26.0.5-1   2:26.0.5-4    -1,40 MiB
cachyos/vulkan-mesa-implicit-layers        1:26.0.5-1   2:26.0.5-4     0,00 MiB
cachyos/vulkan-radeon                      1:26.0.5-1   2:26.0.5-4    -0,02 MiB

Yesterday when libs where updated to 2:26.0.5-2 and 2:26.0.5-4 my games crashed after 10-15min mostly all the time

The first time:after waiting 2 days libs32 got updated and everything was fixed

Today the same happened and by downgrading to 1:26.0.5-1 everything was fixed

So is it something expected how could I prevent it to happen while still being able to update regularly?

It’s not my mirror that is late it seems normal. (I forced all mirror lists to the same source)

In the meantime I will create a small script to check futur versions but I feel I’m missing something

Thanks for any help

Setup:
Ryzen 7 5800X3D
Radeon RX 7800 XT