Hi everyone,
I got this bug for the second time and I believe to be what happen:
If mesa libs and the 32 bit version slightly differ, I get driver crash/page fault like this one
amdgpu 0000:0b:00.0: [gfxhub] page fault (src_id:0 ring:169 vmid:0 pasid:0)
amdgpu 0000:0b:00.0: in page starting at address 0x0000000000000000 from client 10
amdgpu 0000:0b:00.0: GCVM_L2_PROTECTION_FAULT_STATUS:0x00040D52
amdgpu 0000:0b:00.0: Faulty UTCL2 client ID: CPG (0x6)
amdgpu 0000:0b:00.0: MORE_FAULTS: 0x0
amdgpu 0000:0b:00.0: WALKER_ERROR: 0x1
amdgpu 0000:0b:00.0: PERMISSION_FAULTS: 0x5
amdgpu 0000:0b:00.0: MAPPING_ERROR: 0x1
amdgpu 0000:0b:00.0: RW: 0x1
amdgpu 0000:0b:00.0: Dumping IP State
amdgpu 0000:0b:00.0: Dumping IP State Completed
amdgpu 0000:0b:00.0: [drm] AMDGPU device coredump file has been created
amdgpu 0000:0b:00.0: [drm] Check your /sys/class/drm/card1/device/devcoredump/data
amdgpu 0000:0b:00.0: ring gfx_0.0.0 timeout, signaled seq=7663915, emitted seq=7663917
amdgpu 0000:0b:00.0: Process Main pid 10780 thread vkd3d_queue pid 10990
amdgpu 0000:0b:00.0: Starting gfx_0.0.0 ring reset
amdgpu 0000:0b:00.0: MES failed to respond to msg=RESET
amdgpu 0000:0b:00.0: failed to reset legacy queue
amdgpu 0000:0b:00.0: reset via MES failed and try pipe reset -110
amdgpu 0000:0b:00.0: Ring gfx_0.0.0 reset failed
amdgpu 0000:0b:00.0: GPU reset begin!. Source: 1
amdgpu 0000:0b:00.0: MES failed to respond to msg=REMOVE_QUEUE
amdgpu 0000:0b:00.0: failed to unmap legacy queue
Locally I can see that
> sudo pacman -Syyu
cachyos/lib32-mesa 1:26.0.5-1 2:26.0.5-2 2,63 MiB
cachyos/lib32-opencl-mesa 1:26.0.5-1 2:26.0.5-2 0,66 MiB
cachyos/lib32-vulkan-mesa-implicit-layers 1:26.0.5-1 2:26.0.5-2 0,00 MiB
cachyos/lib32-vulkan-radeon 1:26.0.5-1 2:26.0.5-2 0,49 MiB
cachyos/mesa 1:26.0.5-1 2:26.0.5-4 0,09 MiB
cachyos/opencl-mesa 1:26.0.5-1 2:26.0.5-4 -1,40 MiB
cachyos/vulkan-mesa-implicit-layers 1:26.0.5-1 2:26.0.5-4 0,00 MiB
cachyos/vulkan-radeon 1:26.0.5-1 2:26.0.5-4 -0,02 MiB
Yesterday when libs where updated to 2:26.0.5-2 and 2:26.0.5-4 my games crashed after 10-15min mostly all the time
The first time:after waiting 2 days libs32 got updated and everything was fixed
Today the same happened and by downgrading to 1:26.0.5-1 everything was fixed
So is it something expected how could I prevent it to happen while still being able to update regularly?
It’s not my mirror that is late it seems normal. (I forced all mirror lists to the same source)
In the meantime I will create a small script to check futur versions but I feel I’m missing something
Thanks for any help
Setup:
Ryzen 7 5800X3D
Radeon RX 7800 XT
