Amdgpu ring gfx_0.0.0 timeout when playing World of Warcraft Midnight

Hi there!

I get random amdgpu ring gfx_0.0.0 timeout crashes while playing WoW Midnight. Everything was fine during TWW. Every other game is fine. The screens go black then return. Most of the time only the game (and steam) crash in this process, sometimes I get an additional “device wedged” and screens freeze. I launch WoW using wine-cachyos 10.0-20260227 with ntsync on and mangohud.

Here’s what I tried:

  • I tried reverting to mesa 25.3.6, but then PLM won’t show anything, just a blinking cursor in the top left corner.
  • I disabled overclocking (removed kernel parameter), but that didn’t change anything.
  • I tried limiting core clock to 2500 MHz (amdgpu detects 2935 MHz), without success.

Here’s the kernel log:

Mär 15 20:28:32 minotar kernel: amdgpu 0000:33:00.0: amdgpu: MODE1 reset
Mär 15 20:28:32 minotar kernel: amdgpu 0000:33:00.0: amdgpu: GPU mode1 reset
Mär 15 20:28:32 minotar kernel: amdgpu 0000:33:00.0: amdgpu: GPU smu mode1 reset
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: GPU reset succeeded, trying to resume
Mär 15 20:28:33 minotar kernel: [drm] PCIE GART of 512M enabled (table at 0x0000008001300000).
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: VRAM is lost due to GPU reset!
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: PSP is resuming...
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: reserve 0x1300000 from 0x85fc000000 for PSP TMR
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: RAP: optional rap ta ucode is not available
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: SECUREDISPLAY: optional securedisplay ta ucode is not available
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: SMU is resuming...
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: smu driver if version = 0x0000003d, smu fw if version = 0x00000040, smu fw pro>
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: SMU driver if version not matched
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: SMU is resumed successfully!
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: [drm] DMUB hardware initialized: version=0x07002F00
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring gfx_0.0.0 uses VM inv eng 0 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.0.0 uses VM inv eng 1 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.1.0 uses VM inv eng 4 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.2.0 uses VM inv eng 6 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.3.0 uses VM inv eng 7 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.0.1 uses VM inv eng 8 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.1.1 uses VM inv eng 9 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.2.1 uses VM inv eng 10 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring comp_1.3.1 uses VM inv eng 11 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring sdma0 uses VM inv eng 12 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring sdma1 uses VM inv eng 13 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring vcn_unified_0 uses VM inv eng 0 on hub 8
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring vcn_unified_1 uses VM inv eng 1 on hub 8
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring jpeg_dec uses VM inv eng 4 on hub 8
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: ring mes_kiq_3.1.0 uses VM inv eng 14 on hub 0
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: amdgpu: GPU reset(1) succeeded!
Mär 15 20:28:33 minotar kernel: amdgpu 0000:33:00.0: [drm] device wedged, but recovered through reset
Mär 15 20:28:44 minotar kernel: amdgpu 0000:33:00.0: [drm] *ERROR* [CRTC:254:crtc-1] flip_done timed out
Mär 15 20:28:44 minotar kernel: amdgpu 0000:33:00.0: amdgpu: [drm] *ERROR* [CRTC:254:crtc-1] hw_done or flip_done timed out

Here’s the output of sudo cachyos-bugreport.sh: 4328f47 I just noticed that the crash is not part of this bugreport as I did multiple reboots inbetween. I will upload a new one after the next crash (shouldn’t take long).

How can I revert mesa to an older version and not break PLM? What can I do to mitigate this issue if a downgrade is not possible?

Best regards,
Mr nUUb

Here’s a new bugreport: 60a4e99 (this one contains the crash)

Additional information: using DX11 or DX12 doesn’t change anything. RT shadows are disabled.

EDIT: I mentioned downgrading mesa which didn’t work. I installed mesa-2:25.3.6-3-x86_64.pkg.tar.zst, opencl-mesa-2:25.3.6-3-x86_64.pkg.tar.zst, vulkan-mesa-implicit-layers-2:25.3.6-3-x86_64.pkg.tar.zst, vulkan-radeon-2:25.3.6-3-x86_64.pkg.tar.zst from /var/cache/pacman/pkg and rebooted. Is this wrong? Should I downgrade even more packages?

Some options that might help:

  1. Use mesa-git. Sometimes the latest Mesa has fixes for these type of issues. It helped a while back with my Intel iGPU resetting constantly.
  2. Update the DXVK and VKD3D-Proton DLLs by downloading the latest Github Actions master-branch artifacts. These newer pre-release DLLs can fix weird issues that have yet to propagate to a stable-release.
  3. Try the latest RC kernel (linux-cachyos-rc 7.0.rc3-2). The AMDGPU kernel driver will have seen some updates in the latest RC kernel that have yet to propagate to a stable kernel release. It is possible that the reset issue has been fixed there but not yet in available a stable kernel.

Thanks for the suggestions, but I don’t like the idea of running release candidates and even newer software (git packages). I am struggling to downgrade mesa. If I run mesa-git, who guarantees that I can switch to a stable version once the bug gets fixed (required there actually is a bug in mesa)? Since the number of people complaining about amdgpu ring gfx_0.0.0 timeouts after updating to mesa 26 grows daily, I would like to downgrade and see if that helps.

This problem is like 3-4 years old at least and probably will not be fixed in the near future. I have it in DX12 games (translated to VKD3D), while DXVK is not affected. Try to switch to DX11 via ingame settings or underclock the GPU via LACT to see if it will help. Reddit, forums, gitlab are the places where you can see that you are not alone and so many people are affected.

I have found the issue. I have to add RADV_DEBUG=nohiz. The issue is reproducible in the dungeon Murder Row. It happens at random places inside this dungeon, but at least one crash is guaranteed. With this setting, the crashes immediately disappear.