Can't start new apps or close existing ones after lock screen wake

I’m running the KDE flavour of CachyOS. On the previous night I locked my PC. When I returned to it in the morning and unlocked it, I could log back in. However I could not close or open apps, all of them resulted with a notification saying remote network timeout. I could not open a new terminal, but luckily I had one already open. I could not run any sudo commands all of them returned with Executing command "/usr/bin/sudo" timed out.. systemctl commands did the same: Connection timed out. Also I couldn’t switch TTY with ctrl+alt+f3/f4/f5 shortcut for an emergency shell. I could only recover from this by pressing the reset button, because even the safe restart feature through reboot command timed out. By this point the system was running for 3 days. I haven’t used sleep mode. Why does this happen? Can I prevent it somehow?

I think I had the same issue but in different form before the latest update. It would just freeze entirely after I move the mouse to wake up the screen to see the lock screen without the screens ever waking up. So I guess some improvements have been made?

Some more context:

cat /proc/cmdline

root=UUID=5916ee67-621a-4200-8e03-21aa90cdb33b rw quiet amdgpu.gfx_off=0 amdgpu.gpu_recovery=1 initrd=\boot\initramfs-linux-cachyos.img

  • Host: X870 Pro RS
  • Kernel: Linux 6.19.6-2-cachyos
  • DE: KDE Plasma 6.6.2
  • WM: KWin (Wayland)
  • WM Theme: Breeze
  • CPU: AMD Ryzen 9 9950X3D
  • GPU 1: AMD Radeon RX 7900 XTX [Discrete] (it’s a saphire nitro+, it’s running with the second bios which gives slight performance boost - including this because I always see a message regarding it in dmesg :smiley: )
  • GPU 2: AMD Radeon Graphics [Integrated]
  • 128gb memory

I still have SDDM installed, I didn’t migrate to plasma login manager.

I tried to troubleshoot this with AI but it sent me on a goose chase, with commands not even working on a system with non borked state.

Hello,

This is one of those threads I had saved for some reason.
Sorry it is so late.

Did you ever get anywhere with this?

I might also mention that a good start would be to execute;

sudo cachyos-bugreport.sh

and share the resulting link here.

Thanks for coming back to me.
I tried the cachyos bugreport utility, but it didn’t show anything meaningful so I thought I skip posting it. Since then this issue only happened one more time. Two weeks ago I upgraded the system and migrated to plasma login manager. Since then I had no issues. The current uptime as of writing is 5 days, 29 mins.
Other possible suspects in my mind are the rootless podman + jellyfin, and my tv plus the PCON.

  • podman + jellyfin: I watch shows from my tablet in bed from jellyfin running on this system while it’s locked - and it has some relevance to lingering user session with systemd
  • TV plus PCON: display port to hdmi adapter as my tv is connected to my desktop pc with an hdmi cable. Before it caused minor flickering on the other screens when I turn off the tv and the tv transitions to deep sleep. Also during the night the tv can wake from deep sleep to install updates, which might cause this issue somehow. However the latest kernel version 6.19.10-1-cachyos had many fixes around PCON so I’m guessing that fixed the issue, but nothing concrete. :man_shrugging: No more flickering either since then.

Edit: Well it happened again (at 17:33 BST), but this time around it was similar to other reports where you can’t even login anymore and the login screen gets stuck. I could login through ssh, and jellyfin was working fine (I could watch shows from it).
Here is the bug report: c2e5411 (unplugged the tv so it doesn’t appear on the list of screens)

The interesting part is:

Logs
Apr 16 17:33:27 <hostname-redacted> kernel: amdgpu 0000:03:00.0: amdgpu: Fence fallback timer expired on ring sdma0
Apr 16 17:33:30 <hostname-redacted> kernel: amdgpu 0000:03:00.0: amdgpu: device lost from bus!
Apr 16 17:33:30 <hostname-redacted> kernel: amdgpu 0000:03:00.0: amdgpu: SMU: response:0xFFFFFFFF for index:41 param:0x00000000 message:DisallowGfxOff?
Apr 16 17:33:30 <hostname-redacted> kernel: amdgpu 0000:03:00.0: amdgpu: Failed to disable gfxoff!
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 1 (reg 283, index 1) beyond range (38, 930)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 2 (reg 594, index 2) beyond range (77, 969)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 3 (reg 906, index 3) beyond range (116, 1008)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 4 (reg 1217, index 4) beyond range (155, 1047)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 468 (reg 4019, index 1) beyond range (505, 1397)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 469 (reg 4331, index 2) beyond range (544, 1436)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 470 (reg 4642, index 3) beyond range (583, 1475)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 471 (reg 4642, index 4) beyond range (583, 1475)
Apr 16 17:33:33 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 896 (reg 7444, index 1) beyond range (933, 1825)
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 897 (reg 7445, index 2) beyond range (933, 1825)
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 898 (reg 7756, index 3) beyond range (972, 1864)
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 899 (reg 7756, index 4) beyond range (972, 1864)
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Frame ID 900 (reg 8067, index 5) beyond range (1011, 1903)
Apr 16 17:33:34 <hostname-redacted> kernel: xhci_hcd 0000:0d:00.0: Ignore frame ID field, use SIA bit instead
Apr 16 17:33:49 <hostname-redacted> xdg-desktop-portal-kde[2425]: QDBusMarshaller: cannot add a null QDBusVariant
Apr 16 17:33:49 <hostname-redacted> xdg-desktop-portal-kde[2425]: QDBusConnection: Could not emit signal org.freedesktop.impl.portal.Settings.SettingChanged: Marshalling failed: Invalid QVariant passed in arguments
Apr 16 17:34:12 <hostname-redacted> kernel: amdgpu 0000:03:00.0: amdgpu: ring sdma0 timeout, signaled seq=5197980, emitted seq=5197980
Apr 16 17:34:12 <hostname-redacted> kernel: amdgpu 0000:03:00.0: amdgpu: Starting sdma0 ring reset
Apr 16 17:34:12 <hostname-redacted> kernel: amdgpu 0000:03:00.0: amdgpu: Ring sdma0 reset succeeded
Apr 16 17:34:24 <hostname-redacted> systemd[1718]: dbus-:<email-address-redacted>: Failed with result 'exit-code'.

Hmm..

Since AMD is making a mention .. even though its not quite the same .. in the off chance that anything here is related;

Just an update on this: After the latest set of updates the issue never came back. :crossed_fingers: