#415 SDDM/Kwin_wayland core dump - results in black screen on boot
Closed: Fixed by mattm. Opened by mattm.

On Framework AMD / Ryzen 7040 series platform, Fedora Kinoite 39, boots to black screen.

Journalctl shows some errors for amdgpu sdma ring time out and then performs amdgpu reset and then later Process 1282 (kwin_wayland) of user 981 dumped core.

Logging into a TTY and running startplasma-wayland starts correctly and everything appears to be working correctly.

Framework have advised that some late firmware is coming to LVFS, but nothing specific about this. I understand that AMD 7040 is a new platform.

I have attached a dump of the journal. Specific error appear at 13:40:49

I will reconfig sddm to run on X and see if it still crashes and will update the issue
journal.log


put sddm-x11.conf into /etc/sddm.conf.d/ as defined here https://fedoraproject.org/wiki/Changes/WaylandByDefaultForSDDM#How_To_Test and sddm starts correctly and is able to log into the plasma wayland session

confirming core dump appears to be caused by sddm running under wayland

ideally we need the actual backtrace from the coredump. you should be able to get it with coredumpctl gdb <pid>, after running coredumpctl list to find the PID (any unique identifier from the relevant crash should work, actually). Or you can use abrt-gui.

I believe this should be what you need
core.kwin_wayland.981.610284b90f2242af887209a2042e83af.1282.1697254849000000.zst

This is amdgpu crashing:

Oct 14 13:40:37 yukon sddm-helper-start-wayland[1281]: "OpenGL vendor string:                   AMD\nOpenGL renderer string:                 AMD Radeon Graphics (gfx1103_r1, LLVM 16.0.6, DRM 3.54, 6.5.6-300.fc39.x86_64)\nOpenGL version string:                  4.6 (Core Profile) Mesa 23.2.1\nOpenGL shading language version string: 4.60\nDriver:                                 Unknown\nGPU class:                              Unknown\nOpenGL version:                         4.6\nGLSL version:                           4.60\nMesa version:                           23.2.1\nLinux kernel version:                   6.5.6\nRequires strict binding:                no\nGLSL shaders:                           yes\nTexture NPOT support:                   yes\nVirtual Machine:                        no\n"
Oct 14 13:40:37 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=14
Oct 14 13:40:37 yukon kernel: [drm:amdgpu_mes_reg_write_reg_wait [amdgpu]] *ERROR* failed to reg_write_reg_wait
Oct 14 13:40:38 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=14
Oct 14 13:40:38 yukon kernel: [drm:amdgpu_mes_reg_write_reg_wait [amdgpu]] *ERROR* failed to reg_write_reg_wait
Oct 14 13:40:38 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=14
Oct 14 13:40:38 yukon kernel: [drm:amdgpu_mes_reg_write_reg_wait [amdgpu]] *ERROR* failed to reg_write_reg_wait
Oct 14 13:40:38 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=14
Oct 14 13:40:38 yukon kernel: [drm:amdgpu_mes_reg_write_reg_wait [amdgpu]] *ERROR* failed to reg_write_reg_wait
[...]
Oct 14 13:40:47 yukon kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* ring sdma0 timeout, signaled seq=37, emitted seq=39
Oct 14 13:40:47 yukon kernel: [drm:amdgpu_job_timedout [amdgpu]] *ERROR* Process information: process  pid 0 thread  pid 0
Oct 14 13:40:47 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: GPU reset begin!
Oct 14 13:40:47 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:47 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:48 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:48 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:48 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:48 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:48 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:48 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:48 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:48 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:48 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:48 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:48 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:48 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:48 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:48 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:49 yukon kernel: [drm:mes_v11_0_submit_pkt_and_poll_completion.constprop.0 [amdgpu]] *ERROR* MES failed to response msg=3
Oct 14 13:40:49 yukon kernel: [drm:amdgpu_mes_unmap_legacy_queue [amdgpu]] *ERROR* failed to unmap legacy queue
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: MODE2 reset
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: GPU reset succeeded, trying to resume
Oct 14 13:40:49 yukon kernel: [drm] PCIE GART of 512M enabled (table at 0x000000801FD00000).
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: SMU is resuming...
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: SMU is resumed successfully!
Oct 14 13:40:49 yukon kernel: [drm] DMUB hardware initialized: version=0x08002300
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:264
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:272
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:280
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:288
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:264
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:272
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:280
Oct 14 13:40:49 yukon kernel: [drm] REG_WAIT timeout 1us * 1000 tries - dcn314_dsc_pg_control line:288
Oct 14 13:40:49 yukon kernel: [drm] kiq ring mec 3 pipe 1 q 0
Oct 14 13:40:49 yukon kernel: [drm] VCN decode and encode initialized successfully(under DPG Mode).
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: [drm:jpeg_v4_0_hw_init [amdgpu]] JPEG decode initialized successfully.
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring gfx_0.0.0 uses VM inv eng 0 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.0.0 uses VM inv eng 1 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.1.0 uses VM inv eng 4 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.2.0 uses VM inv eng 6 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.3.0 uses VM inv eng 7 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.0.1 uses VM inv eng 8 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.1.1 uses VM inv eng 9 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.2.1 uses VM inv eng 10 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring comp_1.3.1 uses VM inv eng 11 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring sdma0 uses VM inv eng 12 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring vcn_unified_0 uses VM inv eng 0 on hub 8
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring jpeg_dec uses VM inv eng 1 on hub 8
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: ring mes_kiq_3.1.0 uses VM inv eng 13 on hub 0
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: recover vram bo from shadow start
Oct 14 13:40:49 yukon kernel: amdgpu 0000:c1:00.0: amdgpu: recover vram bo from shadow done
[...]
Oct 14 13:40:49 yukon kernel: [drm] Skip scheduling IBs!
Oct 14 13:40:49 yukon sddm-helper-start-wayland[1281]: "amdgpu: The CS has been rejected (-125). Recreate the context.\n"
Oct 14 13:40:49 yukon kernel: [drm] Skip scheduling IBs!
Oct 14 13:40:49 yukon sddm-helper-start-wayland[1281]: "amdgpu: The CS has been rejected (-125). Recreate the context.\n"

this issue was resolved by the beta firmware released today version 3.0.30 it is available on the lvfs beta channel

fwupdmgr enable-remote lvfs-testing
fwupdmgr update

Metadata Update from @mattm:
- Issue close_status updated to: Fixed
- Issue status updated to: Closed (was: Open)

Metadata