Hi,
I have tried Qubes 3.1 RC3 and then dom0 memory corruption bug [1] with Radeon GPU is still present with 4.1.13-9.pvops.qubes.x86_64.
It is very easy to reproduce:
- boot Qubes (via EFI or legacy boot)
- login to XFCE
- start a dom0 terminal and in one tab run: watch 'dmesg -T | tail'
- run glxgears
- within a minute or two you notice the 'BUG: Bad rss-counter state mm:ffff880006671c00 idx:0 val:7' in dmesg
- things start crashing with segfaults (including watch, bash, or whatever you try to run)
Note: if you want long enough the system will just reboot by itself
- stop glxgears
- wait a minute or two, the system is back to normal (no more segfaults)
What I tried to debug this:
- stop all VMs in Qubes including service VMs and have just dom0 running on Xen: still segfaults
- boot Debian jessie+backports dom0 with Xen 4.6 and kernel 4.3.0-0.bpo.1-amd64: no segfaults
- enable some debugging/sanitization flags in Qubes kernel's .config, and build a new .debug kernel.
Some of these wouldn't boot, and some just wouldn't show any new messages in dmesg. E.g. CONFIG_DMA_API_DEBUG and CONFIG_IOMMU_STRESS didn't help.
I'm not entirely sure the bug is with the GPU driver: it is just very easy to reproduce with it.
I noticed a crash once with a large download, and there is that pciback warning about the network card [2], but I wasn't able to reproduce it,
so that might've still been caused by the GPU driver.
Do you have any suggestions on how to debug this further to determine where the bug is?
If you have some suggestions for kernel .config flags, or patches, I can try them as well.
Also has anyone else noticed a similar problem when running glxgears, or is this specific to Radeon GPUs?
[1]
https://github.com/QubesOS/qubes-issues/issues/1680#issuecomment-174201448
[2]
https://gist.github.com/edwintorok/59819a7e4eee2907553c#file-messages-firstboot-L2223
--
Edwin Török | Co-founder and Lead Developer
Skylable open-source object storage: reliable, fast, secure
http://www.skylable.com