On Mon, 15 Jun 2020 16:16:53 -0700 (PDT), Steve Scargall <
steve.s...@intel.com> wrote:
>>> Are you saying that the documentation is correct and that there is no
>supported configuration with four DCPMMs per socket on a motherboard with
>three memory channels per socket?
>
>Yes. The keyword here is '*support*'. ie: the ability to call your vendor
>(Hardware or Software) to file a hardware replacement, bug, issue, or
>enhancement request *and receive help*. Being 'supported' and 'I did <this>
>and it works' are two very different things.
Ok, but the server was sold to me originally with 4 DCPMMs on one socket.
I would think that means that they have to support it in that configuration.
>Documentation is always written from the 'what is supported' perspective,
>and it does not always cover every possible combination or scenario. This
>is not unique to PMem. Given your budget, requirements, and objectives,
>stepping outside of a supported config into "you're on your own" seems an
>acceptable risk to you, and that's fine. Many of us have home labs and do
>things that are not officially supported. It's how we learn and improve
>ourselves. However, for production, we must follow the vendor's
>documentation.
Of course.
>The documentation will be updated as, when, or if new population topologies
>are tested, validation, certified, and supported.
>
>>> However, that doesn't answer my question about the claim in the
>documentation that there can be a maximum of one DCPMM per memory channel.
>That isn't stated as a guideline but as a rule.
>
>You'll note the frequent use of 'optimal [memory] performance' throughout
>the documentation for systems with DDR & PMem. Placing two PMem modules on
>the same channel is not recommended for performance reasons. There's no
>physical, electrical, or protocol reason why it wouldn't work (as you found
>out).
If they had just said it wasn't optimal, I wouldn't have been surprised.
> For quality assurance and supportability, some OEMs & ODMs have
>chosen to implement checks within the BIOS that will fail the memory
>training if they detect a memory config that has not been tested,
>validated, and certified. They cannot support a configuration they've not
>vetted. For BIOS's that implement the population matrix, it's definitely a
>hard rule as the system will stop at POST and not boot until you physically
>reconfigure the DIMMs.
I get warnings in the event log about an invalid DIMM configuration but I never saw them until I set up remote monitoring on another machine so I can
change the fan speed (long story).
>For those that don't implement the check, it's
>expected the user follows the documentation. Implementation of the
>population matrix check within the BIOS is not universal across the
>OEM/ODMs, and there's nothing preventing a future BIOS version, without
>warning or notice, from delivering the feature and breaking a system with
>an unsupported config that previously worked.
Good point. Then I'm not going to take any BIOS updates.
>>> But if that really is a rule, it would be impossible to set up a valid
>configuration with more than three DCPMMs per socket on motherboards having
>three memory channels per socket, which is a common configuration.
>
>See previous responses.
>
>While the integrated memory controllers within the Xeon CPUs support up to
>twelve DIMMs, motherboard designs can, and do, implement fewer for a
>variety of reasons. Regardless of the number of physical slots available,
>production customers do follow the DIMM population matrix and install no
>more than one PMem module per memory channel for performance reasons. Using
>fewer than six interleaved PMem modules delivers lower bandwidth and lower
>total capacity, but there are usually other higher priority factors such as
>space, power, and cooling. Populating more than the supported number of
>slots with PMem will usually tip you over the power & cooling thresholds
>that made you choose a system with fewer DIMM slots in the first place.
>
>(This is not relevant to your specific requirements and objectives, but is
>relevant to production use and requirements).
Yes, I understand that. Some day maybe I'll get some funding so I can buy a supported configuration and still have enough capacity to run my tests.
:-)
Thanks for the detailed response.
------------
Steve Heller