Inaccurate DCPMM population documentation

19 views
Skip to first unread message

steve

unread,
Jun 14, 2020, 2:45:30 PM6/14/20
to pmem
Hi all,

I've just added a second CPU to my X11DAi-N based server, so now I can have over 1 TB of DCPMM in total.

But I've discovered that the documentation on how to populate DCPMMs in the slots is inaccurate. This is true for the documentation from SuperMicro
(the motherboard manufacturer) and from Intel.

What the documentation says is that you can have only one DCPMM per memory channel.

The X11DAi-N motherboard has only three memory channels per socket (as I found out while researching memory population rules).

When I got the server, it had one CPU and 4x 125 GB DCPMMs; clearly that means that you can in fact have two DCPMMs on the same memory channel, not
just one.

This makes me wonder: with a total of 8 slots per socket, how many DCPMMs can I really have on one socket?

I imagine that the rule that you need at least one DRAM stick is actually a rule rather than a guideline, but even accounting for that, could I have
5, 6, or 7 DCPMMs? That would be good to know.
------------
Steve Heller

Steve Scargall

unread,
Jun 15, 2020, 10:55:47 AM6/15/20
to pmem
Hi Steve,

The DIMM population guidelines matrix and notes are intended to provide an optimal performance configuration for the system. The documentation recommends both CPU sockets have the same DIMM population for balanced performance. Of course, there is a difference between what works and what is supported, tested, and validated by the vendor(s). If you ever need to log a support call, they'll likely have you implement the documented population before proceeding and escalating to the next level engineering. Additionally, it could be difficult to reproduce any performance numbers you release as production systems will have the recommended population. 

/Steve


steve

unread,
Jun 15, 2020, 2:59:19 PM6/15/20
to Steve Scargall, pmem
I understand that what I'm doing may not be optimal in a theoretical sense, but being self-funded, I have to get the most bang for my buck.

That means using configurations that aren't what others might use in production, assuming of course that everyone who wants to use pmem in production
has an unlimited budget.

However, that doesn't answer my question about the claim in the documentation that there can be a maximum of one DCPMM per memory channel. That isn't
stated as a guideline but as a rule.

But if that really is a rule, it would be impossible to set up a valid configuration with more than three DCPMMs per socket on motherboards having
three memory channels per socket, which is a common configuration.

Are you saying that the documentation is correct and that there is no supported configuration with four DCPMMs per socket on a motherboard with three
memory channels per socket? If so, that seems very odd (no pun intended).
------------
Steve Heller

Steve Scargall

unread,
Jun 15, 2020, 7:16:53 PM6/15/20
to pmem
>> Are you saying that the documentation is correct and that there is no supported configuration with four DCPMMs per socket on a motherboard with three memory channels per socket?

Yes. The keyword here is 'support'. ie: the ability to call your vendor (Hardware or Software) to file a hardware replacement, bug, issue, or enhancement request and receive help. Being 'supported' and 'I did <this> and it works' are two very different things. 

Documentation is always written from the 'what is supported' perspective, and it does not always cover every possible combination or scenario. This is not unique to PMem. Given your budget, requirements, and objectives, stepping outside of a supported config into "you're on your own" seems an acceptable risk to you, and that's fine. Many of us have home labs and do things that are not officially supported. It's how we learn and improve ourselves. However, for production, we must follow the vendor's documentation.

The documentation will be updated as, when, or if new population topologies are tested, validation, certified, and supported.

>> However, that doesn't answer my question about the claim in the documentation that there can be a maximum of one DCPMM per memory channel. That isn't stated as a guideline but as a rule.

You'll note the frequent use of 'optimal [memory] performance' throughout the documentation for systems with DDR & PMem. Placing two PMem modules on the same channel is not recommended for performance reasons. There's no physical, electrical, or protocol reason why it wouldn't work (as you found out). For quality assurance and supportability, some OEMs & ODMs have chosen to implement checks within the BIOS that will fail the memory training if they detect a memory config that has not been tested, validated, and certified. They cannot support a configuration they've not vetted. For BIOS's that implement the population matrix, it's definitely a hard rule as the system will stop at POST and not boot until you physically reconfigure the DIMMs. For those that don't implement the check, it's expected the user follows the documentation. Implementation of the population matrix check within the BIOS is not universal across the OEM/ODMs, and there's nothing preventing a future BIOS version, without warning or notice, from delivering the feature and breaking a system with an unsupported config that previously worked.

>> But if that really is a rule, it would be impossible to set up a valid configuration with more than three DCPMMs per socket on motherboards having three memory channels per socket, which is a common configuration.

See previous responses. 

While the integrated memory controllers within the Xeon CPUs support up to twelve DIMMs, motherboard designs can, and do, implement fewer for a variety of reasons. Regardless of the number of physical slots available, production customers do follow the DIMM population matrix and install no more than one PMem module per memory channel for performance reasons. Using fewer than six interleaved PMem modules delivers lower bandwidth and lower total capacity, but there are usually other higher priority factors such as space, power, and cooling. Populating more than the supported number of slots with PMem will usually tip you over the power & cooling thresholds that made you choose a system with fewer DIMM slots in the first place. (This is not relevant to your specific requirements and objectives, but is relevant to production use and requirements).

HTH


steve

unread,
Jun 15, 2020, 7:28:30 PM6/15/20
to Steve Scargall, pmem
On Mon, 15 Jun 2020 16:16:53 -0700 (PDT), Steve Scargall <steve.s...@intel.com> wrote:

>>> Are you saying that the documentation is correct and that there is no
>supported configuration with four DCPMMs per socket on a motherboard with
>three memory channels per socket?
>
>Yes. The keyword here is '*support*'. ie: the ability to call your vendor
>(Hardware or Software) to file a hardware replacement, bug, issue, or
>enhancement request *and receive help*. Being 'supported' and 'I did <this>
>and it works' are two very different things.

Ok, but the server was sold to me originally with 4 DCPMMs on one socket.
I would think that means that they have to support it in that configuration.

>Documentation is always written from the 'what is supported' perspective,
>and it does not always cover every possible combination or scenario. This
>is not unique to PMem. Given your budget, requirements, and objectives,
>stepping outside of a supported config into "you're on your own" seems an
>acceptable risk to you, and that's fine. Many of us have home labs and do
>things that are not officially supported. It's how we learn and improve
>ourselves. However, for production, we must follow the vendor's
>documentation.

Of course.

>The documentation will be updated as, when, or if new population topologies
>are tested, validation, certified, and supported.
>
>>> However, that doesn't answer my question about the claim in the
>documentation that there can be a maximum of one DCPMM per memory channel.
>That isn't stated as a guideline but as a rule.
>
>You'll note the frequent use of 'optimal [memory] performance' throughout
>the documentation for systems with DDR & PMem. Placing two PMem modules on
>the same channel is not recommended for performance reasons. There's no
>physical, electrical, or protocol reason why it wouldn't work (as you found
>out).

If they had just said it wasn't optimal, I wouldn't have been surprised.

> For quality assurance and supportability, some OEMs & ODMs have
>chosen to implement checks within the BIOS that will fail the memory
>training if they detect a memory config that has not been tested,
>validated, and certified. They cannot support a configuration they've not
>vetted. For BIOS's that implement the population matrix, it's definitely a
>hard rule as the system will stop at POST and not boot until you physically
>reconfigure the DIMMs.

I get warnings in the event log about an invalid DIMM configuration but I never saw them until I set up remote monitoring on another machine so I can
change the fan speed (long story).

>For those that don't implement the check, it's
>expected the user follows the documentation. Implementation of the
>population matrix check within the BIOS is not universal across the
>OEM/ODMs, and there's nothing preventing a future BIOS version, without
>warning or notice, from delivering the feature and breaking a system with
>an unsupported config that previously worked.

Good point. Then I'm not going to take any BIOS updates.

>>> But if that really is a rule, it would be impossible to set up a valid
>configuration with more than three DCPMMs per socket on motherboards having
>three memory channels per socket, which is a common configuration.
>
>See previous responses.
>
>While the integrated memory controllers within the Xeon CPUs support up to
>twelve DIMMs, motherboard designs can, and do, implement fewer for a
>variety of reasons. Regardless of the number of physical slots available,
>production customers do follow the DIMM population matrix and install no
>more than one PMem module per memory channel for performance reasons. Using
>fewer than six interleaved PMem modules delivers lower bandwidth and lower
>total capacity, but there are usually other higher priority factors such as
>space, power, and cooling. Populating more than the supported number of
>slots with PMem will usually tip you over the power & cooling thresholds
>that made you choose a system with fewer DIMM slots in the first place.
>
>(This is not relevant to your specific requirements and objectives, but is
>relevant to production use and requirements).

Yes, I understand that. Some day maybe I'll get some funding so I can buy a supported configuration and still have enough capacity to run my tests.
:-)

Thanks for the detailed response.
------------
Steve Heller

Jan K

unread,
Jun 16, 2020, 4:23:02 AM6/16/20
to pmem
> There's no physical, electrical, or protocol reason why it wouldn't work
> (as you found out). For quality assurance and supportability, some OEMs
> & ODMs have chosen to implement checks within the BIOS that will fail
> the memory training if they detect a memory config that has not been
> tested, validated, and certified. They cannot support a configuration
> they've not vetted. For BIOS's that implement the population matrix,
> it's definitely a hard rule as the system will stop at POST and not boot
> until you physically reconfigure the DIMMs. For those that don't
> implement the check, it's expected the user follows the documentation.
> Implementation of the population matrix check within the BIOS is not
> universal across the OEM/ODMs, and there's nothing preventing a future
> BIOS version, without warning or notice, from delivering the feature and
> breaking a system with an unsupported config that previously worked.

In my opinion implementing a check that stops the machine at POST without
a "physical, electrical, or protocol reason why it wouldn't work" is a
deliberate evil from the user point of view.

A warning that the configuration is unsupported / unrecognized is definitely
desired. But stopping the machine? What is the rationale reason for that?
If one wants to shoot oneself in the foot, just let him. He had been warned.

The only thing you can achieve by failing in place of warning is to stop
people from using that particular hardware, for the setup they want works
on equivalent hardware without such stiff checks. Even if you do not plan
to experiment, a tag of "overly restrivtive" may be the factor that decides
against one vendor in favor of the other.

Regards,
Jan

2020-06-16 1:16 GMT+02:00, Steve Scargall <steve.s...@intel.com>:
>>> Are you saying that the documentation is correct and that there is no
> supported configuration with four DCPMMs per socket on a motherboard with
> three memory channels per socket?
>
> Yes. The keyword here is '*support*'. ie: the ability to call your vendor
> (Hardware or Software) to file a hardware replacement, bug, issue, or
> enhancement request *and receive help*. Being 'supported' and 'I did <this>
> --
> You received this message because you are subscribed to the Google Groups
> "pmem" group.
> To unsubscribe from this group and stop receiving emails from it, send an
> email to pmem+uns...@googlegroups.com.
> To view this discussion on the web visit
> https://groups.google.com/d/msgid/pmem/1f590234-e57b-40fb-9175-04b81e5e495eo%40googlegroups.com.
>

steve

unread,
Jun 16, 2020, 8:24:40 AM6/16/20
to Jan K, pmem
On Tue, 16 Jun 2020 10:22:58 +0200, Jan K <jan.z....@gmail.com> wrote:

>> There's no physical, electrical, or protocol reason why it wouldn't work
>> (as you found out). For quality assurance and supportability, some OEMs
>> & ODMs have chosen to implement checks within the BIOS that will fail
>> the memory training if they detect a memory config that has not been
>> tested, validated, and certified. They cannot support a configuration
>> they've not vetted. For BIOS's that implement the population matrix,
>> it's definitely a hard rule as the system will stop at POST and not boot
>> until you physically reconfigure the DIMMs. For those that don't
>> implement the check, it's expected the user follows the documentation.
>> Implementation of the population matrix check within the BIOS is not
>> universal across the OEM/ODMs, and there's nothing preventing a future
>> BIOS version, without warning or notice, from delivering the feature and
>> breaking a system with an unsupported config that previously worked.
>
>In my opinion implementing a check that stops the machine at POST without
>a "physical, electrical, or protocol reason why it wouldn't work" is a
>deliberate evil from the user point of view.
>
>A warning that the configuration is unsupported / unrecognized is definitely
>desired. But stopping the machine? What is the rationale reason for that?
>If one wants to shoot oneself in the foot, just let him. He had been warned.

Yes, I'd like to hear the rationale for that too.

>The only thing you can achieve by failing in place of warning is to stop
>people from using that particular hardware, for the setup they want works
>on equivalent hardware without such stiff checks. Even if you do not plan
>to experiment, a tag of "overly restrivtive" may be the factor that decides
>against one vendor in favor of the other.

If you have an unlimited budget, maybe you wouldn't mind how restrictive they are.
But how many organizations (not to mention individuals) have such a budget?
------------
Steve Heller

Steve Scargall

unread,
Jun 16, 2020, 11:17:33 AM6/16/20
to pmem
On Tuesday, June 16, 2020 at 2:23:02 AM UTC-6, Jan K wrote:
In my opinion implementing a check that stops the machine at POST without
a "physical, electrical, or protocol reason why it wouldn't work" is a
deliberate evil from the user point of view.

A warning that the configuration is unsupported / unrecognized is definitely
desired. But stopping the machine? What is the rationale reason for that?
If one wants to shoot oneself in the foot, just let him. He had been warned.

The only thing you can achieve by failing in place of warning is to stop
people from using that particular hardware, for the setup they want works
on equivalent hardware without such stiff checks. Even if you do not plan
to experiment, a tag of "overly restrivtive" may be the factor that decides
against one vendor in favor of the other.

Regards,
Jan

The rationale is .... it's not supported for the intended primary use-case which is production environments. I understand the desire to tinker, but talk to your server/BIOS vendor if you are dissatisfied with a decision they made for >99% of their paying customer base.

Memory training information messages, warnings, and fatal error(s) are reported by the BIOS very early in the POST process (before you get the option of pressing F2 to enter the BIOS). You can see these on the console and BMC (or equivalent). Depending on the error severity, the BIOS will decide to stop and wait for manual intervention or continue. The messages will be accompanied by beeps and LED combinations to help troubleshoot the problem if the message isn't clear. There should be a lookup table in the manuals to help decode the beeps and LEDs.

There is a 'Memory configuration' -> 'Halt on mem Training Error' BIOS option, 'Enabled' by default, which may help get around certain situations. 

st...@steveheller.org

unread,
Jun 16, 2020, 11:27:05 AM6/16/20
to Steve Scargall, pmem
If that's the rationale, it is extremely shortsighted on their part.
How does the software used in production environments come to exist?
Often not in production environments.

And even if a non-supported configuration is a fatal error in a
production environment,
I still don't see why it would hurt to have an option to make it a
warning.

> Memory training information messages, warnings, and fatal error(s) are
> reported by the BIOS very early in the POST process (before you get
> the option of pressing F2 to enter the BIOS). You can see these on the
> console and BMC (or equivalent). Depending on the error severity, the
> BIOS will decide to stop and wait for manual intervention or continue.
> The messages will be accompanied by beeps and LED combinations to help
> troubleshoot the problem if the message isn't clear. There should be a
> lookup table in the manuals to help decode the beeps and LEDs.
>
> There is a 'Memory configuration' -> 'Halt on mem Training Error' BIOS
> option, 'Enabled' by default, which may help get around certain
> situations.

So I assume that option should be disabled in order to continue?

Reply all
Reply to author
Forward
0 new messages