On 12/09/2015 01:26 AM, Jakob Bohm wrote:
> On 09/12/2015 02:06,
supe...@casperkitty.com wrote:
>> On Tuesday, December 8, 2015 at 4:50:58 PM UTC-6, Jakob Bohm wrote:
>>> Really, what aspect of the standard would require that, after a test
>>> has clearly proven equality?
>>
>> Unless realloc returns null, doing rvalue conversion on the old pointer
>> will invoke Undefined Behavior. Using memcmp to compare the old pointer
>> and the new one would not invoke Undefined Behavior, but the Standard
>> allows for the possibility that pointers may compare bitwise equal without
>> being semantically equivalent (e.g. given
>>
>> int hey(int x)
>> {
>> int foo[4][4];
>> int *p1 = foo[0]+4;
>> int *p2 = foo[1];
>> if (x > 10)
>> printf("Wow!");
>> if (x < 23)
>> *p2 = *p1;
>> return memcmp(p1,p2);
>> }
>>
>> the two pointers would likely compare bitwise-equal, but a compiler would
>> be allowed to make the call to "printf" unconditional.
>>
>
> I don't see where the undefined behavior is (other than calling memcmp
> with the wrong number of arguments and thus probably no relevant
> prototype).
The array pointed at by p1 is foo[0], and has a length of only 4.
Therefore, p1 points one past the end of that array. Such a pointer can
be compared for equality or relative order with other pointers, but
dereferencing such a pointer has undefined behavior. Because of the
contiguity requirements imposed by the C standard on arrays, p1 should
be pointing at the same location at p2, but because dereferencing p1 has
undefined behavior, an implementation could implement run-time pointer
validity tests that would fail upon the execution of *p1 in any context
other than &*p1.
That's not very likely in real life - run-time pointer validity tests
are expensive, and likely only to be implemented for debugging runs. A
more realistic possibility is that an optimizer could see (in a similar
program, not this one) the expressions p1[i] and p2[j], and notice that
since they refer to elements of two different non-overlapping arrays
(foo[0] and foo[1], respectively), will conclude that there is no need
to check for the possibility that p1[i] and p2[j] are aliases for each
other - any values for i and j that make them aliases would render the
behavior of any code that cares about whether they alias each other
undefined.
> If the standard allows foo[0] + 4 to point somewhere other than foo[1],
> then the assignment to *p2 would be undefined, otherwise a clearly
> defined NOP.
No, it's undefined behavior, not a NOP.
...
> For example if the last line was changed to
>
> return (x++) - (x++);
>
> the value of that subtraction has always been undefined, but no
> implementation that should be considered compliant would trigger
> undefined random execution, just unpredictable value. ...
I've seen an explanation how code like that could fail in ways other
than simply returning an unpredictable value, though this explanation is
more likely to apply if x is a large type, such as long double.
Basically, implementation of x++ requires the issuing of multiple
machine instructions. There might be some reason why (x++) - (y++) could
be optimized by interleaving the instructions for x++ with those for
y++, which would not be problematic. The result of interleaving the
instructions for x++ and a second x++, however, could be catastrophic,
and not simply an unexpected value, because of the way those
instructions interact with each other.
Note: I'm describing from memory an explanation provided by someone else
a long time ago - I cannot provide an example of a specific platform for
which these things are all true - but it seems entirely plausible to me
that there could be such a platform, which is all that matters as far as
explaining why the standard was written this way.
Because the behavior of such code is undefined, the optimizer is NOT
required to make sure to avoid that problem. Would you consider such an
implementation to be non-compliant? If so, on what grounds? "undefined
behavior" means that "this standard imposes no requirements" - which is
a pretty comprehensive statement.
> ... It could even
> vary between executions if the result ends up depending on picosecond
> variations in the propagation of incremented x values in a RISC CPU
> that normally requires delay slots between the two increments. ...
That sounds like it might be precisely the kind of catastrophic
consequence that other person was talking about.
> ... However
> the compiler would not be allowed to generate machine code that crashes
> the CPU or triggers an invalid operation exception if that is what the
> CPU specification says that too close accesses of the same register can
> do.
When the behavior is undefined, the C implementation is, by definition,
unconstrained by any requirements imposed by the C standard. It is, in
particular, not constrained to generate code which avoids causing a
crash or an invalid operation exception. A C implementation is normally
so constrained, when any other requirement is applicable, because that
other requirement cannot, in general, be met if there's a crash or an
invalid operation exception.
>>> Also a standard version of the common msize() function (it has a longer
>>> name in GNU libc) would also be nice.
>>
>> On some platforms it would be hard to guarantee anything useful about the
>> semantics; a solution to that would be to have a return value that means
>> "who knows"?
>>
>
> Any platform where free() is not a NOP should be able to look at what
> free() would do if presented with that same pointer and conclude a size
> from that. If the heap code was optimized for the absence of an
> msize() function, implementing an msize() function might involve a
> costly search of the data that malloc() would use to determine if there
> is a malloc-consumable memory address before the free-extending to
> address that free() would calculate if freeing the pointer being
> queried.
>
> But I fail to see how, in practical terms, one could implement malloc()
> and free() without holding enough data to implement msize().
That depends upon what msize() is defined as doing. There's no man page
for any such function on my system, so I can't be sure. If it's defined
as returning the actual size of the allocated block of memory, I have to
agree with you. However, an implementation is free to allocate more
memory than requested, and as I understand it, most implementations
routinely do so, and the malloc() family of functions has no need to
keep track of any number other than the the amount of memory actually
allocated. So being unable to return the requested size seems entirely
reasonable to me.