gnu/gcc/56c81c325cccf432954ddea4c73069e2fa047922 [PATCH] match.pd: Fold ((X + C1) | (C2 - X)) < 0 to an unsigned range test
The sign bit of (X + C1) | (C2 - X) is the disjunction of the two sign
bits, so the test is X < -C1 || X > C2, that is (unsigned) (X + C1) >
C1 + C2. With U for X + C1 and D for C1 + C2 the second operand is
D - U, and for 0 <= D < 2 ** (prec - 1) the disjunction is negative
exactly when U > D unsigned; D non-negative in the type is the only
guard, and no no-overflow assumption is involved. D is tested with
wi::ge_p (..., SIGNED), not tree_int_cst_sgn: after over_widening the
constants have an unsigned type, for which tree_int_cst_sgn is 1
whatever the bit pattern.
Matched in every spelling that reaches match.pd: the comparison, the
comparison of a same-precision conversion of it (what narrow types and
unsigned data reduce to), the sign-bit extraction x264 uses in
encoder/cavlc.c, and the C1 == 0 and C2 == 0 forms, where the addition
is absent and the subtraction is a negate. Those need matchers of their
own: genmatch cannot make a binary operand optional, and negate does not
fit a (for) with minus. The disjunction needs single_use, or both
operands stay live and the fold only adds a compare (7 -> 9 x86_64).
ABS_EXPR <X> > C and <= C are the same range test, and so is the
ABSU_EXPR pair; ABSU_EXPR needs no assumption about the minimum value,
ABS_EXPR needs TYPE_OVERFLOW_UNDEFINED. C == TYPE_MAX is left to the
existing fold to a constant; == and != are not range tests and are
untouched. The addition is built in the unsigned type, so the fold does
not invent a signed overflow for the sanitizer to report.
All of it is GIMPLE-only, as single_use is always true on GENERIC, and
declines under -fsanitize=signed-integer-overflow, where the operand it
drops carries a check. The nearest pattern, "Optimize (a | b) < 0 to
a < 0 if b is non-negative", does not fire as neither operand is known
non-negative. Not matched: ((X + C1) | ~X) < 0, spelled with a
BIT_NOT_EXPR rather than a minus, and the arithmetic-shift 0/-1 mask
spelling, which is also how vectors reach match.pd.
aarch64 -O2: the x264 site 6 -> 4 insns, the range test 6 -> 4, abs > C
5 -> 4; x86_64 6 -> 5, 6 -> 5, 7 -> 5. Checked against a bit-exact
model over every 16-bit X for 26 constant pairs in all four spellings,
both ABS forms and ABSU_EXPR, and over 32-bit edge and random values.
Bootstrapped and regression tested on x86_64-pc-linux-gnu.
Assisted-by: Claude Opus 5 (Anthropic)
gcc/ChangeLog:
* match.pd (ior_range_test, ior_range_test_1, ior_range_test_0):
New matchers.
(((X + C1) | (C2 - X)) < 0): New simplification.
((unsigned) ((X + C1) | (C2 - X)) >> (prec - 1)): Likewise.
(ABS_EXPR <X> > C): Likewise.
gcc/testsuite/ChangeLog:
* gcc.dg/torture/abs-gt-1.c: New test.
* gcc.dg/torture/ior-range-test-1.c: New test.
* gcc.dg/tree-ssa/ior-range-test-1.c: New test.
* gcc.dg/tree-ssa/ior-range-test-2.c: New test.
* gcc.dg/tree-ssa/ior-range-test-3.c: New test.
Signed-off-by: Dominic P <gcc@gcc.dp11.uk>
6 files changed