gnu/gcc/f4d579b7a2ace16c9b0b7912940ff6809f8ddf9c genattrtab: table-drive attributes that are functions of one other attribute
A machine description often defines one attribute purely in terms of
another, so that a scheduling model can group the several hundred values
of `type' into the handful its pipeline actually distinguishes. In
arm/types.md:
(define_attr "mul32" "no,yes"
(if_then_else
(eq_attr "type"
"smulxy,smlaxy,smulwy,smlawx,mul,muls,mla,mlas,smlawy,smuad,\
smuadx,smlad,smladx,smusd,smusdx,smlsd,smlsdx,smmul,smmulr,\
smmla,smlald,smlsld")
(const_string "yes")
(const_string "no")))
That is a total function from `type' to `mul32', and nothing else. But
genattrtab does not represent it that way. optimize_attrs substitutes
the definition of `type' into it and folds the result separately for
every insn code, so what comes out is a switch over recog_memoized:
attr_mul32
get_attr_mul32 (rtx_insn *insn ATTRIBUTE_UNUSED)
{
attr_type cached_type ATTRIBUTE_UNUSED;
switch (recog_memoized (insn))
{
case -1:
if (GET_CODE (PATTERN (insn)) != ASM_INPUT
&& asm_noperands (PATTERN (insn)) < 0)
fatal_insn_not_found (insn);
/* FALLTHRU */
if (((cached_type = get_attr_type (insn)) == TYPE_SMULXY)
|| (cached_type == TYPE_SMLAXY)
... 20 more ...
|| (cached_type == TYPE_SMLSLD))
{
return MUL32_YES;
}
else
{
return MUL32_NO;
}
case 424: /* *mulsi_neg_uxtw */
case 423: /* *muldi_neg */
... 10 more ...
case 413: /* mulsi3 */
return MUL32_YES;
default:
return MUL32_NO;
}
}
Two things are worth noticing. The `case -1:' arm, reached for asm
statements whose `type' is only known at run time, already contains the
mapping in its original form: evaluate `type' once, then decide. Every
other arm is that same decision, precomputed for one insn code and
re-emitted. So the switch is a partially evaluated copy of a function
that the file already knows how to write, keyed on the wrong thing.
Emit the mapping directly instead:
static const unsigned char mul32_from_type[] = {
MUL32_NO, MUL32_NO, ..., MUL32_YES, ..., MUL32_NO,
};
attr_mul32
get_attr_mul32 (rtx_insn *insn ATTRIBUTE_UNUSED)
{
return (attr_mul32) mul32_from_type[get_attr_type (insn)];
}
The result is better in three ways. It is one array read rather than a
search over insn codes, so it does not grow when the port gains
patterns, only when `type' gains values. It has a single control-flow
path, so the host compiler has nothing to optimise. And the asm case
needs no special handling at all: get_attr_type still issues
fatal_insn_not_found, and whatever `type' it returns indexes the same
table as any other insn.
This is correct by construction rather than by testing. The machine
description says the attribute is a function of `type', and a table
indexed by `type' is that function. The pass therefore only has to recognise
the shape, which it does before optimize_attrs destroys it: a cond, or
an if_then_else chain, in which every test selects a set of values of
one single other attribute and every value including the default is
constant. check_attr_test has already rewritten (eq_attr "type"
"a,b,c") into an ior chain over single values, so a test is matched as a
boolean combination of eq_attr.
Anything else, in particular match_test, match_operand, eq_attr_alt and
attr_flag, falls back to the existing expansion. An attribute that a
define_insn sets directly is not a function of anything, so that is
checked too.
One extension is needed for the case that motivates all this.
cortex_a57_neon_type is written as a cond over `type', except that its
last arm tests is_neon_type, which is itself a function of `type'. So a
test of an attribute already known to be a function of the driver counts
as a test of the driver, and resolving it is a lookup in that
attribute's own table. get_attr_order already supplies the topological
order that guarantees the dependency is processed first. Without this,
the largest of the transformed attributes is missed.
Across the tree 107 of 558 attributes qualify, 62 of them on s390, where
the driver is `mnemonic' with over a thousand values. On aarch64 there
are eight, and the switches they replace are far from uniform in size:
attribute switch lines table + getter lines
mul32 39 466
widen_mul64 37 466
is_mve_type 75 466
is_neon_type 4406 466
cortex_a53_advsimd_type 4472 466
cortex_a57_neon_type 5146 466
exynos_m1_neon_type 4708 466
tsv110_neon_type 4523 466
A table is always |type| entries, so the three attributes that only a
few patterns use get bigger in source. Emitting a table only when it is
the smaller of the two would need a heuristic, and the shape of the
generated code, not its size, is the point, so all of them are
converted. In total insn-attrtab.cc goes from 4.26MB in 92424 lines to
3.28MB in 72584 lines. It compiles, together with insn-dfatab.cc and
insn-latencytab.cc, in 12% less time, with peak memory 313MB
against 332MB.
Run time is unchanged. Compiling a 39-file C corpus with
-mcpu=cortex-a57 takes the same time either way, within a run-to-run
spread of about 1.5%. Note that attribute values are not memoised, so a
derived attribute now runs get_attr_type's switch rather than its own
inlined copy of it. The two are close enough in size that this does not
show up.
Verified by compiling that corpus with -mcpu=generic, cortex-a57,
cortex-a53, exynos-m1, tsv110, neoverse-v2 and neoverse-n1: 273
compilations, and the assembly does not change. On the same compiler all 39
files differ between -mcpu=generic and -mcpu=cortex-a57, so the pipeline
models these attributes feed are being exercised.
Bootstrapped on aarch64-none-linux-gnu.
gcc/ChangeLog:
* genattrtab.cc (attr_value): Add enum_index.
(attr_desc): Add num_values, derived_from and derived_table.
(get_attr_value, add_attr_value, find_attr): Maintain them.
(attr_value_index, eq_attr_value_set, simple_enum_attr_p)
(sole_tested_attr, find_derived_attrs, write_derived_attr_get): New
functions.
(write_attr_get): Use write_derived_attr_get where it applies.
(main): Call find_derived_attrs before optimize_attrs.
Signed-off-by: Kyrylo Tkachov <ktkachov@nvidia.com>
1 file changed