)]}'
{
  "commit": "f4d579b7a2ace16c9b0b7912940ff6809f8ddf9c",
  "tree": "84c7fd34061e85fb374181c2110760cbca40248c",
  "parents": [
    "6e33c9be872cc7d2474b4224d1a16515e1b0d2b8"
  ],
  "author": {
    "name": "Kyrylo Tkachov",
    "email": "ktkachov@nvidia.com",
    "time": "Fri Aug 07 11:23:20 2026 +0200"
  },
  "committer": {
    "name": "Kyrylo Tkachov",
    "email": "ktkachov@nvidia.com",
    "time": "Tue Sep 15 15:39:53 2026 +0200"
  },
  "message": "genattrtab: table-drive attributes that are functions of one other attribute\n\nA machine description often defines one attribute purely in terms of\nanother, so that a scheduling model can group the several hundred values\nof `type\u0027 into the handful its pipeline actually distinguishes.  In\narm/types.md:\n\n  (define_attr \"mul32\" \"no,yes\"\n    (if_then_else\n      (eq_attr \"type\"\n       \"smulxy,smlaxy,smulwy,smlawx,mul,muls,mla,mlas,smlawy,smuad,\\\n        smuadx,smlad,smladx,smusd,smusdx,smlsd,smlsdx,smmul,smmulr,\\\n        smmla,smlald,smlsld\")\n      (const_string \"yes\")\n      (const_string \"no\")))\n\nThat is a total function from `type\u0027 to `mul32\u0027, and nothing else.  But\ngenattrtab does not represent it that way.  optimize_attrs substitutes\nthe definition of `type\u0027 into it and folds the result separately for\nevery insn code, so what comes out is a switch over recog_memoized:\n\n  attr_mul32\n  get_attr_mul32 (rtx_insn *insn ATTRIBUTE_UNUSED)\n  {\n    attr_type cached_type ATTRIBUTE_UNUSED;\n\n    switch (recog_memoized (insn))\n      {\n      case -1:\n        if (GET_CODE (PATTERN (insn)) !\u003d ASM_INPUT\n            \u0026\u0026 asm_noperands (PATTERN (insn)) \u003c 0)\n          fatal_insn_not_found (insn);\n        /* FALLTHRU */\n        if (((cached_type \u003d get_attr_type (insn)) \u003d\u003d TYPE_SMULXY)\n            || (cached_type \u003d\u003d TYPE_SMLAXY)\n            ... 20 more ...\n            || (cached_type \u003d\u003d TYPE_SMLSLD))\n          {\n            return MUL32_YES;\n          }\n        else\n          {\n            return MUL32_NO;\n          }\n\n      case 424:  /* *mulsi_neg_uxtw */\n      case 423:  /* *muldi_neg */\n      ... 10 more ...\n      case 413:  /* mulsi3 */\n        return MUL32_YES;\n\n      default:\n        return MUL32_NO;\n      }\n  }\n\nTwo things are worth noticing.  The `case -1:\u0027 arm, reached for asm\nstatements whose `type\u0027 is only known at run time, already contains the\nmapping in its original form: evaluate `type\u0027 once, then decide.  Every\nother arm is that same decision, precomputed for one insn code and\nre-emitted.  So the switch is a partially evaluated copy of a function\nthat the file already knows how to write, keyed on the wrong thing.\n\nEmit the mapping directly instead:\n\n  static const unsigned char mul32_from_type[] \u003d {\n    MUL32_NO, MUL32_NO, ..., MUL32_YES, ..., MUL32_NO,\n  };\n\n  attr_mul32\n  get_attr_mul32 (rtx_insn *insn ATTRIBUTE_UNUSED)\n  {\n    return (attr_mul32) mul32_from_type[get_attr_type (insn)];\n  }\n\nThe result is better in three ways.  It is one array read rather than a\nsearch over insn codes, so it does not grow when the port gains\npatterns, only when `type\u0027 gains values.  It has a single control-flow\npath, so the host compiler has nothing to optimise.  And the asm case\nneeds no special handling at all: get_attr_type still issues\nfatal_insn_not_found, and whatever `type\u0027 it returns indexes the same\ntable as any other insn.\n\nThis is correct by construction rather than by testing.  The machine\ndescription says the attribute is a function of `type\u0027, and a table\nindexed by `type\u0027 is that function.  The pass therefore only has to recognise\nthe shape, which it does before optimize_attrs destroys it: a cond, or\nan if_then_else chain, in which every test selects a set of values of\none single other attribute and every value including the default is\nconstant.  check_attr_test has already rewritten (eq_attr \"type\"\n\"a,b,c\") into an ior chain over single values, so a test is matched as a\nboolean combination of eq_attr.\nAnything else, in particular match_test, match_operand, eq_attr_alt and\nattr_flag, falls back to the existing expansion.  An attribute that a\ndefine_insn sets directly is not a function of anything, so that is\nchecked too.\n\nOne extension is needed for the case that motivates all this.\ncortex_a57_neon_type is written as a cond over `type\u0027, except that its\nlast arm tests is_neon_type, which is itself a function of `type\u0027.  So a\ntest of an attribute already known to be a function of the driver counts\nas a test of the driver, and resolving it is a lookup in that\nattribute\u0027s own table.  get_attr_order already supplies the topological\norder that guarantees the dependency is processed first.  Without this,\nthe largest of the transformed attributes is missed.\n\nAcross the tree 107 of 558 attributes qualify, 62 of them on s390, where\nthe driver is `mnemonic\u0027 with over a thousand values.  On aarch64 there\nare eight, and the switches they replace are far from uniform in size:\n\n  attribute                 switch lines   table + getter lines\n  mul32                               39                    466\n  widen_mul64                         37                    466\n  is_mve_type                         75                    466\n  is_neon_type                      4406                    466\n  cortex_a53_advsimd_type           4472                    466\n  cortex_a57_neon_type              5146                    466\n  exynos_m1_neon_type               4708                    466\n  tsv110_neon_type                  4523                    466\n\nA table is always |type| entries, so the three attributes that only a\nfew patterns use get bigger in source.  Emitting a table only when it is\nthe smaller of the two would need a heuristic, and the shape of the\ngenerated code, not its size, is the point, so all of them are\nconverted.  In total insn-attrtab.cc goes from 4.26MB in 92424 lines to\n3.28MB in 72584 lines.  It compiles, together with insn-dfatab.cc and\ninsn-latencytab.cc, in 12% less time, with peak memory 313MB\nagainst 332MB.\n\nRun time is unchanged.  Compiling a 39-file C corpus with\n-mcpu\u003dcortex-a57 takes the same time either way, within a run-to-run\nspread of about 1.5%.  Note that attribute values are not memoised, so a\nderived attribute now runs get_attr_type\u0027s switch rather than its own\ninlined copy of it.  The two are close enough in size that this does not\nshow up.\n\nVerified by compiling that corpus with -mcpu\u003dgeneric, cortex-a57,\ncortex-a53, exynos-m1, tsv110, neoverse-v2 and neoverse-n1: 273\ncompilations, and the assembly does not change.  On the same compiler all 39\nfiles differ between -mcpu\u003dgeneric and -mcpu\u003dcortex-a57, so the pipeline\nmodels these attributes feed are being exercised.\n\nBootstrapped on aarch64-none-linux-gnu.\n\ngcc/ChangeLog:\n\n\t* genattrtab.cc (attr_value): Add enum_index.\n\t(attr_desc): Add num_values, derived_from and derived_table.\n\t(get_attr_value, add_attr_value, find_attr): Maintain them.\n\t(attr_value_index, eq_attr_value_set, simple_enum_attr_p)\n\t(sole_tested_attr, find_derived_attrs, write_derived_attr_get): New\n\tfunctions.\n\t(write_attr_get): Use write_derived_attr_get where it applies.\n\t(main): Call find_derived_attrs before optimize_attrs.\n\nSigned-off-by: Kyrylo Tkachov \u003cktkachov@nvidia.com\u003e\n",
  "tree_diff": [
    {
      "type": "modify",
      "old_id": "a2cf08d53058358aba8624adb8821f598847edff",
      "old_mode": 33188,
      "old_path": "gcc/genattrtab.cc",
      "new_id": "503c50637aa0313c2718f919277af144d3043a77",
      "new_mode": 33188,
      "new_path": "gcc/genattrtab.cc"
    }
  ]
}
