Benchmark Data

30 production operators with 440 hot shapes, importance weights from 1,303 production profiles, roofline bounds across 2 hardware platforms, and deployed kernel baselines.

30 operators
ID ▴▾ Operator ▴▾ dtype ▴▾ Shapes ▴▾ Importance ▴▾
030 unified_attention bf16 25
39.5%
009 fused_moe bf16 23
11.4%
002 block_scaled_mm fp8_e4m3 24
9.3%
006 fp8_blockscale_fused_moe fp8_e4m3 6
5.1%
024 paged_attention_decode bf16 8
4.4%
026 reshape_and_cache bf16 8
4.4%
029 topk_filter fp32 8
3.5%
013 gated_delta_rule_update bf16 17
3.5%
011 fused_qkv_rope fp16 6
3.4%
027 rms_norm bf16 56
2.8%
018 mla_decode_attention bf16 3
2.3%
001 attention_forward bf16 21
2.2%
003 causal_conv1d bf16 9
2.2%
022 moe_topk_gating_softmax fp32 10
2.1%
010 fused_qk_rmsnorm fp16 4
1.9%
028 silu_and_mul bf16 37
1.8%
005 chunk_gated_delta_rule_state bf16 25
1.6%
008 fused_add_rms_norm bf16 19
1.1%
019 moe_align_block_size int32 6
1.1%
023 mrope bf16 5
1.0%
004 chunk_delta_rule_output bf16 16
0.8%
007 fp8_dynamic_per_token_quant fp8_e4m3 20
0.8%
017 linear_sigmoid_mul bf16 9
0.7%
025 per_token_group_quant_fp8 fp8_e4m3 19
0.5%
014 gated_rms_norm bf16 7
0.5%
015 l2_norm bf16 13
0.5%
020 moe_count_and_sort int32 5
0.5%
012 fused_rmsnorm_quant fp8_e4m3 11
0.3%
021 moe_sum_reduce bf16 7
0.3%
016 layer_norm bf16 13
0.1%

Operator Importance Distribution

Importance score reflects trace-derived GPU time share across production workloads