| layout | blog_detail | ||
|---|---|---|---|
| title | λ² μ΄μ§μ μ΅μ νλ‘ Helionμ μλ νλ κ°μνκΈ° | ||
| author | Ethan Che, Oguz Ulgen, Max Balandat, Jongsok Choi, Jason Ansel | ||
| ext_author | Junghwan Park (λ°μ ν) | ||
| category |
|
||
| date | 2026-02-24 04:00:00 -0800 | ||
| org_title | Accelerating Autotuning in Helion with Bayesian Optimization | ||
| org_link | https://pytorch.org/blog/accelerating-autotuning-in-helion/ |
μ΄μ λΈλ‘κ·Έ κΈμμ μκ°νλ―μ΄, Helionμ μ΅μν PyTorch μ€νμΌμ λ¬Έλ²μΌλ‘ κ³ μ±λ₯ ML 컀λμ μμ±ν μ μκ² ν΄μ£Όλ κ³ μμ€ DSLμ΄λ©°, 볡μ‘ν μ΅μ ν μμ
μ μλ νλ(autotuning) μμ§μ μμν©λλ€. μ΄ μλ νλ(autotuner)λ λΈλ‘ ν¬κΈ°(block size), 루ν μμ(loop order), λ©λͺ¨λ¦¬ μ κ·Ό ν¨ν΄ λ± κ΅¬ν μ νμ§λ‘ μ΄λ£¨μ΄μ§ λ°©λν κ³ μ°¨μ 곡κ°μ νμνμ¬ λμ νλμ¨μ΄μμ μ±λ₯μ κ·Ήλννλ ꡬμ±(configuration)μ μ°Ύμλ
λλ€. κ·Έ κ²°κ³Ό Helionμ torch.compileμ λ¬Όλ‘ , Tritonμ΄λ CuTe DSLλ‘ μ κ΅νκ² μμΌλ‘ μμ±ν 컀λ보λ€λ μλΉν μλ ν₯μμ λ¬μ±ν μ μμ΅λλ€.
As introduced in a previous blog post, Helion is a high-level DSL that empowers developers to write high-performance ML kernels using a familiar PyTorch-like syntax, delegating the complex task of optimization to its autotuning engine. This autotuner explores a vast, high-dimensional space of implementation choicesβblock sizes, loop orders, memory access patternsβto discover configurations that maximize performance on the target hardware. As a result, Helion can achieve significant speedups over torch.compile and even highly-optimized, hand-written kernels in Triton or CuTe DSL.
κ·Έλ¬λ μλ νλμΌλ‘ μ»λ μ±λ₯ ν₯μμλ λκ°κ° λ°λ¦ λλ€. λ°λ‘ κΈ΄ μ€μ μμ μκ°(wall-clock time) μ λλ€. μΌλ°μ μΈ μλ νλ μΈμ μ μμ² κ°μ ν보 ꡬμ±μ νκ°νλ©΄μ 10λΆ μ΄μμ΄ κ±Έλ¦¬λ©°, 볡μ‘ν 컀λμ κ²½μ° μ μκ° λ¨μκΉμ§ λμ΄λκΈ°λ ν©λλ€. μΆμ μ΄ν κΈ΄ μλ νλ μκ°μ μ¬μ©μ λΆλ§μΌλ‘ κΎΈμ€ν μ κΈ°λμ΄ μμΌλ©°, 컀λ κ°λ° μ£ΌκΈ°μμ κ°μ₯ ν° κ³ μΆ© μ€ νλμμ΅λλ€. Helionμ νμ λ¨κ³ μλ₯Ό μ€μ΄λ λ± μλ νλ κ³Όμ μ λ¨μΆν μ μλ μ νμ§λ₯Ό μ 곡νμ§λ§, μ΄λ λ³΄ν΅ μ»€λ μ±λ₯ μ νλ‘ μ΄μ΄μ Έ λ°λμ§νμ§ μμ μ μΆ©μ κ°μν©λλ€.
However, the performance gains from auto-tuning comes with a cost: long wall-clock times. A typical autotuning session can take 10+ minutes, evaluating thousands of candidate configurations, and can even take on the order of hours for complex kernels. Since its launch, long autotuning times have consistently surfaced as a user complaint and one of the biggest pain points in the kernel development cycle. While Helion provides developers options to shorten the auto-tuning process, e.g. by reducing the number of search steps, this typically leads to a loss in kernel performance, forcing an undesirable trade-off.
μ΄λ² κΈμμλ μλ νλ κ²½νμ κ°μ νκΈ° μν μ§ν μ€μΈ λ Έλ ₯μ λ€λ£Ήλλ€. νΉν μ΄λ¬ν λ¬Έμ λ₯Ό ν΄κ²°νκΈ° μν΄ κ°λ°ν μλ‘μ΄ νμ μκ³ λ¦¬μ¦ LFBO Pattern Searchλ₯Ό μκ°ν©λλ€. μ΄ μκ³ λ¦¬μ¦μ λ¨Έμ λ¬λ(ML) κΈ°λ²μ νμ©νμ¬ μλ νλ μμ§μ ν¨μ¨μ λμ λλ€. νμ μκ³ λ¦¬μ¦μ΄ ML λͺ¨λΈμ νμ΅μμΌ ν보 ꡬμ±μ μ§λ₯μ μΌλ‘ κ±Έλ¬λμΌλ‘μ¨ νκ°νλ ν보μ μλ₯Ό ν¬κ² μ€μ λλ€. μ€μν μ μ, μ΄ λͺ¨λΈμ΄ νμ κ³Όμ μμ μμ§λ λ°μ΄ν°λ§ μ¬μ©νλ©° μ¬μ©μκ° λ³λμ λ°μ΄ν°λ₯Ό μ 곡ν νμκ° μλ€λ κ²μ λλ€.
In this blog post, we discuss our ongoing efforts to improve the autotuning experience. In particular, we discuss a new search algorithm LFBO Pattern Search we developed to address these issues, which employs techniques from machine learning (ML) to improve efficiency of the autotuning engine. The search algorithm trains an ML model to intelligently filter candidate configurations, substantially reducing the number of candidates evaluated. Importantly, the model only uses data collected during the search process, and doesn't need the user to provide any additional data.
MLμ νμ©νλ©΄ μ±λ₯μ ν¬μνμ§ μμΌλ©΄μλ μλ νλ μκ°μ μλΉν μ€μΌ μ μμ΅λλ€.
Using ML, we can reduce autotuning time substantially without sacrificing performance:
- λ²€μΉλ§ν¬μ© NVIDIA B200 컀λ λͺ¨μμμ, μλ νλ μκ°μ 36.5% μ€μ΄λ λμμ 컀λ μ§μ° μκ°(latency)μ νκ· 2.6% κ°μ νμ΅λλ€.
- AMD MI350 컀λμμλ μλ νλ μκ°μ 25.9% μ€μ΄λ©΄μ 컀λ μ§μ° μκ°μ 1.7% κ°μ νμ΅λλ€.
- On our set of benchmark NVIDIA B200 kernels, we reduce autotuning time by 36.5% while improving kernel latency by 2.6% on average.
- On AMD MI350 kernels, we reduce autotuning time by 25.9% while improving kernel latency by 1.7%.
μΌλΆ 컀λμμλ κ°μ ν¨κ³Όκ° νΉν λλλ¬μ§λλ€. B200 layer-norm 컀λμμλ μ€μ μμ μκ°μ΄ μ΅λ 50% κ°μνκ³ , B200 Helion FlashAttention 컀λμμλ 컀λ μ§μ° μκ°μ΄ 15% μ΄μ κ°μ λκΈ°λ νμ΅λλ€. μ΄μ²λΌ ν₯μλ μ±λ₯ λλΆμ, μ΄ μκ³ λ¦¬μ¦μ μ΄ κΈμ μ°λ νμ¬ κΈ°λ³Έ νμ μκ³ λ¦¬μ¦μ λλ€.
For some kernels the improvements are especially significant: we see up to a 50% reduction in wall-clock time for B200 layer-norm kernels, and even a >15% improvement in kernel latency for B200 Helion FlashAttention kernels. Due to its enhanced performance, it is the default search algorithm at the time of writing.
μλ νλ μμ§μ 컀λ ꡬμ±λ€μ νμνλ©΄μ κ·Έ μ§μ° μκ°μ λ²€μΉλ§ν¬νκ³ , κ·Έ κ²°κ³Όλ₯Ό λ°νμΌλ‘ λ€μμ λ²€μΉλ§ν¬ν κ΅¬μ± μ§ν©μ κ²°μ ν©λλ€. λ¨μΌ ꡬμ±μ μ»΄νμΌνκ³ μ§μ° μκ°μ μΈ‘μ νλ λ°λ μ μ΄ μ λκ° κ±Έλ¦¬μ§λ§, μλ νλ μμ§μ κ°λ₯ν μ΅κ³ μ μ±λ₯μ μ»κΈ° μν΄ λ³΄ν΅ μμ² κ°μ ꡬμ±μ νμν©λλ€. μ΅μ μ 컀λ ꡬμ±μ μ°Ύλ μΌμ μ€κ³ 곡κ°μ λ΄μ¬λ μ¬λ¬ μμΈμΌλ‘ μΈν΄ μ΄λ €μ΄ μ΅μ ν λ¬Έμ μ λλ€.
The autotuning engine searches through kernel configurations, benchmarking their latency and using the outcomes to determine the next set of configs to benchmark. While compiling and measuring the latency of a single configuration takes on the order of seconds, the autotuning engine typically searches through thousands of configurations to achieve the best possible performance. Finding the optimal kernel configuration is a challenging optimization problem due to several factors inherent to the design space:
- κ³ μ°¨μμ μ‘°ν© κ³΅κ°(High-Dimensional, Combinatorial Space): λΈλ‘ ν¬κΈ°, μΈλ‘€ ν©ν°(unroll factor) λ± κ°λ₯ν λͺ¨λ μ‘°ν©μ 곡κ°μ κ³ μ°¨μμ΄λ©° λ°©λν©λλ€. LayerNormμ²λΌ λ¨μν 컀λμ‘°μ°¨ 8μ²μ‘°(8 quadrillion, 10^16) κ°κ° λλ ꡬμ±μ κ°μ§ μ μμ΅λλ€. λ€λ§ νμ 곡κ°μ κ±°λνμ§λ§, μ’μ μ±λ₯μ λ΄λ ꡬμ±μ κ·Ήν μΌλΆμ λΆκ³Όν©λλ€.
- κΈ΄ μ»΄νμΌ μκ°(Long Compile Times): μ΄λ€ 컀λ ꡬμ±μ μ»΄νμΌμ μλΉν μκ°μ΄ κ±Έλ €, μλ νλ κ³Όμ μ μ€μ μμ μκ°μ λΆνμνκ² λ립λλ€.
- κ΅¬μ± μ€λ₯μ νμμμ(Config Errors and Timeouts): νμ 곡κ°μλ μ»΄νμΌ μ€λ₯κ° λκ±°λ, λΆμ νν κ²°κ³Όλ₯Ό λ΄κ±°λ, μ»΄νμΌμ λ무 μ€λ 걸리λ ꡬμ±λ ν¬ν¨λ μ μμ΅λλ€.
- High-Dimensional, Combinatorial Space: The space of all possible combinations of block sizes, unroll factors, etc. is high-dimensional and vast. Even a simple kernel like LayerNorm has more than 8 quadrillion (10^16) possible configurations. However, while the search space is large, only a small fraction of configs have good performance.
- Long Compile Times: Certain kernel configurations can take a significant amount of time to compile, unnecessarily extending the autotuning process's wall-clock time.
- Config Errors and Timeouts: The search space can also include configs that have compilation errors, produce inaccurate results, or take too long to compile.
μ΄μ μ κΈ°λ³Έ νμ μ λ΅(Pattern Search)μ μ¬λ¬ κ°μ μ λ§ν ꡬμ±('νμ μ¬λ³Έ(search copies)')μμ μμνμ¬, λ¨μΌ λ§€κ°λ³μλ₯Ό λ³νν λͺ¨λ κ²½μ°λ₯Ό λΉ μ§μμ΄ νκ°νλ λ°©μμΌλ‘ μ΄μ ꡬμ±λ€μ νμν©λλ€. μ² μ νκΈ΄ νμ§λ§ μ΄ λ°©μμ λΉν¨μ¨μ μ λλ€. μ΄μ ꡬμ±μ λλΆλΆμ μ±λ₯μ μ ν κ°μ νμ§ λͺ»νλλ°λ νλνλ μ»΄νμΌνκ³ λ²€μΉλ§ν¬νκΈ° λλ¬Έμ λλ€. κ²λ€κ° μ΄λμ λ¨μΌ λ§€κ°λ³μ λ³κ²½μΌλ‘ μ ννλ©΄, κ³ μ°¨μ νμ 곡κ°μ λΉ λ₯΄κ² κ°λ‘μ§λ₯΄λ λ₯λ ₯μ΄ λ¨μ΄μ§λλ€.
The previous default search strategy (Pattern Search) starts from multiple promising configurations ('search copies') and explores neighboring configs by exhaustively evaluating all single-parameter perturbations. While thorough, this approach is inefficient: the vast majority of neighbors offer no performance improvement, yet each is compiled and benchmarked. Furthermore, restricting moves to single-parameter changes limits the algorithm's ability to traverse the high-dimensional search space quickly.
κ°λ₯λ μλ λ² μ΄μ§μ μ΅μ ν ν¨ν΄ νμ / Likelihood-Free Bayesian Optimization Pattern Search
μ΄λ¬ν λΉν¨μ¨μ ν΄κ²°νκΈ° μν΄, λ€μμ νκ°ν μ μ μ§λ₯μ μΌλ‘ μ ννλ νλ₯ μ λ리 λͺ¨λΈ(surrogate model, μ: κ°μ°μμ νλ‘μΈμ€(Gaussian Process))μ νμ©νλ λ¨Έμ λ¬λμ ν λΆμΌμΈ λ² μ΄μ§μ μ΅μ ν(Bayesian Optimization)μμ μκ°μ μ»μμ΅λλ€(botorchλ Ax κ°μ λΌμ΄λΈλ¬λ¦¬μμ μ¬μ©ν μ μμ΅λλ€). μΆκ°λλ μ€μ μμ μκ°μ μ΅μννκΈ° μν΄, λ κ°λ²Όμ΄ λΆλ₯(classification) λͺ¨λΈμ λ리 λͺ¨λΈλ‘ μ¬μ©νλ κ°λ₯λ μλ λ² μ΄μ§μ μ΅μ ν(Likelihood-Free Bayesian Optimization)(LFBO)λ₯Ό λμ νμ΅λλ€. Pattern Searchμ μ§μ νμ ν΄λ¦¬μ€ν±κ³Ό LFBO λΆλ₯κΈ° λͺ¨λΈμ κ²°ν©νμ¬, μμ νμ λμ κ°μ₯ μ λ§ν νλ³΄λ§ κ±Έλ¬μ λ²€μΉλ§ν¬ν©λλ€.
To address these inefficiencies, we take inspiration from Bayesian Optimization, a sub-domain of machine learning which utilizes a probabilistic surrogate model (e.g. a Gaussian Process) to intelligently select which points to evaluate next (available in libraries such as botorch and Ax). To minimize additional wall-clock time, we adapt Likelihood-Free Bayesian Optimization (LFBO), which uses a lighter-weight classification model as a surrogate. We combine the local search heuristic of Pattern Search with the LFBO classifier model to filter only the most promising candidates to benchmark, instead of exhaustive search.
LFBOPatternSearch μκ³ λ¦¬μ¦μ λ€μκ³Ό κ°μ΅λλ€.
The LFBOPatternSearch algorithm is as follows:
- PatternSearchμ λ§μ°¬κ°μ§λ‘, λ¨Όμ 무μμλ‘ μμ±ν κ΅¬μ± μ§ν©μ λ²€μΉλ§ν¬νμ¬ κ°μ₯ μ λ§ν μμμ ꡬμ±('νμ μ¬λ³Έ(search copies)')μ μλ³ν©λλ€.
- νμ μ¬λ³ΈμΌλ‘λΆν° μ¬λ¬ λ§€κ°λ³μμ κ±Έμ³ λ¬΄μμ λ³νμ κ°ν΄ ν보λ₯Ό μμ±νλ©°, PatternSearchλ³΄λ€ λ λκ² νμν©λλ€.
- μ§κΈκΉμ§ μμ§ν μ§μ° μκ° λ°μ΄ν°λ‘ λΆλ₯ λͺ¨λΈ(λλ€ ν¬λ μ€νΈ(RandomForest))μ νμ΅μν΅λλ€. μ§μ° μκ°μ μ§μ μμΈ‘νλ λμ , ν΄λΉ ꡬμ±μ΄ μ§μ° μκ° κΈ°μ€ μμ 10%μ λλμ§λ₯Ό λνλ΄λ μ΄μ§ λ μ΄λΈμ μμΈ‘ν©λλ€.
- ML λͺ¨λΈμ μμΈ‘μ λ°νμΌλ‘ ν보μ μμλ₯Ό λ§€κΉλλ€. μΌλ°μ μΈ LFBOμ λ¬λ¦¬, νμμ μ₯λ €νκΈ° μν΄ μ΄λ―Έ μμκ° λ§€κ²¨μ§ ν보μμ μ μ¬λμ λν ν¨λν°λ μΆκ°ν©λλ€.
- κ·Έμ€ μμ 10%λ₯Ό μ νν΄ μ»΄νμΌνκ³ λ²€μΉλ§ν¬ν©λλ€. μ±λ₯μ΄ κ°μ₯ μ’μ ꡬμ±μ λ°νμΌλ‘ νμ μ¬λ³Έμ κ°±μ νκ³ , μΈ‘μ λ μ§μ° μκ°μ λ°μ΄ν°μ μ μΆκ°ν©λλ€.
- Similar to PatternSearch, we first benchmark a set of randomly generated configs, and identify a small set of the most promising configurations ('search copies').
- We generate candidates from the search copies, by making random perturbations across multiple parameters, exploring more widely than PatternSearch.
- We train a classification model (RandomForest) on latency data collected so far. Instead of predicting latency directly, we predict a binary label indicating whether the config is in the top 10% in terms of latency.
- We rank the candidates based on ML model predictions. Unlike typical LFBO, we also add a penalty for similarity to previously ranked candidates to encourage exploration.
- We select the top 10% of them to compile and benchmark. We update the search copies based on the best performing configs and add the latencies to the dataset.
κ΄μ°°λ μ€μ μμ μκ°κ³Ό μ§μ° μκ° κ°μ μ λ¬μ±νλ λ° κ²°μ μ μΈ μν μ ν λͺ κ°μ§ ν΅μ¬ μ€κ³ κ²°μ μ μ΄ν΄λ³΄κ² μ΅λλ€.
We discuss some key design decisions, which are critical for achieving the improvements in wall-clock time and latency we observed:
λΆλ₯ λ νκ·(Classification vs Regression): λͺ¨λΈμ΄ μ§μ° μκ°μ μ§μ μμΈ‘νλλ‘ νμ΅μν€λ νκ·(regression) κΈ°λ° λ°©λ²μ μμ€ν /μ»΄νμΌλ¬ μ°κ΅¬μμ λΉμ© λͺ¨λΈλ§(cost modeling)μ μ¬μ€μ νμ€μ λλ€. κ·Έλ¬λ μ’λ λμλ λͺ¨λ ꡬμ±μ μ§μ° μκ°μ νμ΅νλ €κ³ μ μ°λ λμ , λΆλ₯ κΈ°λ° μ κ·Όμ΄ κ°μ₯ μ±λ₯μ΄ μ’μ ꡬμ±μ λͺ¨λΈμ μλμ λ μ μ§μ€μν¨λ€λ μ μ νμΈνμ΅λλ€. λμ§Έλ‘, λΆλ₯ μμ€(classification loss)μ μ€λ₯κ° λκ±°λ μ»΄νμΌ νμμμμ΄ λ°μνλ ꡬμ±(μ΄λ€μλ μμ λ μ΄λΈμ΄ λΆμ¬λ©λλ€)μ νΌνλλ‘ λͺ¨λΈμ΄ νμ΅νκ² ν΄μ€λλ€. λ°λ©΄ μ΄λ° ꡬμ±λ€μ μ ν¨ν μ§μ° μκ° λ°μ΄ν°κ° μμΌλ―λ‘ νκ· κΈ°λ° μ κ·Όμ΄ νμ΅ν κ±°λ¦¬κ° μμ΅λλ€.
Classification vs Regression: Regression-based methods, i.e. training the model to predict latency directly, is the de-facto approach for cost modeling in systems / compiler research. However, we find that a classification-based approach better focuses model capacity on the most performant configs instead of trying to learn the latency of all configs, good or bad. Second, the classification loss enables the model to learn to avoid configs that error out or suffer compile timeouts (as these are assigned negative labels). However, these points do not have any valid latency data for a regression-based approach to learn from.
λ€μμ± μ₯λ €(Encouraging Diversity): λ³΄ν΅ κ΅¬μ±λ€μ λ³λ ¬ μ¬μ μ»΄νμΌ(pre-compilation)μ νμ©νκΈ° μν΄ λ°°μΉ(batch) λ¨μλ‘ μ»΄νμΌλ©λλ€. λλ€ ν¬λ μ€νΈ λΆλ₯κΈ°λ μλ‘ λμ³ μλ μ μ¬ν ꡬμ±μ λ°λ³΅μ μΌλ‘ μ νν μ μλλ°, μ΄λ μλ‘μ΄ μ 보λ₯Ό κ±°μ μ£Όμ§ λͺ»νλ μ€λ³΅ μνμ λ°°μΉ μμ°μ λλΉνκ² λ§λλλ€. μ΄λ₯Ό μννκΈ° μν΄, λλ€ ν¬λ μ€νΈ λͺ¨λΈμ 리ν λ Έλ λμ μΆν(leaf node co-occurrence)μ λ°νμΌλ‘ μ μ¬λ μ μλ₯Ό κ³μ°νκ³ , μ΄λ―Έ μμκ° λ§€κ²¨μ§ κ΅¬μ±κ³Όμ μ μ¬λμ ν¨λν°λ₯Ό λΆμ¬ν©λλ€.
Encouraging Diversity: Typically configs are compiled in batch, to take advantage of parallelized pre-compilation. The Random Forest classifier may repeatedly select similar configurations that cluster, which can waste the batch budget on redundant samples that provide little new information. To mitigate this, we compute a similarity score based on leaf node co-occurrence from the Random Forest model, and penalize similarity to previously ranked configs.
LFBO Pattern Searchμ λμμ μ΄ν΄λ³΄λ©΄, 컀λΒ·νλμ¨μ΄ μ’ λ₯Β·νμ(shape) μ λ°μ κ±ΈμΉ μ±λ₯ κ°μ μ΄ λ μ μ νκ° νμλ‘ λ λμ λ°νμμ κ°μ§ ꡬμ±μ μ°Ύμλ΄λ λ₯λ ₯μμ λΉλ‘―λλ€λ κ²μ μ€μ λ‘ νμΈν μ μμ΅λλ€. μλλ B200 layer-norm 컀λμ λν μλ νλ νΈλ μ΄μ€ μμλ‘, μκ° κ²½κ³Όμ λ°λΌ μλ νλκ° μ»μ μ΅μ ꡬμ±μ μ§μ° μκ°μ 보μ¬μ€λλ€. LFBOκ° μλ νλμ λ μΌμ°(μ½ 9λΆ λμ μ½ 5λΆ) μλ£ν λΏλ§ μλλΌ, Pattern Searchμ λΉν΄ ν¨μ¬ ν° νμ μ±λ₯ λμ½μ μ΄λ£¨λ©° λ λμ ꡬμ±μ λ λΉ λ₯΄κ² μ°ΎμλΈλ€λ κ²μ λ³Ό μ μμ΅λλ€.
When we investigate the behavior of LFBO Pattern Search, we see indeed that improvements in performance across kernels, hardware types, and shapes are due its ability to find configurations with better runtime using fewer evaluations. Below is a plot of example auto-tuning traces for a B200 layer-norm kernel, displaying the latency of the best configuration obtained by the autotuner over time. We see not only that LFBO completes auto-tuning earlier (~5 min instead of ~9 min), it finds better configurations faster with much larger jumps in performance compared to Pattern Search.
LFBOκ° μ΄λ₯Ό λ¬μ±νλ λ°©μμ Pattern Searchλ³΄λ€ λ λκ² νμνλ κ²μ λλ€. μλλ λμΌν B200 layer-norm 컀λμ λν΄ LFBO Pattern Searchμ Pattern Searchκ° μνλ§ν ꡬμ±μ 보μ¬μ£Όλ κ·Έλνλ‘, (ꡬμ±μ΄ κ³ μ°¨μμ΄λ―λ‘) μκ°νλ₯Ό μν΄ μ£Όμ±λΆ λΆμ(Principal Component Analysis, PCA)μ μ μ©νμ΅λλ€. LFBO Pattern Searchλ ꡬμ±μ μ λ°λ μ λλ μλ§νΌ νκ°νλ©΄μλ, κ·Έ μνλ§ν ꡬμ±λ€μ΄ Pattern Searchμ κ²λ³΄λ€ λ λκ² νΌμ Έ μμμ λ³Ό μ μμ΅λλ€. Pattern Searchμ ꡬμ±λ€μ λ¨μΌ λ§€κ°λ³μ λ³νλ§ νκΈ° λλ¬Έμ νκ³³μ μ¬νκ² λμ³ μμ΅λλ€. λΆλ₯κΈ°μ μλ΄λ₯Ό λ°μ LFBO Pattern Searchλ λ ν¬λ©΄μλ λ νμ νλ λμ½μ ν μ μμ΅λλ€.
We see that LFBO accomplishes this by exploring more widely than Pattern Search. Below is a plot of configs sampled by LFBO Pattern Search and Pattern Search for the same B200 layer-norm kernel, where we apply Principal Component Analysis (PCA) for visualization (as configs are high-dimensional). We see that while LFBO Pattern Search evaluates less than half of the number of configs, its sampled configs are more spread out than Pattern Search's which are highly clumped together due to Pattern Search making only single parameter perturbations. Guided by the classifier, the LFBO Pattern Search is able to make larger, but more targeted jumps.
λ§μ§λ§μΌλ‘, λ€λ₯Έ λ리 λͺ¨λΈλ€, νΉν λλ€ ν¬λ μ€νΈ, κ·ΈλλμΈνΈ λΆμ€ν νΈλ¦¬(Gradient-Boosting Tree), λ€μΈ΅ νΌμ νΈλ‘ (Multi-Layer Perceptron, MLP)μ μ¬μ©νλ νκ· κΈ°λ° μ κ·Όλ€κ³Ό ν¨κ» μ λΈλ μ΄μ (ablation) μ€νμ μννμ΅λλ€. μ΄λ (PatternSearchμμ μμ§ν) μλ νλ λ‘κ·Έ λ°μ΄ν°μ μ μ¬μ©νμ΅λλ€. μλ νλ μ±λ₯κ³Ό κ°μ₯ μ§μ μ μΌλ‘ μ°κ΄λ μ§νμΈ, λ리 λͺ¨λΈμ μ¬μ©ν΄ λ€μ ν보 λ°°μΉλ₯Ό κ±Έλ¬λΌ λ κΈ°λλλ 컀λ μ§μ° μκ° κ°μ μΉλ₯Ό κ³μ°ν©λλ€. μλμμλ λ리 λͺ¨λΈμ΄ μ ννλλ‘ νμ©λ ν보 λΉμ¨ λλΉ κΈ°λ κ°μ μΉ(μ§μ° μκ°μ μλμ % κ°μ )λ₯Ό κ·Έλνλ‘ λνλμ΅λλ€. LFBO κΈ°λ° λ°©λ²μ΄ κ°μ₯ ν° κΈ°λ κ°μ μΉλ₯Ό μ 곡νλ©°, λ€μμ±μ κ³ λ €ν μ νμ΄ μλ―Έ μλ κ°μ μ λνλ€λ κ²μ νμΈνμ΅λλ€. νΉν ν보μ 10%λ§ μ νν μ μμ λ, νκ· κΈ°λ° λ°©λ²μ λ¨μ 무μμ μ νκ³Ό λλ±νκ±°λ μ€νλ € λ λμ μ±λ₯μ 보μμ΅λλ€. νκ·κ° νμ μμ λ§€κΈ°κΈ°(ranking) μ±λ₯κ³Ό μΌμΉνμ§λ μκΈ° λλ¬Έμ λλ€.
Finally, we perform an ablation with other surrogate models, in particular regression-based approaches involving a Random Forest, Gradient-Boosting Tree, and Multi-Layer Perceptron (MLP), using a dataset of autotuner logs (collected from PatternSearch). We compute a metric that is most directly correlated with autotuner performance: the expected improvement in kernel latency when using the surrogate to filter the next batch of candidates. Below we plot the expected improvement (in terms of relative % improvements in latency) compared to the percent of candidates the surrogate is allowed to select. We find that the LFBO-based methods deliver the largest expected improvement, with meaningful improvements from diverse selection. Notably, when we only can select 10% of candidates, the regression-based methods perform equivalent or even worse than simple random selection, as regression is not always aligned with ranking performance.
μ΄λ² κΈμμλ λ¨Έμ λ¬λ(ML)μ΄ μ΄λ»κ² μλ νλ μμ§μ κ°μνκ³ Helionμμμ 컀λ μμ± κ²½νμ κ°μ ν μ μλμ§λ₯Ό 보μ¬μ£Όμμ΅λλ€. νμ κ³Όμ μμ μμ§ν μ§μ° μκ° λ°μ΄ν°λ₯Ό νμ©νλ©΄, μλ νλλ₯Ό λ μ λ§ν ꡬμ±μ μ§μ€μμΌ μκ°μ μ μ½νκ³ λ λΉ λ₯Έ 컀λ ꡬμ±μ λ°κ²¬ν μ μμ΅λλ€. μ°λ¦¬λ κ°ν νμ΅(reinforcement learning, RL)κ³Ό λκ·λͺ¨ μΈμ΄ λͺ¨λΈ(large language models, LLMs)μ κΈ°λ²μ ν¬ν¨νμ¬, μλ νλλ₯Ό κ°ννκΈ° μν μΆκ°μ μΈ ML κΈ°λ²μ μ μ©νλ λ° μ κ·Ήμ μΈ κ΄μ¬μ κ°κ³ μμΌλ©°, μ΄λ€ ννμ κΈ°μ¬λ νμν©λλ€.
In this blog post, we illustrate how machine learning (ML) can accelerate the autotuning engine and improve the kernel authoring experience in Helion. By using the latency data collected during the search process, we can focus the autotuner on more promising configurations, saving time and discovering faster kernel configs. We are actively interested in applying additional ML techniques to enhance the auto-tuner, including methods from reinforcement learning (RL) and large language models (LLMs), and welcome any contributions.





