QwenLM releases D2K-Bench, an open-source GPU-kernel co-design benchmark
QwenLM (Alibaba) released D2K-Bench, an open-source benchmark with 26 GPU-kernel co-design tasks, reference implementations, per-task evaluation environments, and a design/code capability scorer. The repository accompanies the paper 'D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?' and is available on GitHub.
China context
- Original name
- D2K-Bench
- Outside China
- Open weights · github.com
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Developers outside China can use the open-source benchmark to evaluate their own LLM agents on GPU kernel generation tasks without needing access to proprietary Chinese infrastructure.
- For investors
- The release signals Alibaba's continued investment in developer tools and open-source AI, which may influence competitive dynamics in the AI coding assistant market.
D2K-Bench is a GPU-kernel co-design benchmark containing 26 tasks with reference implementations, per-task evaluation environments, packed task tables, example agent workspaces, and a design/code capability scorer. Each task includes expert design guidance split into three cumulative levels: L1 high-level algorithmic insights, L2 dataflow design, and L3 low-level optimization tricks. The capability scorer reports JS design and JS impl, each the arithmetic mean of the three level scores (L1 + L2 + L3) / 3, using shared 0/25/50/75/100 anchors. The repository includes dataset, tasks, example workspace, docker, script, and capability directories, along with setup and validation scripts.
The benchmark provides a structured evaluation of LLM agents' ability to translate expert design guidance into efficient GPU kernels. It includes a three-level design contract (algorithmic insights, dataflow design, low-level optimization) and a scoring system that averages per-level scores. The repository supports offline validation, Docker-based GPU runtime, and packed Parquet task tables for frontier solvers (Triton, CUDA, CUTLASS).
Developers and researchers evaluating LLM agents for GPU kernel generation gain a standardized, open-source benchmark with 26 tasks and a reproducible scoring methodology, reducing the cost and effort of building custom evaluation pipelines.
For organizations developing or using LLM agents for code generation, D2K-Bench offers a ready-made evaluation framework to assess and compare agent performance on GPU kernel design tasks, informing tool selection and R&D priorities.
Observable next signals include adoption of D2K-Bench in academic papers or leaderboards, community contributions to the repository, and independent evaluations of LLM agents using the benchmark. Verification question: Has any independent evaluator published results using D2K-Bench?