Event date · · Alibaba

QwenLM releases D2K-Bench, an open-source GPU-kernel co-design benchmark

Alibaba 阿里巴巴Chinese AIOpen weights
FACT STATEMENT

QwenLM (Alibaba) released D2K-Bench, an open-source benchmark with 26 GPU-kernel co-design tasks, reference implementations, per-task evaluation environments, and a design/code capability scorer. The repository accompanies the paper 'D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?' and is available on GitHub.

China context

Original name
D2K-Bench
Outside China
Open weights · github.com
Claims
Company-reported; not yet independently evaluated
For builders
Developers outside China can use the open-source benchmark to evaluate their own LLM agents on GPU kernel generation tasks without needing access to proprietary Chinese infrastructure.
For investors
The release signals Alibaba's continued investment in developer tools and open-source AI, which may influence competitive dynamics in the AI coding assistant market.
What happened

D2K-Bench is a GPU-kernel co-design benchmark containing 26 tasks with reference implementations, per-task evaluation environments, packed task tables, example agent workspaces, and a design/code capability scorer. Each task includes expert design guidance split into three cumulative levels: L1 high-level algorithmic insights, L2 dataflow design, and L3 low-level optimization tricks. The capability scorer reports JS design and JS impl, each the arithmetic mean of the three level scores (L1 + L2 + L3) / 3, using shared 0/25/50/75/100 anchors. The repository includes dataset, tasks, example workspace, docker, script, and capability directories, along with setup and validation scripts.

Technical significance

The benchmark provides a structured evaluation of LLM agents' ability to translate expert design guidance into efficient GPU kernels. It includes a three-level design contract (algorithmic insights, dataflow design, low-level optimization) and a scoring system that averages per-level scores. The repository supports offline validation, Docker-based GPU runtime, and packed Parquet task tables for frontier solvers (Triton, CUDA, CUTLASS).

Industry impact

Developers and researchers evaluating LLM agents for GPU kernel generation gain a standardized, open-source benchmark with 26 tasks and a reproducible scoring methodology, reducing the cost and effort of building custom evaluation pipelines.

Decision value

For organizations developing or using LLM agents for code generation, D2K-Bench offers a ready-made evaluation framework to assess and compare agent performance on GPU kernel design tasks, informing tool selection and R&D priorities.

What to watch

Observable next signals include adoption of D2K-Bench in academic papers or leaderboards, community contributions to the repository, and independent evaluations of LLM agents using the benchmark. Verification question: Has any independent evaluator published results using D2K-Bench?

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  3. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  4. Shanghai AI LaboratoryShanghai AI Laboratory open-sources AdvancedMathBench for proof generation and verification
  5. Shanghai AI LaboratoryInternLM released a Qwen3-based model that grades mathematical proofs

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.