Event date · · ModelBest

ModelBest released UltraData-Code-L2-Classifier on Hugging Face

ModelBest 面壁智能Chinese AIOpen weights
FACT STATEMENT

ModelBest released UltraData-Code-L2-Classifier, a suite of language-specific file-level scorers for 11 programming languages, on Hugging Face under Apache-2.0. The scorers select algorithmically relevant files from UltraData-Code-L1 to form UltraData-Code-L2, a corpus of approximately 400B tokens retaining 12.23% of L1 files.

China context

Original name
面壁智能
Outside China
Open weights · huggingface.co
Claims
Company-reported; not yet independently evaluated
For builders
Builders outside China can use the Apache-2.0 licensed classifier to filter code datasets for pre-training, which can improve code model performance without building their own selection pipeline.
For investors
The release of a code data curation tool under Apache-2.0 signals ModelBest's strategy to build an open ecosystem around its UltraData-Code pipeline, which may attract developer adoption and commercial interest.
What happened

ModelBest released UltraData-Code-L2-Classifier on Hugging Face under Apache-2.0. The suite includes language-specific file-level scorers for 11 programming languages. These scorers are applied to UltraData-Code-L1 files to select algorithmically relevant files, forming UltraData-Code-L2, a corpus of approximately 400B tokens that retains 12.23% of L1 files. The model card reports that under controlled 10B-token continual pre-training of a 1B model, training on L2 instead of L1 raises pass@1 on EvalPlus by 7.80 points and on MultiPL-E by 5.13 points, exceeding Stack-Edu by 4.37 and 3.05 points respectively.

Technical significance

The L2 selection framework encodes each file with a 1024-dimensional embedding from Qwen3-Embedding-0.6B, predicts file roles (ALGO, WEB, TOOL, DATA, TEST, CONFIG, EXCLUDE), constructs dual-cue supervision from ALGO labels and language-specific heuristics, scores relevance and quality, and applies calibrated language-specific thresholds. The reported benchmark gains are company_reported and not independently verified in the evidence.

Industry impact

Developers outside China can now use the Apache-2.0 licensed UltraData-Code-L2-Classifier to filter code datasets for pre-training, which can improve code model performance without building their own selection pipeline. The release adds a concrete open-weights tool to the code data curation landscape, directly competing with proprietary data pipelines.

Decision value

The classifier enables organizations to build higher-quality code training datasets, reducing compute costs by selecting a smaller, more relevant subset of code files. The Apache-2.0 license allows commercial use.

What to watch

The model card mentions a tech report 'Coming Soon' and links to the UltraData-Code dataset and MiniCPM5 series. Verification questions: whether the reported benchmark gains replicate in independent evaluations, and whether the L2 corpus itself will be released for download.

Latest in Chinese AI

  1. MiniMaxMiniMax open-sources MiniMax-Code-MiniApps repository for community-built plugins
  2. DeepSeekDeepSeek open-sources dsh-libreoffice-kit 0.1.0 for font-friendly Office conversion and rendering in Node.js
  3. DeepSeekDeepSeek open-sources DeepEP-Ascend and DeepGEMM-Ascend for Huawei Ascend NPUs
  4. Shanghai AI LaboratoryShanghai AI Laboratory open-sources AdvancedMathBench for proof generation and verification
  5. Shanghai AI LaboratoryInternLM released a Qwen3-based model that grades mathematical proofs

All China AI Events

AIGC Newsletter

China AI, with sources and context.

Analysis of Chinese AI models, companies and policy, and what you can use outside China.