ModelBest released UltraData-Code-L2-Classifier on Hugging Face
ModelBest released UltraData-Code-L2-Classifier, a suite of language-specific file-level scorers for 11 programming languages, on Hugging Face under Apache-2.0. The scorers select algorithmically relevant files from UltraData-Code-L1 to form UltraData-Code-L2, a corpus of approximately 400B tokens retaining 12.23% of L1 files.
China context
- Original name
- 面壁智能
- Outside China
- Open weights · huggingface.co
- Claims
- Company-reported; not yet independently evaluated
- For builders
- Builders outside China can use the Apache-2.0 licensed classifier to filter code datasets for pre-training, which can improve code model performance without building their own selection pipeline.
- For investors
- The release of a code data curation tool under Apache-2.0 signals ModelBest's strategy to build an open ecosystem around its UltraData-Code pipeline, which may attract developer adoption and commercial interest.
ModelBest released UltraData-Code-L2-Classifier on Hugging Face under Apache-2.0. The suite includes language-specific file-level scorers for 11 programming languages. These scorers are applied to UltraData-Code-L1 files to select algorithmically relevant files, forming UltraData-Code-L2, a corpus of approximately 400B tokens that retains 12.23% of L1 files. The model card reports that under controlled 10B-token continual pre-training of a 1B model, training on L2 instead of L1 raises pass@1 on EvalPlus by 7.80 points and on MultiPL-E by 5.13 points, exceeding Stack-Edu by 4.37 and 3.05 points respectively.
The L2 selection framework encodes each file with a 1024-dimensional embedding from Qwen3-Embedding-0.6B, predicts file roles (ALGO, WEB, TOOL, DATA, TEST, CONFIG, EXCLUDE), constructs dual-cue supervision from ALGO labels and language-specific heuristics, scores relevance and quality, and applies calibrated language-specific thresholds. The reported benchmark gains are company_reported and not independently verified in the evidence.
Developers outside China can now use the Apache-2.0 licensed UltraData-Code-L2-Classifier to filter code datasets for pre-training, which can improve code model performance without building their own selection pipeline. The release adds a concrete open-weights tool to the code data curation landscape, directly competing with proprietary data pipelines.
The classifier enables organizations to build higher-quality code training datasets, reducing compute costs by selecting a smaller, more relevant subset of code files. The Apache-2.0 license allows commercial use.
The model card mentions a tech report 'Coming Soon' and links to the UltraData-Code dataset and MiniCPM5 series. Verification questions: whether the reported benchmark gains replicate in independent evaluations, and whether the L2 corpus itself will be released for download.