Event date · · Together Computer & Stanford CRFM

RedPajama: an Open Dataset for Training Large Language Models: RedPajama Open Dataset Promotes Transparency in LLM Training

FACT STATEMENT

Submitted in November 2024. The paper releases RedPajama-V1 (an open reproduction of LLaMA training data) and RedPajama-V2 (over 100 trillion tokens of raw web text with quality signals). The datasets have been used in production models such as Snowflake Arctic, Salesforce XGen, and AI2 OLMo. Ablation experiments with a 1.6B parameter model demonstrate how quality signals can effectively filter high-quality subsets.

What happened

This work addresses the lack of transparency and high-quality data in open-source LLM training. RedPajama-V2 provides raw web data and quality signals (e.g., perplexity, language model scores), enabling researchers to filter data autonomously rather than relying on black-box filtering. This lowers the barrier to high-quality data and promotes reproducibility and fair competition in model training. Its 100 trillion token scale and multi-domain coverage provide a solid foundation for subsequent models.

Technical significance

RedPajama-V1 reproduces the LLaMA training data distribution, including sources like CommonCrawl, C4, GitHub, and Books. RedPajama-V2 extracts raw web text from CommonCrawl and computes multiple quality signals (e.g., language model perplexity, URL quality, content duplication). The paper conducts ablation experiments on a 1.6B parameter decoder-only model, finding that filtering based on quality signals significantly improves downstream task performance, with different signal combinations yielding varying effects. Technical limitations include potential bias introduced by quality signals and the substantial computational resources required for large-scale data cleaning.

Industry impact

For AI companies, RedPajama offers an open alternative to commercial datasets, reducing reliance on proprietary data. Data annotation and cleaning tool providers can develop value-added services based on its quality signals. Model training platforms (e.g., Hugging Face, Together AI) can integrate the dataset as a standard option. In the long term, it may drive the industry toward consensus standards for data quality assessment.

Decision value

It is recommended that AI infrastructure companies (e.g., Databricks, Snowflake) adopt RedPajama as one of the default training datasets and offer automated data filtering services based on quality signals. Investment opportunities include open-source projects that use this dataset to train domain-specific models.

What to watch

Attention must be paid to dataset copyright and ethical issues (e.g., whether web content contains personal privacy); whether the effectiveness of quality signals varies with model scale; and whether the community can sustain and update the dataset. If RedPajama becomes a de facto standard, it will accelerate the catch-up of open-source models with closed-source ones.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.