Event date · · arXiv

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

FACT STATEMENT

An arXiv paper (2609.05381v1) published on 2026-09-04 audits 22 frontier language models on 12 molecular property regression benchmarks for verbatim retrieval of published values. It finds retrieval is widespread but benchmark-specific: on five datasets more than 50% of LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. Experiments at two reasoning levels show reasoning changes retrieval, with the same experiments flagged 89% more often at the higher reasoning level than at the lowest. Testing interruption of retrieval in the most contaminated cases shows the strongest models still recognize a combination of transformed SMILES strings and original labels. Suppressing retrieval moves prediction errors of different models closer together in relative terms, while differing use of verbatim retrieval spreads them apart, indicating general predictive capability is not determined solely by retrieval.

What happened

The paper audits 22 frontier LLMs on 12 molecular property regression benchmarks and finds verbatim retrieval of published values is widespread but benchmark-specific. On five datasets, over 50% of models show verbatim retrieval; on others, retrieval appears only in isolated cells. Reasoning level affects retrieval: the same experiments are flagged 89% more often at higher reasoning than at the lowest. Interrupting retrieval in contaminated cases shows strongest models still recognize transformed SMILES plus original labels. Suppressing retrieval brings model errors closer together, while differing retrieval use spreads them apart, suggesting general predictive capability is not solely due to retrieval.

Technical significance

The study distinguishes prediction from retrieval by auditing digit-level verbatim reproduction of published values. It uses two reasoning levels and finds higher reasoning increases retrieval flags by 89%. Interruption tests with transformed SMILES strings and original labels show models can still retrieve, indicating robust memorization. Suppressing retrieval reduces inter-model error variance, implying retrieval contributes to performance differences. This suggests benchmark contamination via memorization is a significant confounder in molecular property evaluation.

Industry impact

For AI in scientific applications, this highlights a risk that LLM performance on molecular benchmarks may be inflated by memorization rather than generalization. Companies developing or using LLMs for drug discovery or materials science should audit for retrieval contamination. The finding that reasoning level changes retrieval suggests deployment settings (e.g., chain-of-thought) may alter contamination risk. Benchmark developers may need to create retrieval-resistant evaluation protocols.

Decision value

For enterprises using LLMs in molecular property prediction, this paper provides a method to assess whether model outputs are genuine predictions or retrieved values. It can inform model selection and evaluation practices, reducing risk of overestimating model capability. The interruption technique may be adapted for internal audits. The finding that suppressing retrieval aligns model errors suggests that retrieval contributes to apparent model differentiation, which could affect competitive benchmarking.

What to watch

Expect increased scrutiny of LLM benchmark contamination in scientific domains. Future work may develop retrieval detection methods and contamination-resistant benchmarks. Model developers may implement memorization mitigation techniques. The observed reasoning-level effect could lead to guidelines on when to use high-reasoning modes in scientific evaluations. Watch for follow-up studies on other property types and model families.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.