Can LLMs Discover Scientific Laws in Real and Parallel Worlds?
A benchmark called SCILAWS-BENCH is introduced for scientific law discovery, built from published research and real scientific data. It comprises 118 problems drawn from 381 scientific papers, covering 291 candidate laws and roughly 8M real data points across six scientific disciplines. Each problem is instantiated in two settings: SCILAWS-REAL asks models to propose laws from fixed real observations and evaluates held-out predictive fit and scientific validity; SCILAWS-PARALLEL asks models to actively query residual-calibrated worlds and recover synthesized hidden laws.
Researchers introduce SCILAWS-BENCH, a benchmark for evaluating whether large language models can discover scientific laws. The benchmark includes 118 problems from 381 papers, 291 candidate laws, and about 8 million real data points across six disciplines. It has two settings: SCILAWS-REAL, where models propose laws from fixed real observations and are evaluated on predictive fit and scientific validity, and SCILAWS-PARALLEL, where models actively query residual-calibrated worlds to recover synthesized hidden laws.
The benchmark addresses limitations of existing evaluations that simplify discovery or reuse published targets familiar to LLMs. SCILAWS-PARALLEL uses residual-calibrated worlds, suggesting a controlled environment where models can interactively test hypotheses. The scale (8M data points, 291 laws) indicates a rigorous test of generalization and scientific reasoning.
This work signals growing interest in using LLMs for scientific discovery, a key application area for AI. A standardized benchmark could drive competition among model developers to improve scientific reasoning capabilities, potentially impacting research productivity across disciplines.
The benchmark could become a standard for evaluating AI in science, influencing procurement decisions in research institutions and R&D departments. Companies developing scientific AI assistants may use it to demonstrate capabilities, potentially opening new markets in drug discovery, materials science, and climate modeling.
If LLMs perform well on SCILAWS-BENCH, we may see increased investment in AI-driven scientific discovery tools. Conversely, poor performance could highlight fundamental limitations in current models' ability to generalize beyond training data. Watch for follow-up papers reporting model scores on this benchmark.