Paper 2508.21148
A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
- Published
- Aug 2025
- Research lab
- Independent
- Citations
- 26
- GitHub
- 458 stars
01 In brief
Summary
This survey reframes the development of Scientific Large Language Models (Sci-LLMs) as a co-evolution between models and their data substrate, providing a data-centric synthesis across six scientific domains (physics, chemistry, materials science, life sciences, astronomy, and Earth science).
It introduces a unified taxonomy of scientific data and a hierarchical model of scientific knowledge, highlighting the multimodal, cross-scale, and domain-specific challenges that differentiate scientific corpora from general NLP datasets.
The survey systematically reviews recent Sci-LLMs, from general-purpose foundations to specialized models, and analyzes over 270 pre-/post-training datasets and over 190 evaluation benchmarks.
It identifies persistent issues in scientific data development, such as scarcity of experimental data, over-reliance on text, and multi-level biases, and discusses emerging solutions like semi-automated annotation pipelines and expert validation.
The paper outlines a paradigm shift toward closed-loop systems where autonomous agents based on Sci-LLMs actively experiment, validate, and contribute to a living knowledge base, providing a roadmap for building trustworthy, continually evolving AI systems for scientific discovery.
The evolution of Sci-LLMs is traced through four phases: transfer learning, scaling, instruction-following, and agentic science, with a noted trend toward smaller parameter sizes (7B-13B) and the dominance of open-source base models like LLaMA and Qwen.
Evaluation is shifting from static exams to process- and discovery-oriented assessments, with advanced protocols like LLM-as-a-Judge and test-time learning.
The survey concludes by outlining future directions including integrated data ecosystems, automated data standardization, comprehensive evaluation systems, and autonomous scientific agents, emphasizing the need for ethical governance and responsible AI innovation.
The work is a…
02 From the paper
Abstract
Scientific Large Language Models (Sci-LLMs) are transforming how knowledge is represented, integrated, and applied in scientific research, yet their progress is shaped by the complex nature of scientific data. This survey presents a comprehensive, data-centric synthesis that reframes the development of Sci-LLMs as a co-evolution between models and their underlying data substrate. We formulate a unified taxonomy of scientific data and a hierarchical model of scientific knowledge, emphasizing the multimodal, cross-scale, and domain-specific challenges that differentiate scientific corpora from general natural language processing datasets. We systematically review recent Sci-LLMs, from general-purpose foundations to specialized models across diverse scientific disciplines, alongside an extensive analysis of over 270 pre-/post-training datasets, showing why Sci-LLMs pose distinct demands -- heterogeneous, multi-scale, uncertainty-laden corpora that require representations preserving domain invariance and enabling cross-modal reasoning. On evaluation, we examine over 190 benchmark datasets and trace a shift from static exams toward process- and discovery-oriented assessments with advanced evaluation protocols. These data-centric analyses highlight persistent issues in scientific data development and discuss emerging solutions involving semi-automated annotation pipelines and expert validation. Finally, we outline a paradigm shift toward closed-loop systems where autonomous agents based on Sci-LLMs actively experiment, validate, and contribute to a living, evolving knowledge base. Collectively, this work provides a roadmap for building trustworthy, continually evolving artificial intelligence (AI) systems that function as a true partner in accelerating scientific discovery.