Paper 2605.18661
AI for Auto-Research: Roadmap & User Guide
- Published
- May 2026
- Research lab
- Independent
- Citations
- 5
- GitHub
- 473 stars
01 In brief
Summary
This paper surveys AI-assisted research across the complete academic lifecycle, organized into four phases (Creation, Writing, Validation, Dissemination) and eight stages.
It finds that AI excels at structured, retrieval-grounded tasks but remains unreliable for novel ideas, research-level experiments, and scientific judgment.
Key findings include: artifact generation outpaces verification; human-governed collaboration is the most credible deployment mode; and AI use is becoming a governance problem rather than a detection problem.
The paper provides a taxonomy, benchmark suite, tool inventory, and practitioner playbook, with resources maintained on a project page.
02 From the paper
Abstract
AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-horizon agents can execute experiments, draft manuscripts, and simulate critique with minimal human input. Yet this productivity frontier exposes a deeper integrity problem: under scientific pressure, even frontier LLMs still fabricate results, miss hidden errors, and fail to judge novelty reliably. Studying developments through April 2026, we present an end-to-end analysis of AI across the complete research lifecycle, organized into four epistemological phases: Creation (idea generation, literature review, coding & experiments, tables & figures), Writing (paper writing), Validation (peer review, rebuttal & revision), and Dissemination (posters, slides, videos, social media, project pages, and interactive agents). We identify a sharp, stage-dependent boundary between reliable assistance and unreliable autonomy: AI excels at structured, retrieval-grounded, and tool-mediated tasks, but remains fragile for genuinely novel ideas, research-level experiments, and scientific judgment. Generated ideas often degrade after implementation, research code lags far behind pattern-matching benchmarks, and end-to-end autonomous systems have not yet consistently reached major-venue acceptance standards. We further show that greater automation can obscure rather than eliminate failure modes, making human-governed collaboration the most credible deployment paradigm. Finally, we provide a structured taxonomy, benchmark suite, and tool inventory, cross-stage design principles, and a practitioner-oriented playbook, with resources maintained at our project page.