1.Field Maps and Surveys
Pre-reading
Additional reading
- Agentic AI for Scientific Discovery: A Survey of Progress, Challenges, and Future Directions.Paper
AI Scientists are evolving from autonomous workflow agents that automate a research pipeline into persistent scientific systems that maintain long-horizon research state, verify novelty against the prior literature, track experimental provenance, execute in real or simulated laboratories, and collaborate with human scientists across iterative discovery loops.
The view that dominated 2025 — an AI Scientist as an autonomous agent that runs an end-to-end research pipeline (autoresearch, AI-Researcher) — has given way to a view in which the AI Scientist is treated as a persistent scientific system (LabOS, DeepScientist, Lumilab, Rethinking the AI Scientist). The hard problems have shifted from agent orchestration to research-state evolution, novelty verification, scientific infrastructure, embodied execution, epistemic reliability, and scientist-in-the-loop governance.
The literature is fragmented across NLP, ML, agent, and domain-specific venues, and is organized around outdated framings like “end-to-end automation” or “multi-agent debate.” These framings obscure the actual frontier.
This tutorial reframes AI Scientist research around six concerns: persistent systems, literature discovery and novelty verification, scientific infrastructure, embodied scientific interaction, reliability and provenance, and human governance.
You leave with a compact field map of the six concerns, a reusable vocabulary that separates agent architecture from scientific infrastructure, a curated reading list, classroom-ready slides, and a minimal AI Scientist walkthrough notebook that reads a paper, generates a hypothesis, runs an experiment, and produces a report.
From motivating cases through persistent systems, novelty verification, scientific infrastructure, embodied science, reliability, and human governance.
Concrete cases: a citation-hallucinating drafting agent; a coding agent that cannot reproduce its own result; a biomedical agent that proposes a CRISPR screen and hands the protocol to a human collaborator; and a long-horizon system with branching, failed paths, and human takeover.
The central reframing of the tutorial. An AI Scientist is not a one-shot pipeline but a persistent system that maintains long-horizon state: hypothesis lineage, experiment graphs, scientific memory, research branching. We compare DeepScientist, Lumilab, and Rethinking the AI Scientist against autoresearch and AI-Researcher.
Idea generation is cheap; novelty validation is hard. Learned idea comparators that outperform GPT-5; literature similarity filtering integrated into end-to-end systems; AutoResearchBench; citation grounding and faithfulness; hypothesis validation in AI co-scientist.
Recent systems are best understood as research operating systems: persistent execution environments, sandboxing, experiment caching, artifact lineage, and provenance tracking. Biomni as canonical action space; SARE; MiroFlow. Closes with a live walkthrough in which attendees examine common failure modes (hallucinated novelty, broken provenance, and irreproducible results) in real execution traces.
Scientific reasoning is not text reasoning. AI-XR co-scientists in the laboratory (LabOS); embodied scientific benchmarks (LabUtopia); agentic wet-lab automation (CRISPR-GPT, Biomni). The NLP problems that surface when language models talk to instruments, robots, and human collaborators.
Not whether a system produces a research artifact, but whether the artifact is epistemically reliable. Provenance, falsification, calibrated uncertainty, reproducibility, hallucinated novelty, benchmark leakage. ResearchGym, ResearchClawBench, Scientist-Bench, AutoResearchBench, VeriTrace.
Human governance is not a safety appendix but a core architectural concern. Four design patterns: branching workflows, checkpoint approval, interactive refinement, co-scientist loops. LabOS, DeepScientist, Rethinking the AI Scientist. Closes with Q&A and open problems.
Accepted at AACL-IJCNLP 2026 · November 6–10, 2026 · Hengqin, China.
Researchers, Ph.D. students, advanced master's students, and industry practitioners working on language agents, tool learning, RAG, deep research, evaluation, scientific NLP, and applied LM systems.
The conference will be held in hybrid format and physically in Hengqin, China, from November 6 to 10, 2026. The specific tutorial date and room will be posted when announced.
Organized by tutorial workflow — pre-readings recommended before attending, additional readings for broader background. Links will be added as releases become public.
Four presenters and a senior advisor · agents, multi-agent systems, AI Scientists, and computer vision.
Presenters
Advisory Board