Publications

You can also find my articles on my Google Scholar profile.

Conference Papers


SCOPE: Stepwise Calibration of Reasoning Confidence with Outcome-Grounded Prefix Evidence

Huaqin Zhao*, Yifan Zhou*, Bolun Sun, Chao Cao, Yiwei Li, Zhaojun Ding, Lu Zhang, Tianming Liu, Jin Lu

* Equal contribution (co-first authors).

Conference on Empirical Methods in Natural Language Processing (EMNLP Main), 2026

A framework that calibrates confidence after every reasoning step using outcome-grounded prefix evidence, enabling reliable inference-time risk control and confidence-weighted aggregation.

StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction

Xiangyuan Xue*, Yifan Zhou*, Zidong Wang, Shengji Tang, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin

* Equal contribution (co-first authors).

ICML 2026 FAGEN Workshop

A framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning for long-horizon decision making.

CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards

Xiangyuan Xue, Yifan Zhou, Guibin Zhang, Zaibin Zhang, Yijiang Li, Chen Zhang, Zhenfei Yin, Philip Torr, Wanli Ouyang, Lei Bai

International Conference on Learning Representations (ICLR), 2026

A framework that enables LLM-based agents to self-evolve through interaction rewards in a multi-agent setting.

Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning

Zelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang, Guancheng Wan, Yifan Zhou, Qiang He, Xiangyuan Xue, Heng Zhou, Yutao Fan, Zhongzhi Li, Zaibin Zhang, Guibin Zhang, Chen Zhang, Zhenfei Yin, Philip Torr, Lei Bai

Annual Meeting of the Association for Computational Linguistics (ACL Main), 2026

A systematic empirical investigation of scaling behaviors in RL-based post-training, focusing on mathematical reasoning.

Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents

Bolun Sun, Yifan Zhou, Haiyun Jiang

International Conference on Learning Representations (ICLR), 2025

An interactive LLM-based dialogue agent that helps users understand privacy policies, setting new benchmarks in privacy policy analysis.

Journal Articles


RadFabric: Agentic AI System with Reasoning Capability for Radiology

Wenting Chen, Yi Dong, Zhaojun Ding, Yucheng Shi, Yifan Zhou, Fang Zeng, Yijun Luo, Tianyu Lin, Yihang Su, Yichen Wu, Kai Zhang, Zhen Xiang, Tianming Liu, Ninghao Liu, Lichao Sun, Yixuan Yuan, Xiang Li

npj Digital Medicine, 2026

A multi-agent, multimodal radiology framework that integrates visual, anatomical, and textual reasoning for chest X-ray interpretation.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

Guibin Zhang, Hejia Geng, Xiaohang Yu, Zhenfei Yin, Zaibin Zhang, Zelin Tan, Heng Zhou, Zhongzhi Li, Xiangyuan Xue, Yijiang Li, Yifan Zhou, Yang Chen, Chen Zhang, Yutao Fan, Zihu Wang, Songtao Huang, Francisco Piedrahita-Velez, Yue Liao, Hongru Wang, Mengyue Yang, Heng Ji, Jun Wang, Shuicheng Yan, Philip Torr, Lei Bai

Transactions on Machine Learning Research (TMLR), 2026

A survey formalizing the shift from conventional LLM RL to agentic RL, where LLMs act as autonomous decision-making agents.

Tutorials


AI Scientist: Persistent Scientific Systems with Memory, Verification, and Human Governance

Wanghan Xu, Yifan Zhou, Zhenfei Yin, Yingcheng Wu, Philip Torr

AACL-IJCNLP 2026 Tutorial

A tutorial on persistent AI scientist systems, covering long-horizon memory, novelty verification, scientific infrastructure, embodied experimentation, epistemic reliability, and human governance.