Publications
Conference Papers
SCOPE: Stepwise Calibration of Reasoning Confidence with Outcome-Grounded Prefix Evidence
Conference on Empirical Methods in Natural Language Processing (EMNLP Main), 2026
A framework that calibrates confidence after every reasoning step using outcome-grounded prefix evidence, enabling reliable inference-time risk control and confidence-weighted aggregation.
StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction
ICML 2026 FAGEN Workshop
A framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning for long-horizon decision making.
CoMAS: Co-Evolving Multi-Agent Systems via Interaction Rewards
International Conference on Learning Representations (ICLR), 2026
A framework that enables LLM-based agents to self-evolve through interaction rewards in a multi-agent setting.
Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical Reasoning
Annual Meeting of the Association for Computational Linguistics (ACL Main), 2026
A systematic empirical investigation of scaling behaviors in RL-based post-training, focusing on mathematical reasoning.
Empowering Users in Digital Privacy Management through Interactive LLM-Based Agents
International Conference on Learning Representations (ICLR), 2025
An interactive LLM-based dialogue agent that helps users understand privacy policies, setting new benchmarks in privacy policy analysis.
Journal Articles
RadFabric: Agentic AI System with Reasoning Capability for Radiology
npj Digital Medicine, 2026
A multi-agent, multimodal radiology framework that integrates visual, anatomical, and textual reasoning for chest X-ray interpretation.
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Transactions on Machine Learning Research (TMLR), 2026
A survey formalizing the shift from conventional LLM RL to agentic RL, where LLMs act as autonomous decision-making agents.
Tutorials
AI Scientist: Persistent Scientific Systems with Memory, Verification, and Human Governance
AACL-IJCNLP 2026 Tutorial
A tutorial on persistent AI scientist systems, covering long-horizon memory, novelty verification, scientific infrastructure, embodied experimentation, epistemic reliability, and human governance.
