Yao Xiao

PhD student at SUTD · Research Intern at MiroMind. I work on LLM post-training, with hands-on experience building search agents and code agents.

Yao Xiao

About

I am a researcher focusing on LLM post-training — supervised fine-tuning, reinforcement learning, and reward modeling for aligning large language models with downstream tasks. I am especially interested in agentic systems, and have built and trained both search agents and code agents.

I am currently a PhD student at the Singapore University of Technology and Design (SUTD), and a research intern at MiroMind. I received my Master's degree from Shanghai Jiao Tong University (SJTU) and my Bachelor's degree from Tianjin University.

I believe the simplest methods that scale tend to be the ones that last.

Feel free to reach out at xiao_yao@mymail.sutd.edu.sg.

News

Publications

Author names in bold indicate me. See also my Google Scholar.

Apodex-1.0: A Verification-Centric Agent Team for Discoverative Intelligence

Apodex Team (incl. Yao Xiao)

Tech Report 2026 · Apodex AI

MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification

MiroMind Team (incl. Yao Xiao)

arXiv 2026 · Preprint

Document Reconstruction Unlocks Scalable Long-Context RLVR

Yao Xiao, Lei Wang, Yue Deng, Guanzheng Chen, Ziqi Jin, Jung-jae Kim, Xiaoli Li, Roy Ka-Wei Lee, Lidong Bing

arXiv 2026 · Preprint

Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficulty

Yao Xiao, Jung-jae Kim, Roy Ka-Wei Lee, Lidong Bing

ACL 2026 · Annual Meeting of the Association for Computational Linguistics

MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization

Xingxuan Li, Yao Xiao, Dianwen Ng, Hai Ye, et al.

arXiv 2025 · Technical Report

Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization

Yao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng, Lidong Bing, Xiaoli Li, Roy Ka-Wei Lee

ACL 2025 · Annual Meeting of the Association for Computational Linguistics (Long Papers)

Decomposed Prompt Tuning via Low-Rank Reparameterization

Yao Xiao, Lu Xu, Jiaxi Li, Wei Lu, Xiaoli Li

EMNLP 2023 · Findings of the Association for Computational Linguistics