Apodex-1.0: A Verification-Centric Agent Team for Discoverative Intelligence
Tech Report 2026 · Apodex AI
PhD student at SUTD · Research Intern at MiroMind. I work on LLM post-training, with hands-on experience building search agents and code agents.
I am a researcher focusing on LLM post-training — supervised fine-tuning, reinforcement learning, and reward modeling for aligning large language models with downstream tasks. I am especially interested in agentic systems, and have built and trained both search agents and code agents.
I am currently a PhD student at the Singapore University of Technology and Design (SUTD), and a research intern at MiroMind. I received my Master's degree from Shanghai Jiao Tong University (SJTU) and my Bachelor's degree from Tianjin University.
I believe the simplest methods that scale tend to be the ones that last.
Feel free to reach out at xiao_yao@mymail.sutd.edu.sg.
Author names in bold indicate me. See also my Google Scholar.
Apodex-1.0: A Verification-Centric Agent Team for Discoverative Intelligence
Tech Report 2026 · Apodex AI
MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification
arXiv 2026 · Preprint
Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficulty
ACL 2026 · Annual Meeting of the Association for Computational Linguistics
MiroMind-M1: An Open-Source Advancement in Mathematical Reasoning via Context-Aware Multi-Stage Policy Optimization
arXiv 2025 · Technical Report