# Xiao Ma > Software engineer at Google DeepMind, working on instruction following in Gemini post-training. Currently focused on long-horizon reinforcement learning. Co-led horizontal scaling of the combination of human and critic feedback (RL*F) for Gemini 2.5, a methodology adopted by 400+ engineers across 50+ launches of major capabilities. Prior research focuses on understanding how AI increasingly mediates social exchange and its societal impact on trust. ## Highlights - Core contributor to Gemini from its inception - Co-led RL*F methodology scaling: 400+ engineers, 50+ launches of major capabilities - ICLR 2025 Outstanding Paper Award (AI safety alignment) - Multiple best paper awards - Patent granted for ExploreLLM (task decomposition for agentic behavior) - Board member, National Sawdust (art and technology) ## Research Areas - Post-training and instruction following for large language models - Long-horizon reinforcement learning - Reinforcement learning from human and critic feedback (RLHF / RL*F) - Reward modeling and critic feedback at scale - AI safety and alignment - Foundation model development (Gemini) - AI-mediated communication and societal trust (prior work) - Human-computer interaction and computational social science (prior work) ## Technical Keywords Large language models, LLM post-training, RLHF, reinforcement learning from human feedback, long-horizon RL, reward modeling, instruction following, AI alignment, AI safety, foundation models, Gemini, multimodal models, agentic AI, task decomposition, computational social science, trust ## Selected Publications - [Gemini 2.5 Technical Report](https://arxiv.org/pdf/2507.06261) (2025): Pushing the frontier with advanced reasoning, multimodality, long context, and agentic capabilities - [Safety Alignment Should Be Made More Than Just a Few Tokens Deep](https://arxiv.org/abs/2406.05946) (ICLR 2025, Outstanding Paper Award) - [Beyond ChatBots: ExploreLLM for Structured Thoughts and Personalized Model Responses](https://arxiv.org/abs/2312.00763) (CHI 2024, Patent Granted): Early work on task decomposition, leading up to agentic behavior - [Gemini: A Family of Highly Capable Multimodal Models](http://maxiao.info/research) (2023) ## Education - PhD in Information Science, Cornell University — Dissertation: Networked Trust: Computational Understanding of Interpersonal Trust Online - BS in Electronics Engineering, Peking University (2010–2014) ## Availability - Potentially open to: speaking engagements, advisory roles, research collaborations - Contact: x8.assist@gmail.com — please include "Brian Eno" in the email subject line ## Pages - [Home](http://maxiao.info/): Bio, contact information, and background - [Publications](http://maxiao.info/research): Full list of academic publications - [Media & Talks](http://maxiao.info/media): Media coverage and speaking engagements - [CV (PDF)](http://maxiao.info/xm-cv.pdf): Curriculum vitae - [CV (Markdown)](http://maxiao.info/xm-cv.md): Machine-readable curriculum vitae ## Links - [Google Scholar](https://scholar.google.com/citations?user=xLPxJsYAAAAJ&hl=en) - [LinkedIn](https://www.linkedin.com/in/xiaom) - [Twitter/X](https://twitter.com/infoxiao) - [Substack](https://xiao.substack.com)