I’m Wenyue Hua, a Senior Researcher at Microsoft Research, AI Frontiers.
My work runs across three lines. Strategic and social decision-making in LLM agents, from simulation to training: multi-agent simulation of strategic conflict (WarAgent), game-theoretic workflows that make LLM negotiators more rational (Game-theoretic LLM), and reinforcement learning that lets a 4B model match or exceed the GPT-5 family in-domain across six negotiation domains (SocialRL). Trustworthiness: TrustAgent, the Agentic Risk Standard, and EmojiPrompt. Efficiency: Interactive Speculative Planning, Dynamic Speculative Agent Planning, and AgentOpt. Earlier, I worked on generative recommendation (How to Index Item IDs) and reasoning evaluation (NPHardEval, InductionBench, OpenAGI).
Before Microsoft, I was a postdoctoral researcher at UC Santa Barbara with Prof. William Yang Wang (2024 to 2025). I received my Ph.D. in Computer Science from Rutgers University (2020 to 2024), advised by Prof. Yongfeng Zhang. I also hold an MA in Linguistics from Rutgers (advised by Prof. Adam Jardine) and a BS in Mathematics and a BA in Linguistics and Philosophy from UCLA (advised by Prof. Edward Keenan).
I founded NICE (2023), an AI/NLP research community that hosts talks on new papers and projects. If you have work you would like to present, contact nice.ai.academy@gmail.com.
Postdoctoral Research in Computer Science, 2024-2025
Computer Science Department, University of California, Santa Barbara
Ph.D. in Computer Science, 2020-2024
Computer Science Department, Rutgers University, New Brunswick
Master of Arts in Linguistics (Ph.D. track transfer out), 2018-2020
Department of Linguistics, Rutgers University, New Brunswick
B.S. in Mathematics, General & B.A. in Linguistics&Philosophy with Specialization in Computing, 2014-2018
UCLA
Papers I led as first, co-first, or senior author. Full list on the Publications page.
SocialRL: From Passive Delegates to Strategic Negotiators (2026, arXiv 2608.13787). A general RL recipe that trains social reasoning directly. A 4B model matches or exceeds the GPT-5 family in-domain across six negotiation domains. PDF
Game-theoretic LLM: Agent Workflow for Negotiation Games (2024, arXiv 2411.05990). Shows where LLMs depart from rational play in complete- and incomplete-information games, and designs game-theoretic workflows that steer them toward equilibrium and better negotiation outcomes. PDF
WarAgent (2023, arXiv 2311.17227). LLM multi-agent simulation of WWI, WWII, and the Warring States period, used to study the triggers and conditions that lead to war. PDF
TrustAgent (2024, Findings of EMNLP). An agent-constitution framework with pre-planning, in-planning, and post-planning safety strategies. PDF
Quantifying Trust: the Agentic Risk Standard (2026, arXiv 2604.03976). A settlement-layer standard that integrates risk assessment, underwriting, and compensation for AI-mediated transactions. Spotlighted in Fortune. PDF
EmojiPrompt (2025, NAACL). Generative prompt obfuscation so cloud LLMs can complete tasks without seeing raw private content. Co-first author. PDF
Interactive Speculative Planning (2025, ICLR). Speculative execution for agent planning, co-designed with a user interface that treats human interruption as a first-class part of the system rather than an exception. PDF
Dynamic Speculative Agent Planning (2026, ICLR). An asynchronous online RL framework for lossless acceleration of agent planning, exposing a single parameter that trades latency against dollar cost and cutting total cost by 30%. Senior author. PDF
AgentOpt (2026, technical report, arXiv 2604.06296). Client-side model selection for agent pipelines. The cost gap between the best and worst model combinations reaches 13–32x, and Arm Elimination cuts evaluation budget by 24–67% at near-optimal accuracy. Open-source package. PDF · Code
How to Index Item IDs for Recommendation Foundation Models (2023, SIGIR-AP). Systematic study of item ID construction for LLM-based generative recommendation, with sequential, collaborative, semantic, and hybrid indexing. PDF
NPHardEval (2024, ACL). A dynamic reasoning benchmark organized by computational complexity class and refreshed regularly to resist overfitting. Co-first author. PDF
InductionBench (2025, ACL). A benchmark for inductive reasoning, inferring the underlying rule from observations rather than applying a given one. Frontier models fail even the simplest complexity classes of the subregular hierarchy. PDF · Code
Honors
Press
Invited talks
Service