Strategic and Social Decision-Making
Wenyue Hua
Aug 13, 2026
Overview
This line runs from simulation to training. On the simulation side, multi-agent systems reproduce historical conflict and collective behavior at scale, which lets us examine the triggers and conditions that lead to war rather than only its outcomes. On the analysis side, game-theoretic evaluation shows where LLMs depart from rational play as payoff matrices grow and sequential trees deepen, and structured workflows steer them back toward equilibrium and better negotiation outcomes. On the training side, reinforcement learning targets social reasoning directly, so that a small model can match frontier systems at negotiation in domain rather than inheriting the dispositions of a general-purpose assistant.
Publications
AI agents increasingly act on their users’ behalf, but the dispositions that make an assistant pleasant can make it a poor delegate. A friendly frontier model may disclose its principal’s private information unprompted and concede at the first sign of resistance. We present SocialRL, a general recipe that trains social reasoning directly, and apply it to a 4B model across six negotiation and persuasion domains. In-domain training reaches the frontier, with the 4B model matching or exceeding the GPT-5 family per domain on held-out scenarios. Cross-domain transfer follows game structure, and cascade RL and multi-teacher on-policy distillation consolidate the per-domain specialists into a single unified 4B model.
Wenyue Hua,
Zachary Huang,
Tyler Payne,
Safoora Yousefi,
Saleema Amershi,
Asli Celikyilmaz
Role-Playing Agent (RPA) is an increasingly popular type of LLM Agent that simulates human-like behaviors in a variety of tasks. However, evaluating RPAs is challenging due to diverse task requirements and agent designs. This paper proposes an evidence-based, actionable, and generalizable evaluation design guideline for LLM-based RPA by systematically reviewing 1,676 papers published between Jan. 2021 and Dec. 2024. Our analysis identifies six agent attributes, seven task attributes, and seven evaluation metrics from existing literature. Based on these findings, we present an RPA evaluation design guideline to help researchers develop more systematic and consistent evaluation methods.
Chaoran Chen,
Bingsheng Yao,
Ruishi Zou,
Wenyue Hua,
Weimin Lyu,
Yanfang Ye,
Toby Jia-Jun Li,
Dakuo Wang
This paper investigates the rationality of large language models (LLMs) in strategic decision-making contexts, specifically within the framework of game theory. We evaluate several state-of-the-art LLMs across a spectrum of complete-information and incomplete-information games. Our findings reveal that LLMs frequently deviate from rational strategies, particularly as the complexity of the game increases with larger payoff matrices or deeper sequential trees. To address these limitations, we design multiple game-theoretic workflows that guide the reasoning and decision-making processes of LLMs. These workflows aim to enhance the models’ ability to compute Nash Equilibria and make rational choices, even under conditions of uncertainty and incomplete information. Experimental results demonstrate that the adoption of these workflows significantly improves the rationality and robustness of LLMs in game-theoretic tasks. Specifically, with the workflow, LLMs exhibit marked improvements in identifying optimal strategies, achieving near-optimal allocations in negotiation scenarios, and reducing susceptibility to exploitation during negotiations. Furthermore, we explore the meta-strategic considerations of whether it is rational for agents to adopt such workflows, recognizing that the decision to use or forgo the workflow constitutes a game-theoretic issue in itself. Our research contributes to a deeper understanding of LLMs’ decision-making capabilities in strategic contexts and provides insights into enhancing their rationality through structured workflows. The findings have implications for the development of more robust and strategically sound AI agents capable of navigating complex interactive environments.Coda and data are available this url.
Wenyue Hua,
Ollie Liu,
Lingyao Li,
Alfonso Amayuelas,
Julie Chen,
Lucas Jiang,
Lizhou Fan,
Fei Sun,
William Yang Wang,
Xintong Wang,
Yongfeng Zhang
Can we avoid wars at the crossroads of history? This question has been pursued by individuals, scholars, policymakers, and organizations throughout human history. In this research, we attempt to answer the question based on the recent advances of Artificial Intelligence (AI) and Large Language Models (LLMs). We propose WarAgent, an LLM-powered multi-agent AI system, to simulate the participating countries, their decisions, and the consequences, in historical international conflicts, including the World War I (WWI), the World War II (WWII), and the Warring States Period (WSP) in Ancient China. By evaluating the simulation effectiveness, we examine the advancements and limitations of cutting-edge AI systems’ abilities in studying complex collective human behaviors such as international conflicts under diverse settings. In these simulations, the emergent interactions among agents also offer a novel perspective for examining the triggers and conditions that lead to war. Our findings offer data-driven and AI-augmented insights that can redefine how we approach conflict resolution and peacekeeping strategies. The implications stretch beyond historical analysis, offering a blueprint for using AI to understand human history and possibly prevent future international conflicts. Code and data are available at this url.
Wenyue Hua,
Lizhou Fan,
Lingyao Li,
Kai Mei,
Jianchao Ji,
Yingqiang Ge,
Libby Hemphill,
Yongfeng Zhang