I study and train LLM agents that negotiate, cooperate, and act on people’s behalf, and I work on making them safe and efficient enough to deploy. My work runs across three lines: strategic and social decision-making, trustworthiness, and efficiency. Two earlier lines, reasoning evaluation and generative recommendation, still inform how I frame problems, as does my background in formal linguistics.
LLM agents increasingly act on someone’s behalf across from a counterpart whose goals may conflict with their principal’s. I study how such agents negotiate, cooperate, and compete, and how to train the social reasoning that makes them good delegates rather than agreeable ones.
An agent that acts in the world can cause material harm, leak what its principal told it, or simply fail. I work on the mechanisms that make deployment defensible: safety constraints during planning, privacy-preserving communication, and risk standards that pay out when an agent gets it wrong.
Agent pipelines are slow and expensive because they generate long chains of intermediate thoughts through large models. I work on making them cheaper and faster without giving up the quality of the final output, and on giving developers and users control over that tradeoff.
Benchmarks decide what progress looks like, so they should be hard to overfit and should measure something with structure. I build reasoning benchmarks organized by computational complexity, and study where LLM reasoning is genuine and where it reflects patterns already seen in pretraining.
Recommendation foundation models turn recommendation into a language task, generating the item to recommend instead of scoring every candidate. I work on what has to be true for that to function, starting with how items are named.
My background is in formal linguistics, and the questions there still shape how I frame problems: what class of function is a phenomenon in, what can be learned from what kind of evidence, and what a model is provably unable to represent.