Trustworthy Agents: Safety, Risk, and Privacy

Overview

Agent trustworthiness is not a single property. Safety can be built into the planning loop itself, through an agent constitution that injects safety knowledge before a plan is generated, constrains generation while it happens, and inspects the result afterwards. Privacy is a separate axis: cloud models see everything a user sends them, so prompts can be obfuscated generatively and still complete the task. Beyond the model, there is the question of what a user is owed when an agent fails. Treating that as a settlement problem rather than a purely technical one turns trust from an implicit expectation about model behavior into an explicit and enforceable guarantee.

Wenyue Hua
Wenyue Hua
Senior Researcher

Ph.D. in Computer Science, focused on large language models.