Research
Security, auditability, and governance of generative AI agents
Uílton de Oliveira Dutra
Professional Doctorate in Creative Industries · Universidade Feevale · 2025-2029
Advised by Dr. Marta Rosecler Bez
When an AI agent acts on your behalf, what security guarantees remain valid, and how would you prove it? This question underlies my doctoral research. I want to make the security and auditability of generative AI agents assessable and comparable, rather than something each system's builders simply assert. The goal is to move the debate from claims about reliability to evidence for security guarantees by asking what an agent is allowed to do, what it can access, how its actions are constrained, and whether these properties can be demonstrated consistently across architectures.
The question is important because governance of these systems is still catching up with their capabilities. Tool-integrated agents can combine language-model reasoning with external actions. Research systems already interleave reasoning and tool use, and security benchmarks model agents operating email, banking, travel, and other services over untrusted data [2, 11]. When such systems are connected to credentials or confidential data, their errors and compromises can have real consequences. The promise is enormous. These systems can adapt to intentions, automate knowledge work, and reduce barriers to construction. Yet each of these features also creates an attack surface, and the same session that reads confidential files may also encounter instructions supplied by an attacker [3, 11].
My research lies at the intersection of systems security, software architecture, and generative AI governance. Defending these systems responsibly must be a first-rate engineering concern, pursued with the same rigor we bring to their capabilities, rather than added a posteriori as safeguards around an architecture whose guarantees remain unclear. I work on the question from both sides: as CTO at Prodeal, I lead engineering for a virtual data room platform where confidentiality, controlled access, and auditability are the product rather than add-on features.
Many distinctive agent-security failures share a structural cause. A language model predicts successive tokens from its context [1], while an agent repeatedly feeds model outputs and tool observations back into that context [2]. Current language models do not provide a dependable, formal separation between trusted instructions and untrusted data [3, 4]. As agents gain autonomy and authority, enforcing trust boundaries outside the model becomes a defining security challenge.
The field has responded with sandboxes, orchestration layers, access controls, policy engines, capability systems, and audit mechanisms. These mechanisms are valuable, but they are also fragmented and difficult to compare because they operate at different levels and rely on different assumptions. Security is therefore often asserted on a system-by-system basis, with no common way to establish what a given architecture actually guarantees, when those guarantees are valid, or where they fail.
Rather than adding additional architecture to the stack, I work on the evaluative foundations that the field is missing. This means developing principled methods to reason about what secure, verifiable agent execution requires, as well as methods for distinguishing between systems that appear equivalent on paper but offer very different guarantees in practice. Model robustness alone is insufficient; a system must state and enforce its security invariants, threat model, assumptions, and failure boundaries [5, 6]. An architecture should explicitly state what it secures, against what threats, and under what assumptions, rather than treating a list of features as proof of security.
Research area
Agent security is a systems problem. The language model should be treated as an untrusted component, with security guarantees enforced by the surrounding system rather than inferred from or entrusted to model behavior [5, 6]. Phrased this way, the questions become familiar to any security engineer: what is isolated from what, what authority is delegated to whom, what can be proven after the fact, and how far a single compromise can go.
What differentiates agentic systems is the coupling of probabilistic planning with tools, state, and delegated authority. Agents act continuously, chain actions across tools and other agents, and hold long-term state and credentials. Controls designed for human-operated sessions and static roles can require agent-specific mediation, least-privilege delegation, and continuous evidence collection. Understanding where these controls break down and what to replace them with is at the heart of the field I study.
Objectives
A small set of objectives guides the research, pursued throughout the PhD and beyond:
- Make agent security assessable. Establish principled ways to reason about and compare what different agent architectures actually secure.
- Turn assessment into guidance. Produce assessments and decision aids that practitioners, auditors, and engineering teams can use to select and operate agent systems responsibly.
- Connect security to governance. Connect the technical picture to the laws, standards, and risk-management frameworks shaping AI governance, including the EU AI Act [7], the NIST AI Risk Management Framework [8] and its Generative AI Profile [9], and ISO/IEC 42001 [10].
- Bridge production and research. Keep the work grounded in deployed systems and transfer what's learned in production into research, and vice versa.
Method
The research follows a design science approach [12]. The primary artifact is an architecture-neutral assessment framework: a threat model, a set of verifiable security properties, graded assurance levels, and a reproducible evaluation procedure. Its evaluation criteria are fixed before any results are known, rather than chosen afterward to fit them, and the work proceeds in cycles, because the systems being evaluated keep changing while the evaluation is underway. Each reported result preserves its versioned assumptions, procedures, and evidence so that another evaluator can reproduce or challenge it.
Let's collaborate
I am looking for a six-month doctoral research stay abroad, starting in September 2027, hosted by a group working on agent security, system security, or the governance of autonomous AI systems. If this describes your lab, I would welcome a conversation about fit well in advance of the application deadline. I'm also keen to collaborate with teams deploying or operating agents in production who want these deployments studied, and to discuss this work with research, engineering, or policy audiences.
If your work touches secure agent execution, in research, in production, or in policy, I'd be glad to connect:
References
- T. B. Brown et al., "Language models are few-shot learners," in Advances in Neural Information Processing Systems (NeurIPS), vol. 33, 2020, pp. 1877-1901. [Online]. Available: proceedings.neurips.cc
- S. Yao et al., "ReAct: Synergizing reasoning and acting in language models," in Proc. Int. Conf. Learning Representations (ICLR), 2023. [Online]. Available: openreview.net
- K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, "Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection," in Proc. 16th ACM Workshop on Artificial Intelligence and Security (AISec), 2023, pp. 79-90, doi: 10.1145/3605764.3623985.
- E. Zverev, S. Abdelnabi, S. Tabesh, M. Fritz, and C. H. Lampert, "Can LLMs separate instructions from data? And what do we even mean by that?" in Proc. Int. Conf. Learning Representations (ICLR), 2025. [Online]. Available: openreview.net
- M. Christodorescu et al., "Agent security is a systems problem," arXiv:2605.18991v2, May 2026. [Online]. Available: arxiv.org
- E. Debenedetti et al., "Defeating prompt injections by design," arXiv:2503.18813v2, Jun. 2025. [Online]. Available: arxiv.org
- European Parliament and Council, "Regulation (EU) 2024/1689 of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act)," Official Journal of the European Union, L, 2024/1689, 12 Jul. 2024. [Online]. Available: eur-lex.europa.eu
- E. Tabassi, "Artificial intelligence risk management framework (AI RMF 1.0)," NIST, Gaithersburg, MD, USA, Rep. NIST AI 100-1, 2023, doi: 10.6028/NIST.AI.100-1.
- National Institute of Standards and Technology, "Artificial intelligence risk management framework: Generative artificial intelligence profile," NIST, Gaithersburg, MD, USA, Rep. NIST AI 600-1, 2024, doi: 10.6028/NIST.AI.600-1.
- ISO/IEC, ISO/IEC 42001:2023, Information technology - Artificial intelligence - Management system, 2023. [Online]. Available: iso.org
- E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr, "AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents," in Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2024. [Online]. Available: openreview.net
- A. R. Hevner, S. T. March, J. Park, and S. Ram, "Design science in information systems research," MIS Quarterly, vol. 28, no. 1, pp. 75-105, 2004. [Online]. Available: aisel.aisnet.org
