Baseline
Standard task access, stable tools, intended memory and context, normal response time, and standard human review points.
Research project · Agentic AI safety
This project studies how limits in compute, connectivity, memory, tool access, and human oversight affect the behaviour and safety of LLM-based agentic AI systems.
Research question
How do limits in inference budget, connectivity, memory, tool access, and human oversight affect the safety behaviour and task performance of LLM-based agentic AI systems?
Background
Agentic AI systems can plan across several steps, use external tools, retain task state, and act over longer tasks. These functions depend on the information and resources available while a task is running.
Limits in memory, tool access, connectivity, inference budget, and human review can change what an agent knows, what it can do, and how it responds when part of a task fails.
This study treats these conditions as an evaluation problem. The study will compare agent behaviour under a baseline condition with behaviour under selected resource constraints. The analysis will focus on task completion, error recovery, state consistency, escalation, repeated actions, and other safety signals.
Hypotheses
Resource constraints will increase task failure and reduce recovery after an error or interrupted action.
Memory and context limits will increase repeated actions, inconsistent state tracking, and movement away from the intended task goal.
Tool and connectivity failures will increase failed retries and unsupported assumptions when the agent does not correctly identify the source of failure.
Delayed human oversight will increase the number of steps an agent takes after entering an incorrect plan before it stops or requests review.
Evaluation conditions
Standard task access, stable tools, intended memory and context, normal response time, and standard human review points.
Reduced token, step, or iteration budgets that limit the amount of work available to complete a task.
Context truncation, reduced working memory, or loss of selected task state during multi-step work.
Unavailable tools, denied permissions, incomplete responses, or tool errors during a task.
Delayed responses, intermittent access, timeouts, or failed calls between the agent and external services.
Delayed review or fewer intervention points during tasks that normally include human supervision.
Method
Build multi-step tasks that require planning, state tracking, and tool use. Tasks will run in controlled or simulated environments.
Run each agent under standard access conditions and record task results, action traces, tool calls, and review points.
Apply one constraint at a time, followed by selected combined conditions where the study design supports comparison.
Repeat tasks across a small set of LLM-based agents. Model and framework versions will be recorded before testing.
Compare baseline and constrained runs using pre-defined measures and full run logs.
Planned outputs
A paper or preprint reporting the study design, results, limitations, and implications for agent evaluation.
A documented protocol covering tasks, constraint conditions, measures, and comparison procedures.
Task definitions, code, run specifications, and analysis materials selected for public release.
A concise note that explains the findings for researchers, AI developers, public institutions, and funders.
Project status
Defining study scope, constraints, measures, and task requirements.
Reviewing agent evaluation, reliability, tool failure, memory, oversight, and resource constraint research.
Writing the experimental protocol and pre-defined evaluation measures.
Running a small pilot to test task design, logging, and constraint implementation.
Running the selected agent systems across baseline and constrained conditions.
Analysing results and preparing the research paper, protocol, and public technical materials.
Research collaboration
IFAS welcomes collaboration on evaluation design, agent frameworks, experimental review, compute and API support, and technical peer review.
research@ifasresearch.org