Challenges in evaluating and improving long-horizon AI agents
Her research primarily uses AI game village simulations as a context to observe and test these agents, where their interactions and memories persist over time.
AI Engineer Erina Karati outlines the challenges of long-horizon AI agents and introduces Auto Research as a solution.
Her research primarily uses AI game village simulations as a context to observe and test these agents, where their interactions and memories persist over time.
Erina Karati and Arunachalam Manikandan developed Project Paradox at Supercell's AI Innovation Lab.
The framework offers a modular design that allows developers to integrate intelligent autonomous agents into video games, enabling them to interact, compete, or cooperate with other players or agents.
These agents are equipped with internal memory, emotional states, and curiosity, which are designed to guide their actions and decisions within the game environment, supporting custom actions beyond basic object manipulation.
The agents in Project Paradox are designed with a stateful system, where each agent possesses its own RAG-backed memory namespace, ensuring that memories do not blend or bleed between different agents.
Emotional states are tracked using a vector encompassing joy, sadness, fear, anger, and disgust, influencing agent behavior.
A trust matrix is also employed to keep track of belief scores between individual agents and players, while an importance score mechanism helps cache significant events for more efficient retrieval and recall.
Agents frequently lose track of the original source of information when rumors propagate, leading to instances where rumors are mistakenly elevated to factual status.
Furthermore, agents often fail to integrate known facts into their long-term action plans, impacting the coherence and realism of their extended behaviors within the simulated environment.
To address the long-term consistency issues, the team aims to enhance multi-agent systems by moving beyond single-response evaluation and exploring autonomous experimentation.
Inspired by Andrej Karpathy's auto-research concept, they are investigating methods to enable the system to conduct its own experiments, automating the improvement process for complex, long-running social behaviors.
This new approach replaces manual prompt tuning with automated scenario testing, utilizing Project Paradox as a 'lab bench' for multi-agent frameworks.
Auto Research functions as the experimental loop, optimizing agent protocols related to memory, communication, uncertainty, and replanning, rather than merely enhancing RAG retrieval.
Evaluating behavior at the society level, the Auto Research layer operates outside the main simulation to score full run traces against scenario ground truth.
It scores agent behavior and proposes constrained changes to the agent protocol or cognitive policy, then reruns scenarios to verify if society-level behavior has improved, as individual villagers only possess local perspectives and lack a common memory database.
Public fact diffusion tests if agents accurately learn and plan based on disseminated facts, while rumor uncertainty checks if 'might' statements remain possibilities rather than becoming hard facts.
Replanning scenarios assess agents' ability to update actions when their intended routes are obstructed, ensuring adaptive behavior.
A balanced scorecard, which tracks reach and source retention, prevents the hidden failures that arise when relying on vague metrics like agent quality.
Such scorecards include measurements for reach (diffusion), source retention (provenance), and uncertainty preservation (rumors), as over-optimizing a single metric can lead to undesirable behaviors like oversharing or noisy memories.
A key engineering lesson learned is to keep the editable surface for AI agents very small, preventing the auto research layer from arbitrarily rewriting the entire codebase.
This involves freezing harnesses, scenarios, and metrics to maintain stability, while only exposing specific policy parameters like memory writing, communication, and trust rules to ensure changes are made within a controlled policy space rather than through random patches.
Preserving source attribution or storing confidence markers allows agents to hedge uncertain claims, demonstrating that minor adjustments to agent protocols can have significant effects on societal-level behaviors.
For instance, preserving source attribution helps prevent memory loss regarding origin details, storing confidence markers aids agents in hedging uncertain claims, and classifying public facts distinctly improves the proactive sharing of evidence, all contributing to more robust and accurate social dynamics.
Erina Karati emphasized that RAG memory alone is insufficient to guarantee desired long-horizon behavior in agents, as they need to differentiate between firsthand, secondhand, verified, and uncertain information.
It is crucial to separate raw episodic memories from current beliefs and to test behavioral changes through controlled scenarios rather than relying on subjective assessments, while implementing a rollback mechanism is necessary to counteract unintended side effects of optimizations.
The challenges encountered with long-horizon agents are universal across various agent types, including support agents, personal assistants, research agents, and workflow agents.
All these systems share the fundamental problem of maintaining state over time, where that state directly influences future actions, making the framework applicable to diverse AI applications.
This involves freezing the harness, defining clear scenarios, logging traces, scoring actual agent behavior, exposing and searching over a limited policy surface, and only retaining changes that prove effective through rigorous measurement, ensuring systematic improvement through controlled experiments.
Answers come from the transcript, with the exact spot cited.
Want the next article from AI Engineer?
When AI Engineer publishes, we'll write it up like the one you just read and email it to you.
AI Engineer published 87 in the last 7 days.
Akamai Functions Achieve Zero Cold Starts for AIAI Engineeryesterday · 22:37 · 205 views · Created yesterday
From Laptop to Pipeline: Scaling AI AgentsAI Engineeryesterday · 19:39 · 69 views · Created yesterday
Stop Prompting AI: Codify Rules InsteadAI Engineer2 days ago · 15:55 · 77 views · Created 2 days ago