← Back
AI Engineer September 26, 2026 21m

Long-Horizon Agents Need Experiments, Not Just Prompts — Erina Karati

Read full transcript 17 segments
  1. Hello everyone. I'm Iryna, a Hello everyone. I'm Iryna, a former engineer at former engineer at former engineer at Microsoft and Supercell. And Microsoft and Supercell. And Microsoft and Supercell. And today I want to today I want to today I want to talk about talk about talk about automated automated automated research in a research in a research in a multi-agent AI multi-agent AI village. I'll village. I'll village. I'll use use use a video game like AI a video game like AI Village as an example, Village as an example, Village as an example, but the broader question that but the broader question that but the broader question that I I I think think think a lot of AI engineers face is a lot of AI engineers face is a lot of AI engineers face is this. How do you this. How do you this. How do you evaluate and evaluate and evaluate and improve improve improve agents that agents that agents that maintain state maintain state maintain state over over over time? Before time? Before time? Before moving on to the level of moving on to the level of moving on to the level of automated automated automated research, I want to research, I want to research, I want to talk a little about the talk a little about the talk a little about the Paradox project. We Paradox project. We Paradox project. We developed the Paradox project developed the Paradox project developed the Paradox project in in in Supercell's AI Innovation Lab Supercell's AI Innovation Lab Supercell's AI Innovation Lab with my colleague with my colleague with my colleague Arunachalam Arunachalam Arunachalam Manikandan. We have Manikandan. We have Manikandan. We have created a modular created a modular created a modular AI framework that AI framework that AI framework that allows any allows any allows any developer developer developer to integrate to integrate to integrate intelligent intelligent intelligent autonomous agents into a autonomous agents into a autonomous agents into a video game that can video game that can video game that can interact, interact, interact, compete, or compete, or compete, or collaborate with collaborate with collaborate with other players or other players or other players or agents, becoming agents, becoming agents, becoming dynamic game dynamic game dynamic game companions. So, companions. So, companions. So, to give examples to give examples to give examples of what of what of what these agents can do, these agents can do, these agents can do, they can move they can move they can move purposefully. They purposefully. They purposefully. They can head to can head to can head to any place or any place or any place or character, guided by character, guided by character, guided by their own memories, their own memories, their own memories, emotions, or emotions, or emotions, or curiosity. These agents curiosity. These agents curiosity. These agents can interact can interact can interact with the world. They with the world. They with the world. They can pick up can pick up can pick up objects, throw them objects, throw them objects, throw them anywhere, and anywhere, and anywhere, and are also aware of the are also aware of the are also aware of the context of their context of their context of their surroundings, such as

  2. surroundings, such as surroundings, such as objects or other objects or other objects or other characters or characters or characters or agents. I also want to agents. I also want to agents. I also want to point out that point out that point out that game developers game developers game developers can add new can add new can add new actions for these agents in actions for these agents in actions for these agents in our framework, our framework, our framework, not just throwing or not just throwing or not just throwing or placing objects. placing objects. Agents also Agents also obviously react to what is obviously react to what is obviously react to what is happening happening happening around them, and these events around them, and these events around them, and these events influence their influence their influence their own beliefs own beliefs own beliefs and emotions “on the fly.” and emotions “on the fly.” And of course, the picture And of course, the picture would be incomplete if would be incomplete if would be incomplete if agents couldn't agents couldn't agents couldn't start conversations, start conversations, start conversations, right? In this right? In this right? In this scenario, agents scenario, agents scenario, agents can approach can approach can approach other agents or other agents or other agents or even the player, which even the player, which even the player, which makes the game more alive, and makes the game more alive, and makes the game more alive, and these conversations these conversations these conversations are stored in their are stored in their are stored in their memory, influencing memory, influencing memory, influencing their emotions, their emotions, their emotions, beliefs, or goals. beliefs, or goals. And together, these agents And together, these agents make up our make up our make up our multi-agent multi-agent multi-agent framework. Um, yes.

  3. framework. Um, yes. Yes. One second. So, Yes. One second. So, the architecture was the architecture was the architecture was deliberately built deliberately built deliberately built with preservation in mind. The with preservation in mind. The with preservation in mind. The first important first important first important part was the part was the part was the individual individual individual memory for each memory for each memory for each agent. Each agent agent. Each agent agent. Each agent has its own has its own has its own memory space, memory space, memory space, supported by supported by supported by RAG technology. Therefore, RAG technology. Therefore, RAG technology. Therefore, the memory was not the memory was not the memory was not mixed between mixed between mixed between agents. Second, we agents. Second, we agents. Second, we tracked emotions tracked emotions tracked emotions as a small vector. as a small vector. as a small vector. After an event or After an event or After an event or conversation, the system conversation, the system conversation, the system could update could update could update the meanings of the meanings of the meanings of emotions such as joy, emotions such as joy, emotions such as joy, sadness, fear, anger, sadness, fear, anger, sadness, fear, anger, or disgust. Third, or disgust. Third, or disgust. Third, agents had agents had agents had trust scores for trust scores for trust scores for other agents and the other agents and the other agents and the player. You can player. You can player. You can think of it as a think of it as a think of it as a trust matrix, trust matrix, trust matrix, essentially. That is, after essentially. That is, after essentially. That is, after interaction, the language interaction, the language interaction, the language model (LM) decides whether model (LM) decides whether model (LM) decides whether the level of trust should the level of trust should the level of trust should increase, decrease, or increase, decrease, or increase, decrease, or not change at all not change at all not change at all . And . And fourth, each fourth, each fourth, each memory receives an memory receives an memory receives an importance rating.

  4. importance rating. importance rating. To explain this better To explain this better , let's say you had , let's say you had , let's say you had dinner a few days dinner a few days dinner a few days ago, you probably don't ago, you probably don't ago, you probably don't remember exactly what you remember exactly what you remember exactly what you ate, right? But ate, right? But ate, right? But if a murder had happened a few days ago if a murder had happened a few days ago if a murder had happened a few days ago , you , you , you would definitely would definitely would definitely remember it. So, the remember it. So, the remember it. So, the agent or language agent or language agent or language model will assess model will assess model will assess the importance of the event, and the importance of the event, and the importance of the event, and if it exceeds a if it exceeds a if it exceeds a certain threshold, it certain threshold, it certain threshold, it will store that will store that will store that memory in a separate memory in a separate memory in a separate cache so that important cache so that important cache so that important context can be context can be context can be better retrieved better retrieved better retrieved later. Here is an example of later. Here is an example of later. Here is an example of how it works. We are how it works. We are how it works. We are going to ask going to ask going to ask one of the characters to one of the characters to one of the characters to go on a picnic with us. go on a picnic with us. Here, our character Here, our character Blossom decides to Blossom decides to Blossom decides to grab some pastries and go grab some pastries and go grab some pastries and go to the picnic area because to the picnic area because to the picnic area because we we we asked her to. Keep in asked her to. Keep in asked her to. Keep in mind that while she's mind that while she's mind that while she's talking in the background, she's talking in the background, she's talking in the background, she's planning this entire planning this entire planning this entire sequence of actions sequence of actions sequence of actions to execute. And when to execute. And when to execute. And when we talk to her we talk to her we talk to her later, she will also later, she will also later, she will also respond in the respond in the respond in the context of the situation.

  5. context of the situation. Yes. But here an Yes. But here an interesting interesting interesting problem began. As you problem began. As you problem began. As you saw in the last saw in the last saw in the last example, for example, for example, for short-term short-term short-term gameplay, gameplay, gameplay, our architecture our architecture our architecture worked quite worked quite worked quite well. The character could well. The character could well. The character could make a plan, make a plan, make a plan, move around, move around, move around, talk, remember talk, remember talk, remember recent interactions, and recent interactions, and recent interactions, and respond to us or respond to us or respond to us or other characters. But other characters. But other characters. But over longer distances, over longer distances, over longer distances, we noticed that we noticed that we noticed that social social social coherence began coherence began coherence began to weaken. In this to weaken. In this to weaken. In this example, one agent example, one agent example, one agent spreads a rumor about a spreads a rumor about a spreads a rumor about a mango sale to mango sale to mango sale to another agent. And that another agent. And that another agent. And that agent receives this agent receives this agent receives this information and information and information and tells the tells the tells the other about it. Later, after other about it. Later, after other about it. Later, after a series of events that have a series of events that have a series of events that have occurred in between, when occurred in between, when occurred in between, when the player asks the player asks the player asks one of the agents about the one of the agents about the one of the agents about the mango, it doesn't quite mango, it doesn't quite mango, it doesn't quite preserve the preserve the preserve the context we context we context we expected, or expected, or expected, or provide the context provide the context provide the context we want to get. And this is where we want to get. And this is where we want to get. And this is where things things things start to get confusing, start to get confusing, start to get confusing, quite naturally.

  6. quite naturally. The system may The system may remember a general remember a general remember a general topic but lose the topic but lose the topic but lose the source of that topic. source of that topic. source of that topic. A rumor can become A rumor can become A rumor can become a certainty, instead of a certainty, instead of a certainty, instead of remaining remaining remaining just a rumor. The agent just a rumor. The agent just a rumor. The agent may present this as may present this as may present this as fact. Or the agent may fact. Or the agent may fact. Or the agent may know the fact but not know the fact but not know the fact but not mention it mention it mention it when creating a plan of when creating a plan of when creating a plan of action. So action. So action. So the question became: the question became: the question became: how do we improve a how do we improve a how do we improve a multi-agent multi-agent multi-agent system for long-term system for long-term system for long-term social behavior social behavior , not just for a single , not just for a single , not just for a single response? And this is where response? And this is where response? And this is where we wanted to involve we wanted to involve we wanted to involve auto research. As you all auto research. As you all auto research. As you all know, a few know, a few know, a few months ago months ago months ago Karpaty published Karpaty published Karpaty published our auto research, and it our auto research, and it our auto research, and it instantly aroused instantly aroused instantly aroused great curiosity in us. great curiosity in us. Maybe we can Maybe we can make the system make the system make the system conduct conduct conduct experiments experiments experiments on itself, and can on itself, and can on itself, and can we use that for we use that for we use that for our system? So, our system? So, our system? So, we realized that we realized that we realized that instead of manually instead of manually instead of manually configuring configuring configuring prompts or prompts or prompts or watching one watching one watching one nice demo, we nice demo, we nice demo, we could define a could define a could define a set of scenarios, set of scenarios, set of scenarios, run agents, run agents, run agents, collect traces, collect traces, collect traces, evaluate behavior, and evaluate behavior, and evaluate behavior, and change a small change a small change a small portion of the policy, portion of the policy, portion of the policy, keeping only those keeping only those keeping only those changes that really changes that really changes that really improve the outcome improve the outcome . And this is where we . And this is where we . And this is where we try to combine the try to combine the try to combine the Paradox project with Paradox project with Paradox project with external external external research. So at research. So at research. So at this stage, our this stage, our Paradox multi-agent framework is more Paradox multi-agent framework is more like a like a like a lab bench, and lab bench, and lab bench, and external research

  7. external research external research becomes an becomes an becomes an experimental experimental experimental loop around it. loop around it. And importantly, it's And importantly, it's not just about not just about not just about improving RAG search. The improving RAG search. The broader goal is broader goal is to optimize the to optimize the to optimize the agent protocol. agent protocol. For example, how agents For example, how agents record memories, record memories, record memories, retrieve them, retrieve them, retrieve them, report report report uncertainty, uncertainty, uncertainty, update trust, update trust, update trust, cite sources, and cite sources, and cite sources, and replan actions replan actions replan actions based on new based on new based on new facts. Um, yes. In facts. Um, yes. In facts. Um, yes. In this context, our Oh this context, our Oh , yes. In this , yes. In this , yes. In this context, external context, external context, external research is not just research is not just research is not just another agent in the village, as another agent in the village, as another agent in the village, as I said. This is I said. This is I said. This is a metasystem outside the a metasystem outside the a metasystem outside the village. village. village. Villagers, of course, have Villagers, of course, have Villagers, of course, have their own local their own local their own local perspectives. They perspectives. They perspectives. They only know what they have only know what they have only know what they have seen, heard, seen, heard, seen, heard, remembered, or remembered, or remembered, or inferred, inferred, inferred, because there because there because there is no shared is no shared is no shared memory base between them. Information memory base between them. Information memory base between them. Information is only transmitted is only transmitted is only transmitted when other when other when other agents report it correctly agents report it correctly agents report it correctly . The . The . The external external external research level research level research level does a different job here.

  8. does a different job here. It reads full It reads full execution traces execution traces , compares events to , compares events to , compares events to the script, evaluates the script, evaluates the script, evaluates behavior, and behavior, and behavior, and suggests limited suggests limited suggests limited changes to the agents' protocol changes to the agents' protocol changes to the agents' protocol or cognitive or cognitive or cognitive policies. policies. Then he Then he restarts restarts restarts the script and checks the script and checks the script and checks the behavior at the the behavior at the the behavior at the societal level, has societal level, has societal level, has it become better? This is the it become better? This is the it become better? This is the key shift key shift key shift we were trying we were trying we were trying to achieve. So, we were to achieve. So, we were to achieve. So, we were no longer evaluating no longer evaluating no longer evaluating a single response, we a single response, we a single response, we were evaluating the entire were evaluating the entire were evaluating the entire execution cycle. And this is what execution cycle. And this is what execution cycle. And this is what one one one such cycle would look like. such cycle would look like. such cycle would look like. For example, first For example, first For example, first we define a we define a we define a control scenario control scenario , which I will , which I will , which I will explain in more detail explain in more detail explain in more detail later. For example, later. For example, later. For example, one agent one agent one agent learns of a learns of a learns of a public fact or public fact or public fact or hears a rumor. This could hears a rumor. This could hears a rumor. This could be a control be a control be a control scenario. Then we scenario. Then we scenario. Then we run the simulation run the simulation . During our work, we . During our work, we . During our work, we collect collect collect structured data: structured data: structured data: observations, observations, observations, conversations, memory records conversations, memory records conversations, memory records , search , search , search queries, queries, queries, belief updates—anything that is belief updates—anything that is belief updates—anything that is relevant to us in relevant to us in relevant to us in this case. Then this case. Then this case. Then we evaluate this we evaluate this we evaluate this behavior. Did behavior. Did the information spread as we the information spread as we expected? Is the expected? Is the source noted, i.e. does source noted, i.e. does the agent remember who the agent remember who the agent remember who started the rumor? Has started the rumor? Has the uncertainty remained the the uncertainty remained the same? Did same? Did same? Did the agents act based on what the agents act based on what they actually they actually they actually knew? And then the level of knew? And then the level of knew? And then the level of external external external research suggests a

  9. research suggests a research suggests a small change in small change in small change in policy. And this is policy. And this is policy. And this is important. Of course, this important. Of course, this important. Of course, this shouldn't result shouldn't result shouldn't result in rewriting the in rewriting the in rewriting the entire program. He entire program. He entire program. He should only edit the should only edit the should only edit the controlled controlled controlled policy surface. And policy surface. And policy surface. And then we launch then we launch then we launch again. If the estimate again. If the estimate again. If the estimate improves and improves and improves and the constraints the constraints the constraints remain, we remain, we remain, we leave the improvement leave the improvement . And if not, we just . And if not, we just . And if not, we just go back. go back. Regarding controlled Regarding controlled scenarios: scenarios: scenarios: scenario design is important scenario design is important scenario design is important because social because social because social behavior is generally behavior is generally behavior is generally somewhat blurred. somewhat blurred. somewhat blurred. Simply letting Simply letting Simply letting agents roam the agents roam the agents roam the environment may environment may environment may look cool and look cool and look cool and provide interesting provide interesting provide interesting interactions, but it is very interactions, but it is very interactions, but it is very difficult to assess whether difficult to assess whether difficult to assess whether the system has truly the system has truly the system has truly improved. This is why improved. This is why improved. This is why we believe that we believe that controlled controlled scenarios are needed. For example, scenarios are needed. For example, scenarios are needed. For example, one scenario might one scenario might one scenario might test the spread of a test the spread of a test the spread of a public fact.

  10. public fact. public fact. Suppose Agent A Suppose Agent A Suppose Agent A learns that learns that learns that the bakery the bakery the bakery will close tomorrow. Will the will close tomorrow. Will the right agents find out about this? Do right agents find out about this? Do they remember who they remember who they remember who said what? Do said what? Do said what? Do they change their they change their they change their plans based on this plans based on this plans based on this fact? Another scenario fact? Another scenario fact? Another scenario could test could test could test the uncertainty of rumors the uncertainty of rumors . Suppose agent A . Suppose agent A . Suppose agent A hears that agent C may hears that agent C may hears that agent C may leave the village. When leave the village. When leave the village. When this rumor spreads, does this rumor spreads, does this rumor spreads, does "may "may "may leave" suddenly become "is leave" suddenly become "is leaving," or does leaving," or does leaving," or does it remain "may it remain "may it remain "may leave"? That is, is leave"? That is, is leave"? That is, is this becoming a fact or this becoming a fact or this becoming a fact or remains a rumor? remains a rumor? Another scenario could Another scenario could test test test redevelopment. redevelopment. The group has a plan, but The group has a plan, but one agent one agent one agent learns that learns that learns that their chosen route their chosen route their chosen route is blocked. Do is blocked. Do is blocked. Do agents update this agents update this agents update this information and information and information and communicate it to communicate it to communicate it to each other to avoid each other to avoid each other to avoid failed plans or failed plans or failed plans or actions? It's not that actions? It's not that actions? It's not that these specific these specific these specific scenarios are scenarios are scenarios are universal. We universal. We universal. We want to say that want to say that want to say that for for for agents to behave over long agents to behave over long agents to behave over long distances, distances, distances, sets of scenarios are needed.

  11. sets of scenarios are needed. Going back to Going back to our our our mango example: after running mango example: after running mango example: after running one of our one of our one of our research cycles, when a research cycles, when a research cycles, when a player finally player finally player finally asked one of the asked one of the asked one of the agents about a discount on agents about a discount on agents about a discount on mangoes, we found that mangoes, we found that mangoes, we found that this time the agent was able to this time the agent was able to this time the agent was able to respond in respond in respond in context. Compared context. Compared context. Compared to last time. Uh, to last time. Uh, to last time. Uh, yes. And for this yes. And for this yes. And for this report, the exact report, the exact report, the exact formula, in our formula, in our formula, in our opinion, is less important opinion, is less important opinion, is less important than the structure of the than the structure of the than the structure of the evaluation system. evaluation system. You don't need one You don't need one vague metric vague metric like "agent quality like "agent quality ." This will hide all the ." This will hide all the ." This will hide all the interesting failures. interesting failures. interesting failures. Instead, you Instead, you Instead, you need a need a need a balanced balanced balanced scorecard. To scorecard. To scorecard. To spread information, spread information, spread information, you can measure you can measure you can measure reach. For example, reach. For example, reach. For example, how many agents how many agents how many agents know the fact after n know the fact after n know the fact after n steps. To verify steps. To verify steps. To verify sources, measure the sources, measure the sources, measure the retention of retention of retention of information about the information about the information about the source among source among source among agents who agents who agents who know about it. How many know about it. How many know about it. How many of them remember of them remember of them remember where it came from, where it came from, where it came from, etc. For rumors, one can etc. For rumors, one can etc. For rumors, one can measure measure measure the persistence of the persistence of the persistence of uncertainty and the uncertainty and the uncertainty and the level of false level of false level of false confidence. For confidence. For confidence. For planning, you can planning, you can planning, you can measure measure measure the sequence of actions and the sequence of actions and the sequence of actions and the time for the time for the time for replanning. And for replanning. And for replanning. And for privacy, you can privacy, you can privacy, you can measure measure measure restraint. This is restraint. This is restraint. This is important because important because important because optimizing just optimizing just optimizing just one metric can one metric can one metric can lead to

  12. lead to lead to unwanted behavior. unwanted behavior. unwanted behavior. Because, say, if you Because, say, if you Because, say, if you only optimize for only optimize for only optimize for distribution, agents distribution, agents distribution, agents can learn to can learn to can learn to overshare overshare overshare everything at once. And if you everything at once. And if you everything at once. And if you optimize just the optimize just the optimize just the memory playback, memory playback, memory playback, you can create even you can create even you can create even noisier, well, noisier, well, noisier, well, memories. So this memories. So this memories. So this scoring system scoring system scoring system keeps the system within keeps the system within keeps the system within the bounds of honesty and prevents the the bounds of honesty and prevents the auto-research agent from auto-research agent from manipulating manipulating manipulating the system to the system to the system to increase just increase just increase just one indicator. one indicator. one indicator. Another important Another important Another important engineering lesson engineering lesson engineering lesson we learned from this we learned from this we learned from this project is that it's project is that it's project is that it's important important important to keep the to keep the to keep the editing surface very editing surface very editing surface very limited. The outer limited. The outer limited. The outer research layer research layer research layer should not have permission should not have permission should not have permission to arbitrarily to arbitrarily to arbitrarily rewrite the entire rewrite the entire rewrite the entire codebase. codebase. Instead, it is Instead, it is crucial to capture crucial to capture crucial to capture the environment, scenarios, the environment, scenarios, the environment, scenarios, and metrics. So we and metrics. So we and metrics. So we only open up the only open up the only open up the part of the system that part of the system that part of the system that we actually want to we actually want to we actually want to optimize. In the optimize. In the optimize. In the Paradox project, this Paradox project, this Paradox project, this meant things like meant things like meant things like write-to- write-to- write-to- memory policies, memory policies, memory policies, search policies, search policies, search policies, communication communication communication prompts, trust prompts, trust prompts, trust and belief rules, and belief rules, and belief rules, source attribution, source attribution, source attribution, replanning triggers, and so on.

  13. replanning triggers, and so on. This gives the search process This gives the search process room to room to room to improve its behavior improve its behavior , but also prevents , but also prevents evaluation manipulation, as we evaluation manipulation, as we noted earlier. And noted earlier. And noted earlier. And this is the this is the this is the difference between a language difference between a language difference between a language model that writes model that writes model that writes random patches and a random patches and a random patches and a language model that language model that language model that searches within a searches within a searches within a controlled controlled controlled space for policies. Here are space for policies. Here are space for policies. Here are examples of the changes I examples of the changes I examples of the changes I want this cycle want this cycle want this cycle to explore. If to explore. If to explore. If attribution of sources attribution of sources attribution of sources disappears, a disappears, a disappears, a policy change may be to policy change may be to keep the source in keep the source in memory and record memory and record memory and record access rights and a summary access rights and a summary . If rumors become . If rumors become . If rumors become facts, a facts, a facts, a policy change may policy change may policy change may involve involve involve maintaining maintaining maintaining confidence, confidence, confidence, labeling labeling labeling information as information as information as first-hand first-hand first-hand or second-hand, or second-hand, or second-hand, and requiring and requiring and requiring caution when caution when caution when relaying uncertain relaying uncertain relaying uncertain claims. If claims. If claims. If public facts public facts public facts remain remain remain local, a local, a local, a policy change might policy change might policy change might consist of consist of consist of classifying classifying classifying useful public useful public useful public facts differently and encouraging facts differently and encouraging facts differently and encouraging agents to proactively agents to proactively agents to proactively share important share important share important evidence. The key is that these are evidence. The key is that these are evidence. The key is that these are small changes small changes small changes to the agent protocol, to the agent protocol, to the agent protocol, but they can have a but they can have a but they can have a large impact on large impact on large impact on the behavior of the behavior of the behavior of multi-agent multi-agent multi-agent systems at the systems at the systems at the community level. Here I want community level. Here I want community level. Here I want to be careful with to be careful with to be careful with our claims, our claims, our claims, because I cannot say because I cannot say because I cannot say that the system that the system that the system has improved overall without the has improved overall without the cycle's repeated results. We're

  14. cycle's repeated results. We're trying to say trying to say trying to say that this is the right that this is the right that this is the right surface for an surface for an surface for an external loop of external loop of external loop of research research research because it's because it's because it's small enough to small enough to small enough to control, but control, but control, but at the same time at the same time at the same time rich enough to change rich enough to change rich enough to change social behavior social behavior social behavior at least to some at least to some at least to some extent. And the biggest extent. And the biggest extent. And the biggest lesson for me, lesson for me, lesson for me, perhaps, was that perhaps, was that perhaps, was that memory alone is memory alone is memory alone is not enough. You not enough. You not enough. You can add RAG can add RAG memory to an agent, but memory to an agent, but memory to an agent, but still not get the still not get the still not get the desired desired desired long-term long-term long-term behavior, because behavior, because behavior, because agents sometimes agents sometimes agents sometimes need to know need to know need to know where this where this where this information came from. You information came from. You information came from. You need to keep track of need to keep track of need to keep track of whether it was whether it was whether it was first-hand or first-hand or first-hand or second-hand, second-hand, second-hand, verified or verified or verified or unverified. Sometimes unverified. Sometimes unverified. Sometimes it is necessary it is necessary it is necessary to separate "raw to separate "raw " episodic memories " episodic memories " episodic memories from what the agent from what the agent from what the agent believes now, and believes now, and believes now, and test behavior test behavior test behavior through scenarios, rather than through scenarios, rather than through scenarios, rather than just intuition. just intuition. So another lesson So another lesson is that is that is that rolling back changes is also rolling back changes is also rolling back changes is also mandatory. When you mandatory. When you mandatory. When you optimize optimize optimize social behavior, social behavior, social behavior, the change may improve the change may improve the change may improve one thing but hurt one thing but hurt one thing but hurt another. For example, another. For example, another. For example, a policy that is more likely a policy that is more likely a policy that is more likely to disseminate public to disseminate public to disseminate public facts may accidentally facts may accidentally facts may accidentally reveal private reveal private reveal private information. Policies information. Policies information. Policies that improve that improve that improve memory retrieval memory retrieval memory retrieval may increase the may increase the may increase the use of use of use of stale data.

  15. stale data. stale data. Therefore, the cycle should Therefore, the cycle should Therefore, the cycle should work like a work like a work like a ratchet mechanism. ratchet mechanism. ratchet mechanism. Try the change, Try the change, Try the change, evaluate it, and evaluate it, and evaluate it, and only leave when only leave when only leave when performance performance performance has improved and the has improved and the has improved and the safeguards are working safeguards are working . We are sure that this . We are sure that this . We are sure that this applies not only to applies not only to applies not only to gaming agents. Although gaming agents. Although gaming agents. Although I gave an example with a I gave an example with a I gave an example with a game settlement, game settlement, game settlement, we believe this is we believe this is we believe this is also important for others, also important for others, also important for others, say, say, say, support agents. support agents. support agents. Support agents need Support agents need Support agents need to know where to know where to know where every every every rule update is coming from, rule update is coming from, rule update is coming from, right? And does right? And does right? And does it replace the previous it replace the previous it replace the previous answer? answer? Personal Personal assistants, assistants, assistants, for example, need to for example, need to for example, need to remember their remember their remember their commitments and commitments and commitments and adjust them if adjust them if adjust them if the user wants to the user wants to the user wants to change those commitments. change those commitments. Research Research agents need agents need agents need sources, citations, sources, citations, sources, citations, working through working through working through contradictions, and contradictions, and contradictions, and updating hypotheses. updating hypotheses. Programming agents Programming agents need long-term need long-term need long-term context about context about context about tasks, files, tasks, files, tasks, files, colleagues, and changes in colleagues, and changes in colleagues, and changes in requirements. Workflow agents requirements. Workflow agents need need access control, access control, access control, case transfer, and case transfer, and case transfer, and rescheduling when rescheduling when rescheduling when conditions change. All of these conditions change. All of these conditions change. All of these systems have the systems have the systems have the same same same fundamental fundamental fundamental problem. They problem. They problem. They maintain the condition maintain the condition maintain the condition for a long for a long for a long time. And this state time. And this state time. And this state affects future affects future affects future actions. Therefore, we actions. Therefore, we actions. Therefore, we suggest suggest suggest using using using control scenarios control scenarios control scenarios and and and behavior scorecards. So, a behavior scorecards. So, a behavior scorecards. So, a short recipe for short recipe for short recipe for long-acting agents:

  16. long-acting agents: long-acting agents: capture the system, capture the system, capture the system, identify scenarios, identify scenarios, identify scenarios, record traces record traces , evaluate behavior, , evaluate behavior, , evaluate behavior, and expose only a and expose only a and expose only a small portion of small portion of small portion of policy settings. policy settings. Search Search through these changes, through these changes, through these changes, keeping only those keeping only those keeping only those that pass your that pass your that pass your measurements. We measurements. We measurements. We believe this is an believe this is an believe this is an engineering pattern engineering pattern engineering pattern that makes sense. For that makes sense. For long-running agents, the main long-running agents, the main question, we question, we question, we believe, is believe, is believe, is whether the system becomes whether the system becomes whether the system becomes better during better during better during controlled controlled controlled runs? In runs? In runs? In conclusion, Project Paradox conclusion, Project Paradox conclusion, Project Paradox began as an attempt began as an attempt began as an attempt to make game to make game to make game agents "alive" in a 3D agents "alive" in a 3D world, but the deeper world, but the deeper world, but the deeper problem wasn't the problem wasn't the problem wasn't the animation or dialogue. animation or dialogue. The problem was the state: The problem was the state: which agents knew what, which agents knew what, which agents knew what, who told whom what, what who told whom what, what who told whom what, what was true and what was was true and what was was true and what was outdated, and were the outdated, and were the outdated, and were the agents acting according to agents acting according to agents acting according to their memory? Our their memory? Our their memory? Our research gave us a research gave us a research gave us a way to approach way to approach way to approach this more this more this more systematically: not systematically: not systematically: not to trust a single to trust a single demo and not to demo and not to endlessly tweak prompts endlessly tweak prompts manually, but to conduct manually, but to conduct manually, but to conduct controlled controlled controlled experiments, experiments, experiments, keeping only those keeping only those keeping only those changes that passed changes that passed changes that passed the test. Agents with the test. Agents with the test. Agents with long-term long-term long-term planning need planning need planning need experiments, not experiments, not experiments, not just hints, and I just hints, and I just hints, and I hope that is the hope that is the hope that is the main conclusion of main conclusion of main conclusion of my talk. And yes, my talk. And yes, my talk. And yes, please please please contact us. We contact us. We contact us. We would be happy would be happy would be happy to chat if to chat if to chat if you have any questions.

  17. you have any questions. you have any questions. Thank you very much for your attention. Thank you very much for your attention. Thank you very much for your attention. Yes.

Summary

This talk explores automated research in multi-agent AI villages, using AI Village as an example to address the challenge of improving agents that maintain state over time. Key concepts include a modular AI framework for intelligent autonomous agents, their ability to interact, learn from memories and emotions, and communicate, all built with preservation in mind through individual memory, emotion tracking, and trust scores. The practical takeaway is a robust architecture for dynamic, stateful AI agents that can evolve within a game environment.

View original episode ↗