How Should Agents Train Online from Deployment Experience?
- 1MIT
- 2UIUC
- 3Stanford University
- 4Cornell University
- 5MIT-IBM Watson AI Lab
October 2026
Recently, an increasing body of research has focused on how LLM agents can learn from their own interactive experiences post-deployment. For instance, after a Web Agent attempts several online shopping tasks, these experiences can be utilized to avoid previous mistakes and boost success rates in future attempts. This paradigm is commonly referred to as test-time learning or learning from experience.
Back in March this year, the dominant approach to test-time learning was memory: storing past experiences in a memory module and retrieving relevant memories into the context for future queries[1,2,3]. Around that time, we were thinking: Can we learn from deployment experience via updating model parameters? So our initial goal for this project was to do training on deployment experience. Over time, several concurrent works emerged that also explored test-time learning by updating model weights[4,5,6]. Yet, we felt that a critical piece was still missing in these works.
How did these existing works demonstrate the effectiveness of their algorithms? They typically plotted training steps on the x-axis and accuracy on the y-axis; a steady increase in accuracy was taken as an evidence that the model is learning from deployment experience[4,5,6]. But if test-time learning is evaluated this way, how does it differ from standard post-training? Post-training also shows improving accuracy over training steps. We gradually realized that merely demonstrating "accuracy gains" fails to capture the true essence of test-time learning. Guided by this perspective, we developed Fast and Slow Online Learning (FSOL) (originally referred to as Fast and Slow Reinforcement Learning, FSRL, in our draft), which will be presented at the NeurIPS 2026 TTCL Workshop / LCFM Workshop. In this blog, we highlight three main insights from our work.
1. Test-time Learning Should Be Evaluated by Cumulative Performance
Standard post-training typically operates on a static dataset, using SFT or RL to train the model, and then evaluates the final checkpoint on a test set. Although the model traverses many intermediate checkpoints during training, we ultimately only care about the final checkpoint. Under this paradigm, it does not matter if intermediate checkpoints perform poorly.
Test-time learning, however, is fundamentally different. As illustrated in Figure 1, suppose an agent is deployed in production. When the first batch of queries arrives, the agent generates responses, receives feedback, and updates its parameters based on that feedback. When the second batch arrives, it is handled by the updated model, which receives further feedback and updates again. Consequently, every intermediate checkpoint is immediately deployed to handle real incoming queries, and every interaction directly impacts users. Intermediate performance matters immensely: the correctness or quality of every single interaction must count toward the evaluation.
This insight implies that if an online learner eventually achieves high ceiling accuracy at the expense of making numerous mistakes along the learning path, it still cannot be considered a successful online learner because all those errors are directly experienced by users. Therefore, what we optimize is not merely final performance, but cumulative performance across the entire interaction stream. This aligns naturally with the classic paradigm of online learning[7]: data arrives sequentially as a stream, and evaluation is conducted over the entire stream. Hence, we also use online learning as an umbrella term for this learning-from-deployment-experience paradigm.
Assuming binary rewards (\(0\) or \(1\)), the cumulative accuracy up to the \(N\)-th interaction can be defined as:
$$\mathrm{Acc}(N)=\frac{1}{N}\sum_{i=1}^{N} r_i$$
Equivalently, we can evaluate performance using a metric from online learning[7], cumulative regret:
$$\mathrm{Regret}(N)=\sum_{i=1}^{N}\left(1-r_i\right)$$
Here we define the oracle policy that the online learner is compared with as the policy that always succeeds. Intuitively, cumulative regret measures the total number of errors made since deployment.
This perspective also shifts our understanding of learning curves. In Figure 2, we plot the curves of accuracy versus number of queries for algorithms A and B. Evaluated as offline learners, they might appear equally good since they reach the same final accuracy. But as online learners, Algorithm B clearly makes many more mistakes before "learning" the task, incurring a much higher cost of failure. For an online learner, what truly matters is the area under the accuracy–number of queries curve (this integral is proportional to cumulative accuracy) or equivalently the area between the curve and the y=1 line (equal to cumulative regret), rather than just the accuracy at the final checkpoint.
2. Optimizing Cumulative Performance Requires Learning Fast and Learning Well: our method FSOL
Based on the cumulative objective, a natural decomposition emerges: the agent must achieve high performance ceiling (learning well), and it must reach that performance quickly (learning fast with high sample efficiency). Only by satisfying both conditions can we maximize the area under the accuracy–queries curve.
Existing paradigms generally satisfy only one side of this equation. The first category, memory-based methods[1,2,3], stores past successful experiences so that when a similar query is encountered, retrieved experiences can immediately serve as few-shot examples without requiring large volumes of data for stable gradient updates. Thus, memory excels at fast adaptation. However, memory suffers from a clear ceiling: it is bounded by the model's in-context learning capabilities[8,9,10]. When deeper changes in responses patterns are needed, such as longer reasoning chains or reliable tool-use strategies, simply in-context learning yields poor performance, especially for small models.
The second category, RL, directly updates model weights. The core strength of RL lies in fundamentally changing response patterns through weight updates, unlocking a much higher performance ceiling. Yet RL faces a major practical hurdle: it learns slowly. It usually requires thousands of rollouts. In an online setting, this drawback is compounded by the single-rollout constraint. In standard RL post-training (e.g., GRPO), one prompt can be rolled out 8 times to construct a comparative group. In real-world user interactions, however, each query usually comes with only a single feedback signal[6,11]. A shopping agent cannot place 8 separate orders (incurring duplicate costs), nor can a conversational model output 8 alternative responses and ask users to rate each one. Online RL must operate on a single rollout per query (a constraint overlooked by many existing papers on test-time learning). Consequently, every rollout consumes a real user query, making sample efficiency paramount. Standard RL often yields a learning curve flat at first, rising slowly in the middle, and eventually surpassing memory. While acceptable for final-checkpoint metrics, under a cumulative objective, slow learning means more failures directly impact the user.
Inspired by Neural Episodic Control (NEC)[12,13], a classic work in the RL community, we realized that combining memory (analogous to episodic memory in NEC) and RL offers a path to achieving both speed and depth: Fast and Slow Online Learning (FSOL). The core idea is: Memory handles fast adaptation, while RL manages slow internalization. The complete workflow is shown in Figure 3. Memory allows newly acquired successful experiences to immediately influence the next interaction; RL gradually internalizes interaction experiences into model weights over time.
This combination creates a genuine synergy. Early in training, a pure RL policy is weak and yields a low success rate, generating low-quality or failed rollouts with sparse positive learning signals. Integrating memory enables instant retrieval of past successes, elevating rollout quality and guiding exploration toward high-reward regions. Conversely, as the RL policy grows stronger, it generates higher-quality successful trajectories, which continually enrich the memory bank.
Importantly, FSOL is a general framework combining RL and memory rather than a single specific algorithm. In our experiments, we evaluated PPO[14] and REINFORCE++[15] for RL since they adhere to the single-rollout constraint, and raw action sequences as well as Agent Workflow Memory (AWM)[2] for memory. The empirical results demonstrate that across different memory representations and RL algorithms, FSOL consistently outperforms either component used alone, confirming that its gains do not depend on a specific memory structure or RL variant. We evaluated FSOL across four benchmark domains: Writing[16] (also used in[11]), WebShop[17], ALFWorld[18], and Search-augmented QA[19,20] (Table 1). For each benchmark, training data was arranged into a continuous stream and shuffled by four random seeds, allowing the learner to process each query sequentially without overlap. FSOL achieved superior cumulative accuracy across all four benchmarks. The only exception occurred in the AWM + PPO combination on WebShop, primarily because AWM itself was ineffective on WebShop.
| Memory | WebShop | Search | ALFWorld | Writing | Average | |
|---|---|---|---|---|---|---|
| Zero-shot | Zero-shot | 4.0 ± 0.2 | 15.0 ± 0.0 | 13.0 ± 0.1 | 17.9 ± 0.3 | 12.5 |
| Memory | Raw Memory | 34.0 ± 1.0 | 33.8 ± 0.2 | 25.0 ± 0.5 | 68.7 ± 1.9 | 40.4 |
| AWM | 6.1 ± 0.7 | 33.8 ± 0.2 | 20.3 ± 0.6 | – | – | |
| RL | REINFORCE++ | 38.4 ± 5.6 | 35.8 ± 1.9 | 33.7 ± 3.8 | 79.8 ± 0.4 | 46.9 |
| PPO | 47.7 ± 2.7 | 35.1 ± 0.6 | 40.6 ± 3.9 | 49.7 ± 4.8 | 43.3 | |
| FSOL (Ours) | PPO + AWM | 46.7 ± 2.9 | 37.8 ± 0.2 | 48.8 ± 3.0 | – | – |
| PPO + Raw | 56.3 ± 0.9 | 37.5 ± 0.6 | 50.9 ± 3.2 | 96.3 ± 0.5 | 60.2 | |
| REINFORCE++ + Raw | 52.5 ± 1.5 | 38.8 ± 0.9 | 52.1 ± 2.4 | 97.4 ± 0.3 | 60.2 |
To illustrate these dynamics, consider the Writing benchmark in Figure 4. In this task, the LLM continues a prompt in one of three candidate styles. The model does not know which style will be rewarded; a judge model evaluates whether the response matches the target style. Memory stores responses previously judged as matching the target style, while RL employs REINFORCE++. We can observe that memory adapts rapidly but saturates near a 0.7 accuracy plateau. RL starts slowly but eventually surpasses memory. FSOL matches the fast initial trajectory of memory while continuing to climb beyond memory's plateau toward nearly 100% accuracy. Crucially, FSOL achieves the maximum area under the curve, signifying optimal cumulative performance.
3. Evaluating Online Learning Demands More Than IID Streams
Another overlooked problem is: the real-world interaction stream is usually not IID. In Table 1 we only examine IID streams: gather queries, shuffle them, and let the agent learn online over the sequence. This is also the common practice in most online learning papers[4,5,6] (only[11] considers distribution shifts). However, real-world streams are inherently dynamic. User preferences may shift over time, requiring an intelligent agent to track evolving preferences. Daily tasks may interleave writing, coding, and mathematical reasoning. Alternatively, an agent might shift from coding to web shopping for a period and then return to coding, where it should retain prior coding proficiency rather than relearn from scratch. To capture these deployment realities, we evaluated FSOL not only on stationary streams but also constructed dynamic benchmarks featuring distribution shifts, mixed-task streams, and sequential-task streams.
Distribution Shift
We investigated three distinct sources of distribution shift: reward rules, environment dynamics, and query distributions. In the Writing task, user style preference changed across three phases, rendering previously rewarded behaviors incorrect. In WebShop, the checkout sequence shifted from product page → Buy Now to product page → Add to Cart → Checkout, modifying environment dynamics while keeping queries and underlying rewards constant. In ALFWorld, we introduced query distribution shifts by drawing queries from distinct task subsets across phases.
Why are these settings crucial? Imagine that an agent spent 3,000 queries adapting to an old environment, and requires another 3,000 queries upon an environment update. This means the agent would deliver thousands of faulty responses to users every time it encounters an environment update. Thus, fast adaptation is vital to avoid repetitive performance loss. As shown in Figure 5, across all three shift types, FSOL consistently achieved the fastest recovery and maintained superior accuracy.
Mixed Tasks
Next, we evaluated mixed-task streams by interleaving Writing and WebShop queries into a unified sequence. As shown in Figure 6, FSOL maintained higher task-specific running accuracy across both domains compared to standalone memory or RL baselines.
Sequential Tasks
Finally, we evaluated sequential task transitions of the form A → B → A (specifically, Writing → WebShop → Writing). This setup evaluates two classic continual learning properties: plasticity (the ability to learn task B after task A) and retention (the ability to retain task A capability when returning to it). Our goal was to test whether learning WebShop shows plasticity, and whether returning to Writing shows catastrophic forgetting. Figure 7 illustrates the results. All three methods (Memory, RL, and FSOL) maintained strong plasticity: learning WebShop after Writing (solid line in phase 2) progressed at a rate nearly identical to learning WebShop from the base model (dashed line). Furthermore, upon returning to Writing in phase 3, accuracies fluctuated around their end-of-Phase-1 levels, demonstrating excellent retention without catastrophic forgetting.
Takeaways
To summarize, our work highlights three core takeaways for the community:
1. Evaluation for online agent learning should shift from final checkpoint accuracy to cumulative performance, because each interaction during the learning process directly impacts the user.
2. Maximizing cumulative performance demands both fast adaptation and deep skill internalization. Memory provides rapid non-parametric adaptation, while RL delivers slower but broader parametric tuning. FSOL bridges these two timescales to maximize cumulative performance.
3. Online agent benchmarks must look beyond stationary IID streams. Practical deployments encounter distribution shifts, interleaved task compositions, and task re-emergence. A comprehensive evaluation framework should systematically encompass stationary streams, distribution shifts, mixed tasks, and sequential task transitions. Our experimental results confirm that FSOL achieves consistent performance gains across all these realistic scenarios.
For further details, please refer to our paper: Agent Online Learning Beyond Memory. Feedback and discussion are very welcome! Feel free to reach out at hbchen [AT] mit.edu.
Citation
@inproceedings{
chen2026agent,
title={Agent Online Learning Beyond Memory},
author={Huaibo Chen and Maohao Shen and Siru Ouyang and Dylan Zhang and Zexue He and Yingheng Wang and Xiaotong Zhang and Inkit Padhi and Subhajit Chaudhury and Gregory W. Wornell and Prasanna Sattigeri and Zhang-Wei Hong and KAMAL YOUCEF-TOUMI},
booktitle={NeurIPS 2026 Workshop on Towards Test-Time Continual Learning Agents},
year={2026},
url={https://openreview.net/forum?id=LtrVQShaq7}
}
References
- Longtao Zheng, Rundong Wang, Xinrun Wang, and Bo An. Synapse: Trajectory-as-exemplar prompting with memory for computer control. In International Conference on Learning Representations (ICLR), 2024.
- Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried, and Graham Neubig. Agent Workflow Memory. In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025.
- Siru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen, Ke Jiang, Zifeng Wang, Rujun Han, Long Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, and Tomas Pfister. ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory. In The Fourteenth International Conference on Learning Representations (ICLR), 2026.
- Yinjie Wang, Xuyang Chen, Xiaolong Jin, Mengdi Wang, and Ling Yang. OpenClaw-RL: Train Any Agent Simply by Talking. arXiv preprint arXiv:2603.10165, 2026.
- Tianzhu Ye, Li Dong, Qingxiu Dong, Xun Wu, Shaohan Huang, and Furu Wei. Online experiential learning for language models. arXiv preprint arXiv:2603.16856, 2026.
- Linbo Liu, Yuzhe Lu, Youzhi Luo, Sijun Tan, Panpan Xu, and Luke Huan. Continual Learning via Real-Time RL for Agents. rLLM Blog, 2026.
- Elad Hazan. Introduction to online convex optimization. Foundations and Trends in Optimization, 2(3-4):157–325, 2016.
- Ekin Akyürek, Mehul Damani, Adam Zweiger, Linlu Qiu, Han Guo, Jyothish Pari, Yoon Kim, and Jacob Andreas. The Surprising Effectiveness of Test-Time Training for Few-Shot Learning. In Forty-second International Conference on Machine Learning (ICML), 2025.
- Pengrui Han, Peiyang Song, Haofei Yu, and Jiaxuan You. In-context learning may not elicit trustworthy reasoning: A-not-b errors in pretrained language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 5624–5643, 2024.
- Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems (NeurIPS), 35:24824–24837, 2022.
- Zhenyu Hou, Yujiang Li, Jie Tang, and Yuxiao Dong. Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning. arXiv preprint arXiv:2607.07508, 2026.
- Alexander Pritzel, Benigno Uria, Sriram Srinivasan, Adria Puigdomenech Badia, Oriol Vinyals, Demis Hassabis, Daan Wierstra, and Charles Blundell. Neural episodic control. In International Conference on Machine Learning (ICML), pages 2827–2836, 2017.
- Matthew Botvinick, Sam Ritter, Jane X Wang, Zeb Kurth-Nelson, Charles Blundell, and Demis Hassabis. Reinforcement learning, fast and slow. Trends in Cognitive Sciences, 23(5):408–422, 2019.
- John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347, 2017.
- Jian Hu, Jason Klein Liu, Haotian Xu, and Wei Shen. REINFORCE++: Stabilizing critic-free policy optimization with global normalization. arXiv preprint arXiv:2501.03262, 2025.
- Angela Fan, Mike Lewis, and Yann Dauphin. Hierarchical neural story generation. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL), pages 889–898, 2018.
- Shunyu Yao, Howard Chen, John Yang, and Karthik Narasimhan. Webshop: Towards scalable real-world web interaction with grounded language agents. Advances in Neural Information Processing Systems (NeurIPS), 35:20744–20757, 2022.
- Mohit Shridhar, Xingdi Yuan, Marc-Alexandre Côté, Yonatan Bisk, Adam Trischler, and Matthew Hausknecht. ALFWorld: Aligning Text and Embodied Environments for Interactive Learning. arXiv preprint arXiv:2010.03768, 2021.
- Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 2369–2380, 2018.
- Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics (TACL), 7:453–466, 2019.