Self-evolving agents turn observations, feedback, and execution traces into reusable state that shapes later behavior. SAVER traces that influence as it persists, moves, and changes role across carriers — from the transition S×A that changes state, through the V→E→R chain that exposes and repairs it.
The coded pool follows the systematic scoping protocol of the survey (see the paper's survey-protocol appendix). Every count below is a corpus fact, not a field-wide prevalence estimate.
—
Large language models (LLMs) have demonstrated remarkable capabilities yet remain fundamentally static: their parameters cannot adapt to novel tasks, evolving knowledge domains, or dynamic interaction contexts. As LLM-based agents are increasingly deployed in open-ended, interactive environments, this static nature has become a critical bottleneck, motivating a paradigm shift from scaling static models to self-evolving agents that reason, act, and continually learn from data, interactions, and experience. Such agents evolve models, memory, tools, and architecture at intra-test-time or inter-test-time stages, guided by scalar rewards, textual feedback, and single-agent or multi-agent designs.
This shift fundamentally changes the safety problem: when experience becomes reusable state, past events become future causes, and information that was harmless in one context may later influence decisions with greater persistence, authority, or scope. Unlike LLM and agent safety, which asks whether a response is aligned or an action is authorized, self-evolving agent safety asks whether safety properties survive across state adaptations. Existing surveys characterize agent safety through components, evolution mechanisms, or lifecycles, but rarely explain how the same influence changes its safety implications when it crosses substrate boundaries.
To address this gap, we present SAVER, a transition-centered framework that analyzes self-evolving agent safety through the relation S×A→V→E→R: Substrate identifies where reusable influence resides, Adaptation captures how its role changes, Violation characterizes the failed safety attributes, Exposure identifies where the failure becomes observable, and Response evaluates whether the influence can be contained or recovered. SAVER treats the substrate-adaptation transition as the fundamental analysis unit and traces safety attribute violations, exposure surfaces, and response mechanisms across memory, tool and workflow, multi-agent, and model-side research.
Click any field of the SAVER relation to see what question it answers. Substrate and Adaptation form the transition pair; Violation, Exposure, and Response form the directed evidence chain that follows it.
Three paper figures, redrawn as responsive live components — the same sentences, carriers, and gates as the manuscript.
We review 619 coded papers through the transition lens, spanning adversarial and non-adversarial risks across four substrate families, and the governance mechanisms that admit, migrate, activate, monitor, and repair reusable state.
| Substrate Family | Carriers Covered | Papers |
|---|---|---|
Workflow | Workflow rules · Planner & topology state · Shared artifacts · Multi-agent routing | |
Memory | Prompt & context · Long-term memory · Retrieval & knowledge stores | |
Tools & Skills | Tool & API bindings · Skill & procedure libraries · MCP state | |
Model | Parameters & adapters · Policy prompts · Cache state · Steering vectors | |
| Total (coded pool) | 619 | |
Agents that update at least one concrete state carrier after deployment; adversarial and non-adversarial risks (benign accumulation, over-generalized summaries, stale memories, failed rollback); governance mechanisms from update admission to contestable recovery.
A complete agent-safety survey; memory as the only core object; general continual learning of models. Prompt injection, jailbreak, backdoor, and privacy leakage appear as exposure forms rather than primitive categories.
All charts render live from papers.json, the single data surface generated from the manuscript's coded literature. Click a taxonomy node or a family bar to open the Paper Reader with the matching filter. The year trend shows the self-evolving-agent window (2023 onward); foundational and adjacent works predating 2023 remain part of the coded pool but are excluded from that window.
518 coded records with usable time metadata, 2023 through 7 August 2026 — the self-evolving-agent window of the corpus
Substrate → Adaptation → terminal family
Substrate → Adaptation → Violation / Response
registry works vs. reviewed paper-card supplements
six Violation families and four Response stages
All 619 coded records live in the dedicated Paper Reader with search and SAVER filters. Jump straight to a substrate family:
The survey roadmap, rendered in the paper's own layout and colors: the SAVER root, four substrate lanes, recurrent operation paths, and violation / response leaves. Hover a reference number to see the paper; click it to open the source.
Loading roadmap…
The paper's literature-index and coverage tables, converted to live HTML directly from the manuscript LaTeX sources — no screenshots. Reference markers link to the arXiv pages of the cited works.
Loading tables…
Open an issue or pull request in the repository with the paper title, link, and the SAVER coding you propose (substrate / adaptation / outcome family). Updates are reconciled against data/saver_record_literature.csv, the same surface that drives the manuscript figures.
Questions, corrections, or missing evidence: contact xaddwell@zju.edu.cn.
Repository and interactive project page published, driven by the manuscript's coded literature surface (619 coded records; 518 with usable time metadata through 7 August 2026).
The coding pool behind Figure 2 of the paper is frozen at its 7 August 2026 cutoff; new integrations keep the same SAVER coding contract.
@article{chen2026saver,
title={Safety in Self-Evolving Agents: A Survey},
author={Chen, Jiahao and Feng, Zhou and Ma, Oubo and Yan, Yichen and Lin, Ruixiao and Zhang, Hangtao and Du, Linkang and Li, Yiming and An, Hengyu and Liu, Jun and Li, Junhao and Xu, Naen and Du, Mengyao and Song, Yuanyi and Zhou, Chunyi and Du, Tianyu and Su, Yuan and Jin, Zehao and Ma, Qianli and Qi, Leyi and Wang, Yiming and Fu, Zhihui and Wang, Jun and Ma, Zhe and Pu, Yuwen and Li, Jinfeng and Ji, Shouling},
year={2026},
note={Preprint}
}