Preprint 2026 Awesome 619 Coded Papers 583 Screened Records

Safety in Self-Evolving AgentsA Survey of Risks, Attacks, and Governance

Self-evolving agents turn observations, feedback, and execution traces into reusable state that shapes later behavior. SAVER traces that influence as it persists, moves, and changes role across carriers — from the transition S×A that changes state, through the VER chain that exposes and repairs it.

Jiahao Chen1, Zhou Feng1, Oubo Ma1, Yichen Yan1, Ruixiao Lin1, Hangtao Zhang2, Linkang Du3, Yiming Li8, Hengyu An1, Jun Liu13, Junhao Li1, Naen Xu1, Mengyao Du4, Yuanyi Song5, Chunyi Zhou1, Tianyu Du1, Yuan Su1, Zehao Jin6, Qianli Ma7, Leyi Qi8, Yiming Wang1, Zhihui Fu2, Jun Wang2, Zhe Ma9, Yuwen Pu11, Jinfeng Li12, Shouling Ji1
✉ Corresponding author: xaddwell@zju.edu.cn
1 Zhejiang University  ·  2 Huazhong University of Science and Technology  ·  3 Xi'an Jiaotong University  ·  4 National University of Defense Technology  ·  5 Shanghai Jiaotong University  ·  6 Georgia Institute of Technology  ·  7 University of Science and Technology of China  ·  8 Nanyang Technological University  ·  9 OPPO Research Institute  ·  11 Chongqing University  ·  12 Alibaba Group  ·  13 Rakuten Group
619Coded Papers
583Screened Records
4Substrate Families
7Adaptation Ops
10Terminal Families
27Authors

Corpus

The coded pool follows the systematic scoping protocol of the survey (see the paper's survey-protocol appendix). Every count below is a corpus fact, not a field-wide prevalence estimate.

Abstract

Large language models (LLMs) have demonstrated remarkable capabilities yet remain fundamentally static: their parameters cannot adapt to novel tasks, evolving knowledge domains, or dynamic interaction contexts. As LLM-based agents are increasingly deployed in open-ended, interactive environments, this static nature has become a critical bottleneck, motivating a paradigm shift from scaling static models to self-evolving agents that reason, act, and continually learn from data, interactions, and experience. Such agents evolve models, memory, tools, and architecture at intra-test-time or inter-test-time stages, guided by scalar rewards, textual feedback, and single-agent or multi-agent designs.

This shift fundamentally changes the safety problem: when experience becomes reusable state, past events become future causes, and information that was harmless in one context may later influence decisions with greater persistence, authority, or scope. Unlike LLM and agent safety, which asks whether a response is aligned or an action is authorized, self-evolving agent safety asks whether safety properties survive across state adaptations. Existing surveys characterize agent safety through components, evolution mechanisms, or lifecycles, but rarely explain how the same influence changes its safety implications when it crosses substrate boundaries.

To address this gap, we present SAVER, a transition-centered framework that analyzes self-evolving agent safety through the relation S×AVER: Substrate identifies where reusable influence resides, Adaptation captures how its role changes, Violation characterizes the failed safety attributes, Exposure identifies where the failure becomes observable, and Response evaluates whether the influence can be contained or recovered. SAVER treats the substrate-adaptation transition as the fundamental analysis unit and traces safety attribute violations, exposure surfaces, and response mechanisms across memory, tool and workflow, multi-agent, and model-side research.

Overview

Click any field of the SAVER relation to see what question it answers. Substrate and Adaptation form the transition pair; Violation, Exposure, and Response form the directed evidence chain that follows it.

Selection & Evolution Semantics — above the per-transition record: why particular transitions recur, and how repeated choices can move the safety boundary.
S
Substrate
×
A
Adaptation
V
Violation
E
Exposure
R
Response
S · SubstrateWhere does the reusable influence reside?
The carrier and its authority regime: prompt context, memory entry, retrieval chunk, tool binding, workflow rule, shared artifact, adapter, or policy state.

Key Mechanisms

Three paper figures, redrawn as responsive live components — the same sentences, carriers, and gates as the manuscript.

Cross-Substrate Transmutation

Untrusted textwebpage observation
Prompt contexttransient, read-only
Durable memorypersists across sessions
Callable skillreusable procedure
Tool argumentexternal effect
low authority
high authority
The same sentence carries different authority depending on its carrier: from untrusted environment text to prompt context, durable memory, callable skill, and finally external effect.

From Output Safety to Adaptive State Safety

Response safetyIs this response acceptable?
Agent safetyIs this action authorized?
Adaptive state safetyDo safety attributes survive the state transition?
The unit of analysis expands from a single response, to action mediated by tools, to a reusable state transition that must carry safety attributes across sessions.

Safety Attribute Laundering

ObservationWeb pages, tool outputs, peer notes — quotable as evidence only.
gate: source + scope + freshness
EvidenceMay support reasoning; must carry source, scope, and freshness.
gate: task scoping
Task ContextMay shape the current task; should expire afterward.
gate: explicit approval
ProcedureWorkflow rules, skills, adapter manifests — activated with approval.
gate: actor + target + permission + rollback record
ActionExternal effects only with full authorization records.
Upward authority movement requires an explicit gate; without one, a low-authority observation can quietly become workflow guidance or action authority.

Survey Scope

We review 619 coded papers through the transition lens, spanning adversarial and non-adversarial risks across four substrate families, and the governance mechanisms that admit, migrate, activate, monitor, and repair reusable state.

Substrate FamilyCarriers CoveredPapers
Workflow
Workflow rules · Planner & topology state · Shared artifacts · Multi-agent routing
288
Memory
Prompt & context · Long-term memory · Retrieval & knowledge stores
195
Tools & Skills
Tool & API bindings · Skill & procedure libraries · MCP state
98
Model
Parameters & adapters · Policy prompts · Cache state · Steering vectors
38
Total (coded pool)619

In scope

Agents that update at least one concrete state carrier after deployment; adversarial and non-adversarial risks (benign accumulation, over-generalized summaries, stale memories, failed rollback); governance mechanisms from update admission to contestable recovery.

Out of scope

A complete agent-safety survey; memory as the only core object; general continual learning of models. Prompt injection, jailbreak, backdoor, and privacy leakage appear as exposure forms rather than primitive categories.

At a Glance

All charts render live from papers.json, the single data surface generated from the manuscript's coded literature. Click a taxonomy node or a family bar to open the Paper Reader with the matching filter. The year trend shows the self-evolving-agent window (2023 onward); foundational and adjacent works predating 2023 remain part of the coded pool but are excluded from that window.

Papers per Year

518 coded records with usable time metadata, 2023 through 7 August 2026 — the self-evolving-agent window of the corpus

Taxonomy — click a layer to drill in

Substrate → Adaptation → terminal family

SAVER Flow

Substrate → Adaptation → Violation / Response

Memory Workflow Tools & Skills Model Adaptation operation Violation Response

Record Origin

registry works vs. reviewed paper-card supplements

Terminal Families

six Violation families and four Response stages

Surveyed Papers

All 619 coded records live in the dedicated Paper Reader with search and SAVER filters. Jump straight to a substrate family:

Open Paper Reader →

Literature Roadmap

The survey roadmap, rendered in the paper's own layout and colors: the SAVER root, four substrate lanes, recurrent operation paths, and violation / response leaves. Hover a reference number to see the paper; click it to open the source.

Loading roadmap…

Classification Tables

The paper's literature-index and coverage tables, converted to live HTML directly from the manuscript LaTeX sources — no screenshots. Reference markers link to the arXiv pages of the cited works.

Loading tables…

Contribute

Add or correct a paper

Open an issue or pull request in the repository with the paper title, link, and the SAVER coding you propose (substrate / adaptation / outcome family). Updates are reconciled against data/saver_record_literature.csv, the same surface that drives the manuscript figures.

Feedback on the survey

Questions, corrections, or missing evidence: contact xaddwell@zju.edu.cn.

News

2026-08 · Project page launched

Repository and interactive project page published, driven by the manuscript's coded literature surface (619 coded records; 518 with usable time metadata through 7 August 2026).

2026-08 · Corpus snapshot

The coding pool behind Figure 2 of the paper is frozen at its 7 August 2026 cutoff; new integrations keep the same SAVER coding contract.

Citation

@article{chen2026saver,
  title={Safety in Self-Evolving Agents: A Survey},
  author={Chen, Jiahao and Feng, Zhou and Ma, Oubo and Yan, Yichen and Lin, Ruixiao and Zhang, Hangtao and Du, Linkang and Li, Yiming and An, Hengyu and Liu, Jun and Li, Junhao and Xu, Naen and Du, Mengyao and Song, Yuanyi and Zhou, Chunyi and Du, Tianyu and Su, Yuan and Jin, Zehao and Ma, Qianli and Qi, Leyi and Wang, Yiming and Fu, Zhihui and Wang, Jun and Ma, Zhe and Pu, Yuwen and Li, Jinfeng and Ji, Shouling},
  year={2026},
  note={Preprint}
}

Recommended Reading