Research

MachinoAI explainer / AI Agents

WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution

WikiSkill gives agent skill evolution a persistent memory of what worked, what failed, and why.

Liyan Tang, Cyrus Rashtchian, Chun-Sung Ferng, Andrew Tomkins, Da-Cheng Juan, Tu VuAug 27, 2026arXiv / cs.AI14 min readGoogle Research; Virginia Tech
TRENDINGAI AgentsAdvancedAgent MemorySkill EvolutionValidation Gating

01

Abstract

WikiSkill addresses a simple but important problem in self-improving agents: useful experience is often scattered across optimization histories and disappears when a proposed skill is rejected. The paper introduces a persistent wiki that sits between raw execution traces and executable skills. The wiki accumulates failure patterns, successful strategies, rejected proposals, and validation outcomes, while the active skills remain compact and reversible. Across five benchmarks and five models, WikiSkill reports the strongest average performance among the evaluated skill-evolution methods. The paper also finds that evolved skills can transfer across models and that persistent knowledge is critical to the gains.

02

Introduction

Modern LLM agents can execute multi-step tasks, but many require domain-specific procedures rather than general language ability alone. Skills provide a lightweight way to add such procedures without changing model weights. The challenge is how those skills should improve over time.

Prior skill-evolution systems learn from execution traces, but the lessons behind rejected changes or recurring failures can remain fragmented. WikiSkill asks whether agent experience can instead be compiled into persistent, structured knowledge that future skill updates can reuse.

The central design choice is deliberately asymmetric: executable skills can be rolled back, while the wiki is never rolled back. This lets the system reject a bad instruction without forgetting why it was tried.

03

Problem

The paper formalizes iterative skill evolution for an LLM agent with tools, a set of active skills, and train/validation/test task splits. At each iteration, the agent produces execution trajectories; a separate process extracts useful patterns; a proposer creates a candidate skill update; and a validation gate decides whether that update should become active.

The core problem is not merely generating a better skill. It is preserving the evidence accumulated during previous iterations so future proposals do not repeatedly rediscover the same failure modes.

04

Background

Agent skills are modular procedural resources that can contain instructions, scripts, and applicability conditions. They can improve specialized behavior without modifying the underlying model parameters. Existing skill-evolution methods already use agent trajectories to revise skills, but their representations of accumulated experience differ.

WikiSkill is inspired by the idea of compiling experience into persistent knowledge. Its contribution is to make that knowledge an explicit layer rather than keeping it only inside skill text or optimization history.

05

Methodology

WikiSkill uses four components in an iterative loop: an Inference Agent generates task trajectories using the current skills; a Wiki Maintainer consolidates sampled successes and failures into persistent wiki knowledge; a Skill Proposer uses the wiki and selected traces to propose one atomic skill change; and a Gating/Rollback step keeps the change only when validation performance strictly improves.

The inference agent is intentionally not given direct wiki access during task execution. The wiki is used to improve the skill, while the inference agent receives the resulting compact skill.

Formula notes

Mathematical details

Skill-evolution gating rule

Equation
Sk=Sk′ifR(Tval,k)>Rbest;otherwiseSk=Sk−1S_k = S'_k if R(T_val,k) > R_best; otherwise S_k = S_{k-1}

A candidate skill change is accepted only when its validation score is strictly better than the best score seen so far; otherwise the previous skill set is restored.

SkS_k
accepted active skill set at iteration k
Sk′S'_k
candidate skill set after applying the proposal
RbestR_best
best validation score before the proposal
R(Tval,k)R(T_val,k)
validation performance of the candidate

06

Architecture

The system has three layers. The Raw Layer stores immutable execution traces. The Wiki Layer stores persistent patterns, an evolution log, and a skill-impact record. The Skills Layer contains the active procedural instructions used by the inference agent.

This separation is the key architectural idea: the history can grow and remain detailed without forcing the entire history into the inference prompt.

07

Dataset

The evaluation uses five benchmarks: LiveMathematicianBench for mathematical reasoning, SealQA for web search, SpreadsheetBench for spreadsheet manipulation, OfficeQA for long-context document question answering, and ALFWorld for interactive embodied tasks. The experiments use five models spanning Qwen, Gemma, and Gemini families.

08

Training

There is no neural-model training or weight updating in WikiSkill itself. The framework evolves external skill files through agent experience, proposal, and validation. The paper starts from an empty skill set and iteratively develops skills for each dataset.

09

Experiments

The authors evaluate WikiSkill against Trace2Skill, EvoSkill, SkillOpt, and a no-skill baseline across five benchmarks and five inference models. Each skill-evolution method starts with an empty skill set, and the resulting skills are injected into the inference agent at evaluation time. The full evolution process is repeated three times and reported scores are averaged across the runs.

The main result is consistent: WikiSkill achieves the highest average score for all five tested inference models. For example, Gemini-3.5-Flash improves from 49.5% without skills to 68.1% with WikiSkill-evolved skills. Qwen-3.6-27B improves from 39.4% to 63.3%.

10

Baselines

The primary baselines are Trace2Skill, EvoSkill, SkillOpt, and a no-skill agent. These methods share the broad pattern of executing tasks, analyzing trajectories, proposing skill changes, and evaluating them. The comparison is designed to isolate the value of WikiSkill’s persistent knowledge layer.

11

Results

WikiSkill achieves the highest average performance across all five inference models in the main comparison. The reported gains over the strongest competing skill-evolution method are 3.3, 5.1, 10.0, 5.8, and 12.0 percentage points for Qwen-3.5-4B, Qwen-3.5-9B, Qwen-3.6-27B, Gemma-4-31B, and Gemini-3.5-Flash respectively.

The paper also finds that skill evolution complements model scaling. Within the Qwen family, WikiSkill improves average performance by 12.3, 17.5, and 23.9 points for the 4B, 9B, and 27B models. A striking comparison is Qwen-3.5-9B with WikiSkill at 47.4% versus Qwen-3.6-27B without skills at 39.4%.

Cross-model transfer is also promising: on ALFWorld, Qwen-3.5-9B reaches 70.2% using a skill evolved by Qwen-3.6-27B, versus 63.4% with its own evolved skill.

12

Ablation

The ablation directly supports the persistent-knowledge hypothesis. In the Gemini-3.5-Flash study, the default setup averaged 63.7% when the Skill Proposer had wiki access but the Inference Agent did not. Removing wiki access from the proposer and the Wiki Maintainer reduced the average to 48.7%. Giving the Inference Agent wiki access as well reduced the result to 60.9%.

The authors interpret this as evidence that the wiki is most useful as development memory for improving skills, rather than as additional runtime context.

13

Limitations

The paper deliberately evaluates full skill injection rather than skill retrieval, so it does not establish how a large production skill library should select the right skill. The validation sets are small, although the authors repeat evolution three times and use paired bootstrap testing.

The wiki also grows without an automated pruning mechanism. The strict validation gate rejects neutral changes even if they might enable later improvements. Finally, the experiments do not establish behavior on extremely long-running real-world tasks lasting hundreds of actions or hours.

14

Conclusion

WikiSkill reframes agent improvement as a knowledge-management problem as much as a prompt-optimization problem. The model does not need to remember every failure directly; a persistent development layer can organize those failures and use them to improve compact executable skills.

The strongest takeaway is the separation between durable knowledge and reversible instructions. That design produced consistent gains in the paper’s experiments and opens a practical direction for agents that repeatedly perform the same classes of tasks.

15

References

Tang, L., Rashtchian, C., Ferng, C.-S., Tomkins, A., Juan, D.-C., & Vu, T. (2026). WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution. arXiv:2608.27454.

Related methods discussed by the paper include Trace2Skill, EvoSkill, and SkillOpt. External coverage from VentureBeat provides additional discussion of the persistent-wiki design and reported experiments.

Continue reading

Related research

Browse all research