Event date · · Recuris

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

FACT STATEMENT

Recuris, a recursive Experiential-Working Memory architecture, improves task success in 35 of 37 completed model-benchmark pairs across four long-horizon benchmarks and ten models. It adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5 on tau-bench, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6-27B/35B on SkillFlow. The advantage widens to +32.2 points on the longest tasks.

What happened

Recuris introduces a recursive Experiential-Working Memory architecture for long-horizon agent harnesses. Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than full history. Execution becomes structured evidence that localizes failures to specific memory components. A fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory, forming a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of 37 completed model-benchmark pairs, carrying frontier models to SOTA-level task success.

Technical significance

The architecture couples Working Memory and Experiential Memory to ground skill selection in current task state, then uses execution evidence to localize failures and apply validation-gated updates to Skill Memory. This bounded recursive loop avoids full-history conditioning and enables targeted self-improvement. The observed +32.2 point gain on longest tasks suggests the approach scales with interaction horizon.

Industry impact

Long-horizon agent reliability remains a key barrier to enterprise deployment. Recuris demonstrates that memory-centric recursive self-improvement can lift frontier models to SOTA-level task success without retraining, potentially reducing the cost and complexity of building robust agent harnesses.

Decision value

For enterprises deploying long-horizon agents, Recuris offers a path to higher task success rates using existing frontier models, potentially lowering failure rates and operational costs. The validation-gated update mechanism may also reduce risk by localizing and correcting specific memory components.

What to watch

Next observable signals include replication on additional benchmarks, open-source release of the Recuris harness, and integration into commercial agent frameworks. If the +32.2 point gain on longest tasks holds, expect increased investment in memory-evolution techniques for autonomous agents.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.