Shepard Tone Agents
Recursive Composition Without Forgetting

Executive Summary
We propose Shepard Tone Agents (STA), a framework for continuous AI improvement through recursive agent composition and orchestrated evaluation, without modifying model weights. Inspired by the Shepard tone auditory illusion—where multiple frequencies create the perception of continuous ascent—STA uses multiple LLM agents operating at different "frequencies" (reasoning styles, abstraction levels, domains) that interact recursively. An orchestrator manages handoffs, evaluates cumulative performance, and pushes agents beyond current limits, grounded in live, real-time data. Unlike weight-update approaches (SEAL, LoRA, continual learning), STA avoids catastrophic forgetting entirely by operating on stateless agent compositions.
1. Core Problem: The Forgetting Bottleneck
Catastrophic Forgetting
Self-evolving LLMs face a fundamental obstacle: when models update weights to learn new information or tasks, they degrade performance on prior knowledge.
This problem is documented across multiple approaches:
SEAL (Zweiger et al., 2025)
Repeated self-edits cause earlier task performance to decline
Continual Learning
Balancing plasticity vs. stability remains unsolved
Production Systems
Fine-tuning requires replay buffers or regularization
Weight modification is fundamentally at odds with retention.
Our hypothesis: Composition is safer than modification. Rather than change the model, compose multiple agents into sequences that collectively solve harder problems without overwriting any prior capability.
2. The Shepard Tone: Conceptual Model
The Shepard tone is an auditory illusion where a sequence of tones appears to continuously rise in pitch indefinitely. The trick: each octave completes and restarts while the previous octave fades, creating seamless upward motion without actually reaching higher frequencies.
Applied to Agents:
Each agent is a frequency: Operating at a specific abstraction level, reasoning style, or domain specialization
Recursion is the fade: As Agent A reaches a local optimum or plateaus, Agent B takes its output and pushes further
Agent C continues: Another layer of refinement, another frequency
The orchestrator manages the sequence: Detecting when to hand off, scoring cumulative improvement
No forgetting: Each agent call is stateless. Earlier solutions remain unchanged
Result: The system appears to continuously improve (rise in Shepard tone terms) while maintaining all prior capabilities.
3. System Architecture
4. ReasoningBank Integration
ReasoningBank (arXiv:2509.25140v1) introduces powerful concepts for agent memory and test-time scaling that align naturally with STA's recursive composition approach.
1. Reasoning Memory Bank
Rather than storing raw trajectories, distill generalizable reasoning strategies from both successful and failed experiences.
2. Memory-Aware Test-Time Scaling (MaTTS)
Combines parallel and sequential scaling with memory. Parallel mode generates diverse trajectories; sequential mode produces iterative refinements.
3. Learning from Successes and Failures
Agent plateaus become constructive signals rather than noise. Failures guide subsequent agents away from dead ends.
4. Emergent Complexity
Memory items evolve from low-level execution strategies to high-level adaptive checks to compositional reasoning.
5. Self-Evaluation Without Ground Truth
LLM-as-judge approach enables scalable closed-loop learning during test time without delayed feedback.
6. Contrastive Memory Signals
Comparing successful and failed trajectories provides richer memory curation through contrastive pairs.
5. Operational Flow
6. Key Advantages
No Forgetting
- Model weights never modified
- Prior capabilities preserved
- Only orchestration improves
Compositional Emergence
- Complex behavior from interactions
- Arbitrary problem complexity via recursion
- Stateless agents remain interpretable
Computational Efficiency
- No fine-tuning per agent
- Parallel agent calls possible
- Caching of intermediate outputs
Live Data Grounding
- Validated against real data
- Automatic adaptation to drift
- Robust feedback loop
Scalability
- Add agents without disruption
- Learn new sequences without retraining
- Orthogonal to model size
7. Hypothesis & Expected Outcomes
Primary Hypothesis
Recursive agent composition via intelligent orchestration can achieve continuous improvement equivalent to (or exceeding) weight-update methods, while completely avoiding catastrophic forgetting.
Secondary Hypotheses
- •Orchestrator policy converges quickly (< 100 RL steps)
- •Optimal recursion depth is problem-type dependent
- •Live data feedback improves policy generalization
Expected Outcomes
- Accuracy: 15-30% over single-agent baselines
- No forgetting: 0% degradation on prior tasks
- Latency: 2-3x single agent (acceptable tradeoff)
- Works across LLMs and SLMs
8. Related Work & Differentiation
Existing Approaches
Our Differentiation
- First to combine recursive agent composition + learned orchestration + live data grounding
- Explicit focus on zero forgetting
- Stateless agents → interpretability + composability
- Scalable to arbitrary problem complexity via recursion depth
9. Conclusion
Shepard Tone Agents represent a novel path to continuous AI improvement. By abandoning weight updates in favor of intelligent composition, we sidestep the forgetting problem entirely.
Early experiments should validate whether this approach can match or exceed current self-evolution methods while maintaining all prior capabilities.
If successful, STA offers a scalable, interpretable, and forgiving framework for building AI systems that genuinely improve over time.
New to AI Agents?
Explore our comprehensive ABC of AI Agents glossary to understand the fundamental concepts, tools, and practices that power modern agentic systems.
Discover ABC of AI Agents