Shepard Tone Agents

Recursive Composition Without Forgetting

October 2025•SIROCCO Research
Listen to this article

Executive Summary

We propose Shepard Tone Agents (STA), a framework for continuous AI improvement through recursive agent composition and orchestrated evaluation, without modifying model weights. Inspired by the Shepard tone auditory illusion—where multiple frequencies create the perception of continuous ascent—STA uses multiple LLM agents operating at different "frequencies" (reasoning styles, abstraction levels, domains) that interact recursively. An orchestrator manages handoffs, evaluates cumulative performance, and pushes agents beyond current limits, grounded in live, real-time data. Unlike weight-update approaches (SEAL, LoRA, continual learning), STA avoids catastrophic forgetting entirely by operating on stateless agent compositions.

1. Core Problem: The Forgetting Bottleneck

Catastrophic Forgetting

Self-evolving LLMs face a fundamental obstacle: when models update weights to learn new information or tasks, they degrade performance on prior knowledge.

This problem is documented across multiple approaches:

SEAL (Zweiger et al., 2025)

Repeated self-edits cause earlier task performance to decline

Continual Learning

Balancing plasticity vs. stability remains unsolved

Production Systems

Fine-tuning requires replay buffers or regularization

Weight modification is fundamentally at odds with retention.

Our hypothesis: Composition is safer than modification. Rather than change the model, compose multiple agents into sequences that collectively solve harder problems without overwriting any prior capability.

2. The Shepard Tone: Conceptual Model

The Shepard tone is an auditory illusion where a sequence of tones appears to continuously rise in pitch indefinitely. The trick: each octave completes and restarts while the previous octave fades, creating seamless upward motion without actually reaching higher frequencies.

Applied to Agents:

Each agent is a frequency: Operating at a specific abstraction level, reasoning style, or domain specialization

Recursion is the fade: As Agent A reaches a local optimum or plateaus, Agent B takes its output and pushes further

Agent C continues: Another layer of refinement, another frequency

The orchestrator manages the sequence: Detecting when to hand off, scoring cumulative improvement

No forgetting: Each agent call is stateless. Earlier solutions remain unchanged

Result: The system appears to continuously improve (rise in Shepard tone terms) while maintaining all prior capabilities.

3. System Architecture

Input LayerTask/QueryLive Data ContextOrchestratorPolicy: Agent SequenceState ManagerEvaluatorAgent LayerAgent 1ReasoningAgent 2CritiqueAgent 3SynthesisAgent NDomainOutput & FeedbackFinal OutputGround Truth ValidationReward Signal

4. ReasoningBank Integration

ReasoningBank (arXiv:2509.25140v1) introduces powerful concepts for agent memory and test-time scaling that align naturally with STA's recursive composition approach.

1. Reasoning Memory Bank

Rather than storing raw trajectories, distill generalizable reasoning strategies from both successful and failed experiences.

2. Memory-Aware Test-Time Scaling (MaTTS)

Combines parallel and sequential scaling with memory. Parallel mode generates diverse trajectories; sequential mode produces iterative refinements.

3. Learning from Successes and Failures

Agent plateaus become constructive signals rather than noise. Failures guide subsequent agents away from dead ends.

4. Emergent Complexity

Memory items evolve from low-level execution strategies to high-level adaptive checks to compositional reasoning.

5. Self-Evaluation Without Ground Truth

LLM-as-judge approach enables scalable closed-loop learning during test time without delayed feedback.

6. Contrastive Memory Signals

Comparing successful and failed trajectories provides richer memory curation through contrastive pairs.

5. Operational Flow

Input: New task/query + live data context
Orchestrator Initialize:
├─ Load problem context
├─ Query orchestrator policy
└─ Set depth limit (e.g., max 5 recursive calls)
Agent Recursion Loop (i = 1 to max_depth):
├─ Select Agent_i
├─ Call Agent_i with context
├─ Generate output_i
├─ Evaluate contribution
│ ├─ Score improvement
│ └─ Check plateau
└─ If plateau detected → Exit; else continue
Post-Processing:
├─ Final output = best output_i
├─ Validate against ground truth
├─ Compute reward: R = f(accuracy, latency, cost)
└─ Feed R to orchestrator policy learner
Orchestrator Policy Update (RL):
├─ Log: (problem_type, agent_sequence, reward)
├─ Update policy via reinforcement learning
└─ Next iteration → improved sequence selection

6. Key Advantages

No Forgetting

  • Model weights never modified
  • Prior capabilities preserved
  • Only orchestration improves

Compositional Emergence

  • Complex behavior from interactions
  • Arbitrary problem complexity via recursion
  • Stateless agents remain interpretable

Computational Efficiency

  • No fine-tuning per agent
  • Parallel agent calls possible
  • Caching of intermediate outputs

Live Data Grounding

  • Validated against real data
  • Automatic adaptation to drift
  • Robust feedback loop

Scalability

  • Add agents without disruption
  • Learn new sequences without retraining
  • Orthogonal to model size

7. Hypothesis & Expected Outcomes

Primary Hypothesis

Recursive agent composition via intelligent orchestration can achieve continuous improvement equivalent to (or exceeding) weight-update methods, while completely avoiding catastrophic forgetting.

Secondary Hypotheses
  • •Orchestrator policy converges quickly (< 100 RL steps)
  • •Optimal recursion depth is problem-type dependent
  • •Live data feedback improves policy generalization
Expected Outcomes
  • Accuracy: 15-30% over single-agent baselines
  • No forgetting: 0% degradation on prior tasks
  • Latency: 2-3x single agent (acceptable tradeoff)
  • Works across LLMs and SLMs

8. Related Work & Differentiation

Existing Approaches

SEAL:Weight updates → forgetting problem
Mixture-of-Experts:Modular, but requires retraining
Chain-of-Thought:Single-agent reasoning
Multi-agent systems:Usually adversarial or independent, not recursive
Orchestration:No RL or live data feedback

Our Differentiation

  • First to combine recursive agent composition + learned orchestration + live data grounding
  • Explicit focus on zero forgetting
  • Stateless agents → interpretability + composability
  • Scalable to arbitrary problem complexity via recursion depth

9. Conclusion

Shepard Tone Agents represent a novel path to continuous AI improvement. By abandoning weight updates in favor of intelligent composition, we sidestep the forgetting problem entirely.

Early experiments should validate whether this approach can match or exceed current self-evolution methods while maintaining all prior capabilities.

If successful, STA offers a scalable, interpretable, and forgiving framework for building AI systems that genuinely improve over time.

New to AI Agents?

Explore our comprehensive ABC of AI Agents glossary to understand the fundamental concepts, tools, and practices that power modern agentic systems.

Discover ABC of AI Agents
October 2025•SIROCCO Research