Agent Memory Leaderboard Opened Registration for Cycle 2

The evaluation platform now includes textual, coding, and multimodal tracks to standardize how AI systems store and recall data.

Updated on Sept. 28, 2026 in Artificial Intelligence

Isometric editorial illustration of a stack of solid-state server modules, representing a standardized benchmarking framework for AI memory.
The Agent Memory Leaderboard has opened registration for its second cycle, expanding standardized benchmarking to include coding and multimodal AI systems. AI Illustration. Upload story photo >

Live Poll

Do you believe standardized evaluation frameworks will improve the quality of future artificial intelligence systems?

The Agent Memory Leaderboard (AML) has opened registration for its second cycle, expanding its scope to include coding and multimodal memory evaluation. The project aims to address the current lack of consistency in how AI systems use different datasets, retrieval methods, and scoring approaches.

Why it matters

As AI agents move beyond simple prompt-response loops, reliable long-term memory has become a critical bottleneck for performance. This leaderboard forces a standardized evaluation framework onto a field currently fragmented by disparate, non-comparable benchmarking methods.

The leaderboard provides a total prize pool of US$22,000 for open-source teams, with US$3,000 awarded for first-place finishes in each of the three tracks. Participants must provide standardized Add and Search APIs, which the AML uses to automate answer generation and evaluation.

The players

Agent Memory Leaderboard

An open platform dedicated to standardizing benchmarks for AI memory and retrieval capabilities.

The details

To compete, teams provide Add and Search APIs — programming interfaces that allow the leaderboard to interact with a system's internal memory. The AML then handles the downstream answer generation, objective evaluation, and result review. By forcing all entries to use this uniform structure, the leaderboard aims to solve the problem where systems currently rely on wildly different retrieval methods and scoring approaches that make direct performance comparisons impossible.

Timeline

  1. September 28, 2026: Cycle 2 registration officially opened.

  2. October 31, 2026: The final deadline for application submissions.

  3. November 4, 2026: The evaluation cycle is scheduled to close.

  4. mid-November 2026: Results are expected to be announced.

The Tech Race

The Agent Memory Leaderboard seeks to do for agentic retrieval what the MMLU benchmark did for foundational model reasoning. This initiative attempts to replicate the standardization that MMLU brought to language modeling, applying that rigor specifically to the emerging challenge of agent-based memory storage.

Developers and researchers can participate in the evaluation for free to benchmark their own agents against commercial and open-source peers. For the broader industry, this will establish a clearer baseline for which memory architectures actually work, potentially accelerating the development of more capable AI agents.

The takeaway

The move toward standardized retrieval APIs represents the most significant shift toward making AI memory performance measurable rather than anecdotal. Watch for the mid-November results to see which architectures perform best across the new multimodal memory track.

What happens next

Teams interested in participating must submit their applications by the October 31, 2026, deadline to be eligible for the US$22,000 prize pool. The results of the evaluation are expected to be published in mid-November 2026.

Further reading

For more on the current state of agent development, see our Artificial Intelligence coverage.

Source note: This article includes information reported by The Manila times.

Live Poll

Do you believe standardized evaluation frameworks will improve the quality of future artificial intelligence systems?