Developer Tool / AI Evaluation / Benchmarking

agent-memory-leaderboard

Public evaluation code for an Agent Memory Leaderboard, including answer-generation and scoring contracts for comparing LLM agent memory systems.

Clear24/30
Useful27/30
Specific13/20
Complete13/20
agent-memory-leaderboard screenshot

Why it was accepted

The page clearly presents a real public repository for an AI-adjacent developer tool: a fixed harness for evaluating agent memory systems. The README gives concrete runtime details, supported phases, concurrency limits, timeouts, model usage, and Python version/dependency requirements, which is enough for a useful directory listing.

Weakness

The crawl does not show how to run the benchmark end to end, what datasets or memory systems are supported beyond folder names, or any example outputs/results. There is also no project overview outside the technical evaluation contract.

Review status

29 days ago #1637 → 0

Last evaluated 29 days ago. Current rank #1637. Holding steady in the rankings.

Score history

77

Related listings

bdinfo-rs screenshot
92

Developer Tool / Media Analysis

A Rust Blu-ray disc analyzer that reimplements BDInfo for BDMV folders and ISO images, with CLI, desktop, and browser/WASM front ends.

WebLLM screenshot
#7 WebLLM
92

Developer Tool / AI SDK / In-browser LLM inference

WebLLM is a high-performance in-browser LLM inference engine that runs locally in the browser with WebGPU acceleration. It exposes an OpenAI-compatible API, supports streaming and JSON mode, and includes examples for building chat apps and browser extensions.

Molt screenshot
#31 Molt
90

Developer Tool / AI Agent Infrastructure

Molt is an open protocol for letting an AI agent make bounded purchases through disposable, single-use payment credentials. The repo includes a clear quickstart, MCP integration, a threat model, and a working test-mode demo.

Flightdeck screenshot
90

Developer Tool / AI Observability

Self-hosted observability and control plane for production and coding AI agents, with live timelines, fleet-wide feeds, token budgets, MCP allow/block rules, and support for Claude Code plus a Python sensor.