I Sell Memory APIs. I'm Also Building the Benchmark. Here's How I'm Trying Not to Rig It.

Woochan, a developer at Wontopos, has introduced a new, transparent benchmark designed to evaluate the memory capabilities of LLMs. Addressing widespread industry skepticism regarding vendor-published performance metrics, the project aims to establish a fair, reproducible standard. The benchmark utilizes a massive corpus of 1.9 million tokens, testing models across 14 axes, including stale fact detection, contradiction handling, and multilingual recall. To prevent bias, the project is open-source under the Apache 2.0 license, and the author has implemented strict rules, such as mandatory publication of per-question records and independent verification. The benchmark also introduces a nuanced approach to measuring latency, separating network overhead from actual compute performance. By prioritizing transparency and community-driven validation, the author seeks to provide a reliable way for developers to compare memory APIs without the influence of potentially rigged marketing claims.
This is a summary. Read the full article at the original source:
Dev.toRelated stories
I made my AI office manager swear at the agents. And only then did I realize what actually worked
The author shares an unconventional experience in optimizing AI agent performance. After finding that polite and diligent agents were producing medioc…
I made two AIs review each other's code for 30 days. A human still caught the bug in 5 minutes.
A developer recently experimented with using two AI agents to manage code quality: an 'author' agent to write features and a 'skeptic' agent to perfor…
The article 'Aligned to whom?' explores the complex and often ambiguous nature of AI alignment. It challenges the prevailing industry narrative that A…



