A Framework Lost to a Plain Baseline

Token-F1 0.2109 baseline against 0.1323 for Letta. A diagnosis, not a verdict.

2026-05-06 · LAB NOTES

150 generated sessions. Same questions, two conditions: Letta, and a plain baseline with no memory framework at all. Average token-F1 came out 0.2109 for the baseline and 0.1323 for Letta. The headline writes itself. I am not running it, because the honest result is narrower than the headline. ## What the run cannot tell you Retrieval and generation were entangled in this setup. When the answer is wrong, this experiment cannot say whether the framework retrieved the wrong memory or retrieved the right memory and used it badly. Averaged over 150 sessions that distinction is the entire story, and nothing in this design separates it. There are at least four live explanations for the gap and I can rule out none of them: a genuine memory failure, an ingestion problem, an answer-format mismatch penalised by the metric, or a weak answer model dragging both conditions down unevenly. Token-F1 is also a blunt instrument. It rewards overlap with the gold string. A correct answer phrased differently scores badly, which is not a hypothetical concern in this lab. ## What is defensible In this harness, on this task, with this metric, the added machinery did not pay for itself. That is a diagnosis of one configuration. It is a starting point for a controlled run, not a verdict on the framework, and anyone quoting it as "Letta loses" is quoting me wrong. The distinction is not modesty. A verdict ends an investigation and a diagnosis directs one. The next experiment isolates the retrieval layer and scores it on its own, so the answer layer stops hiding inside the number. Diagnosis before verdict.

Back to all writing