A Skeptic's Shortlist for Open-Source Agent Memory
Most shortlists start too early. They collect names before defining the job. I would start with three disqualifiers: the system must work outside one agent framework, keep sensitive session data under operator control, and expose a repeatable way to test retrieval quality.
Eliminate categories before comparing projects
A transcript archive is useful, but it is not automatically reusable memory. A vector database stores vectors, but it does not decide what part of a tool trace deserves preservation. A framework-native memory module may be excellent inside that framework while creating migration cost elsewhere.
For a coding-agent team, the shortlist should separate four jobs:
- capture a session without modifying the reasoning loop;
- extract a bounded reusable unit;
- retrieve it under an explicit scope;
- inject or reuse it with a failure policy.
Semantix belongs on the shortlist only when those four jobs match the requirement. It should not be shortlisted as a general knowledge base, a hosted memory API, or a replacement agent framework.
The 20-minute rejection test
go build -o semantix ./cmd/semantix
go vet ./...
go test ./...
Then ingest one small JSONL file and run the same query through three retrievers:
semantix search --query "fix the failing Go test" --retriever bm25
semantix search --query "fix the failing Go test" --retriever vector
semantix search --query "fix the failing Go test" --retriever hybrid
Record whether the expected slice appears, whether an unrelated project appears, and whether repeated hybrid searches preserve ranking. If the team cannot explain a miss from those outputs, the project is not operationally ready for that team—even if the demo succeeds.
My current verdict
The repository shows a real CLI, a local JSONL-backed store, deterministic hashing, hybrid RRF fusion, and tests for unsafe L3 reuse. Those are inspectable strengths. The weaker side is external validation: there is no published representative real-session corpus proving the “70% relevance gate.” Shortlist it for a controlled pilot, not for an unqualified production recommendation.
Sources and limitations
- Quickstart — commands and supported release paths.
- M0 gate report — what passed, what is conditional, and what remains unverified.
- Source and tests — implementation is the final authority.