MEMORY WITHOUT VECTORS

Everyone reaches for vector databases. We've been running flat markdown files with frontmatter and grep for months. For agent memory at this scale, it's not a compromise. It's the better architecture.

THE VECTOR DEFAULT

The answer to "how do I give my agent memory" is always the same: vectors. Embed into Pinecone or Chroma or Weaviate, retrieve with cosine similarity. RAG. Every tutorial, every framework, every pitch deck.

We built that. It works for searching large corpora. For the kind of memory a coding agent actually needs -- a few hundred facts about the user, the project, and past mistakes -- it's overengineered in ways that actively hurt.

What we ran into:

  • Embedding drift. The same concept embedded at different times produces different vectors. Memory retrieval quality degrades as the store grows and embedding distributions shift.
  • Context collapse. Vector similarity retrieves content that's semantically close but contextually wrong. "The user prefers dark themes" and "never use dark themes" are semantically similar and functionally opposite.
  • Opacity. You can't grep a vector database. You can't open it in a text editor. You can't git diff it. Debugging why the agent remembered the wrong thing requires tooling that doesn't exist yet.
  • Infrastructure overhead. A vector database is a running service. It needs embedding model calls on write, similarity search on read, and periodic maintenance. For a single-user coding agent, this is a server to remember 200 facts.

We scrapped it and started over with the simplest thing that could work.

THE ACTUAL PROBLEM

Agent memory for coding workflows has specific characteristics that differ from general RAG:

  • Small corpus. A power user accumulates maybe 100-300 memories over months. Not millions of documents. Not even thousands.
  • Categorical, not continuous. Memories fall into distinct types: facts about the user, behavioral feedback, project state, external references. These categories matter more than semantic similarity.
  • Retrieval by relevance to task, not similarity to query. When the agent starts a coding session, it doesn't search for "memories similar to this prompt." It loads context relevant to the current project, the user's preferences, recent decisions.
  • Human-readable and editable. The user should be able to read their agent's memories, correct wrong ones, delete outdated ones. Memory is a shared resource between human and agent.

Two hundred entries. Categorical structure. Human editability. Task-based retrieval. This is a filesystem problem. We were treating it like a search problem.

THE ARCHITECTURE

Our memory system has three components: individual memory files, an index, and a set of conventions for when to read and write.

MEMORY FILES

Each memory is a markdown file with YAML frontmatter:

---
name: trading approach
description: Alt swings preferred over major coin grinding, spot-only is platform constraint
type: feedback
---

Alt swings > major coin grinding. Spot-only is an OKX constraint,
not a preference.

**Why:** User corrected assumption that they preferred conservative
major-pair trades. They want higher-volatility alt positions.

**How to apply:** Default to alt pairs when suggesting trades.
Only recommend BTC/ETH for specific macro setups.

The frontmatter fields serve as structured metadata: name for identification, description as a one-line summary used for relevance decisions, type for categorical filtering.

Four memory types, each with distinct semantics:

  • user -- who the person is, their role, expertise, preferences. Shapes how the agent communicates and what it assumes.
  • feedback -- corrections and confirmations. What to stop doing, what to keep doing, and why. These are the highest-priority memories because they represent explicit behavioral guidance.
  • project -- active work context. What's being built, why, constraints, deadlines. These decay fastest and need regular updates.
  • reference -- pointers to external resources. Where to find things outside the local filesystem.

THE INDEX

A single MEMORY.md file serves as the index. It's loaded into the agent's context at the start of every conversation. Each entry is one line:

# Memory Index

## User
- [user_profile.md](user_profile.md) -- Full personal profile: role, interests, location

## Feedback
- [feedback_no_dark_themes.md](feedback_no_dark_themes.md) -- NEVER use dark themes
- [feedback_dangerous_mode.md](feedback_dangerous_mode.md) -- Full bypass permissions

## Projects
- [project_flock.md](project_flock.md) -- Flock v0.9.1: macOS Claude Code multiplexer

The index is capped at 200 lines. It's a table of contents, not a memory store. The agent reads the index, decides which memories are relevant, and reads only those files. This is lazy loading -- you don't embed all 200 memories into every conversation. The index costs maybe 3K tokens. Reading a specific memory costs another 200-500 tokens.

RETRIEVAL

There's no search algorithm. The agent reads the index, uses the name and description fields to decide relevance, and reads the files it needs. This works because:

  1. The index is small enough to fit in context (3K tokens)
  2. The descriptions are written to be maximally informative for relevance decisions
  3. The agent is an LLM -- it's better at judging relevance from natural language descriptions than any cosine similarity score

In practice, the agent reads 5-15 memory files per session, depending on the task. Total memory context: 5-10K tokens. That's a fraction of what a RAG pipeline would inject.

WHY GREP BEATS COSINE SIMILARITY

For this problem, at this scale, exact match and keyword search outperform semantic similarity. Here's why:

Precision matters more than recall. When the agent retrieves a memory about dark themes, it needs exactly the right memory -- not five memories that are vaguely about themes. A grep for "dark theme" in the index returns exactly one result. A vector search returns a ranked list where the correct memory might be third, behind memories about "theme configuration" and "dark mode in the terminal."

The LLM is the semantic layer. Vector search uses embeddings to approximate semantic understanding. But the agent IS a language model. It already understands semantics. Giving it the raw index and letting it decide is using a better semantic engine than any embedding model.

Debuggability. When the agent acts on a wrong memory, you can trace exactly what happened. Open the index, find the memory, read it, fix it. The entire memory system is a directory of text files. grep -r "dark theme" memory/ tells you everything. Try debugging why a vector search returned the wrong result.

WRITE CONVENTIONS

When to write memory is as important as how to store it. We follow strict rules:

Save feedback immediately. When the user corrects the agent ("don't do that", "stop doing X"), that correction becomes a feedback memory before anything else happens. Feedback memories include a Why line (the reason for the correction) and a How to apply line (when the rule kicks in). The why matters because it lets the agent judge edge cases instead of blindly following rules.

Don't save what the code already knows. File paths, architecture patterns, function signatures -- these are in the codebase. Reading the code is authoritative. A memory about "the auth module is in src/auth/" becomes stale the moment someone renames the directory. The code never lies. Memory can.

Convert relative dates. "Next Thursday" in a memory is meaningless two weeks later. All dates are stored as absolutes: "2026-03-05". This seems minor until you have a project memory that says "freeze is next week" and the agent reads it a month later.

Verify before acting. A memory that names a specific file, function, or flag is a claim about the past. Before recommending it, the agent checks that it still exists. "The memory says X exists" is not the same as "X exists now."

MULTI-TIER MEMORY

The flat-file system is one tier in a larger hierarchy:

  1. Project memory files -- the system described above. Per-project, loaded by the agent at session start.
  2. Session debriefs -- end-of-session summaries capturing decisions, open threads, and outcomes. Stored as structured JSON. Queryable for continuity between sessions.
  3. Knowledge base -- an Obsidian vault serving as long-term storage. Session logs, research findings, project documentation. The agent reads and writes to it naturally during sessions.

Each tier has different read/write patterns. Project memory is loaded automatically. Session debriefs are queried on demand ("what was I working on yesterday?"). The knowledge base is referenced when the agent needs deeper context than project memory provides.

WHAT WE'D CHANGE

The system works well at its current scale. The weaknesses we've identified:

Index maintenance is manual-ish. Adding a memory is a two-step process: write the file, update the index. The agent does this automatically, but occasionally the index drifts from the actual files. An fswatch-based sync would eliminate this.

No temporal queries. "What did I decide about X last week?" requires reading session debriefs, not project memory. The boundary between "what's true now" (project memory) and "what happened when" (debriefs) could be cleaner.

Scale ceiling is real. At 500+ memories, the index stops fitting comfortably in context. We haven't hit this yet, but the architecture doesn't scale to thousands of entries without adding a search layer. For the single-user coding agent use case, 200 entries covers months of work. For multi-user or enterprise use, you'd need something else.

THE ARGUMENT

Vectors solve a real problem. If you're building a support agent that searches 50,000 knowledge base articles, use them. That's what they're for.

Agent memory isn't that problem. It's a small, structured, categorical dataset that a human needs to read and edit and a language model needs to reason about. For this problem, flat files with frontmatter and an LLM-readable index beat vectors on precision, debuggability, editability, and cold-start simplicity.

The best retrieval engine for natural language descriptions is a language model. Give it the index. It already knows what's relevant.

← ALL POSTS