Storage
Memory is backed by a vector database (SQLite backend), living in your device. The wrapper is, which manages two collections:the memory store is initialized at startup. a vector database init failures are
wrapped so they raise SystemExit(1) and surface clearly in the logs rather than
hanging the startup.The retrieval loop
- On each
/chat,memory.get_context(query)runs a semantic search over stored facts. - The most relevant facts are injected into the LLM system prompt.
- After responding, a background task asks the LLM to extract new facts from the exchange and writes them back to the vector store.
Why vector search
Fuzzy recall
“What’s my sister’s birthday?” matches a stored fact phrased entirely differently.
Scales quietly
Facts accumulate over time; only the top matches enter the prompt, keeping it small.
Related services
LongTermMemory: higher-level memory orchestration.WorkflowContext: the per-workflow shared variable context used for step-to-step$variablesubstitution (distinct from long-term memory).