- Update Quick Start to use with_memory() API (replaces with_fast_memory)
- Add hierarchical scoping section (USER → SESSION → AGENT → TURN)
- Add temporal versioning section with supersession examples
- Add comparison with state of the art (Letta, Mem0) with feature matrices
- Document Memory API for direct access (.memory.search, .add, .get_all)
- Add advanced HierarchicalMemory API usage examples
- Update configuration options for embedders and storage
- Add protocol-based architecture diagram
- Update memory categories (PREFERENCE, FACT, CONTEXT, ENTITY, DECISION, INSIGHT)
- Add troubleshooting and best practices sections
Features:
- with_fast_memory(): Zero-latency inline extraction (Letta-style)
- Memory extracted as part of LLM response, no extra API calls
- Semantic retrieval with local embeddings (sub-50ms)
- with_memory(): Background extraction for non-blocking memory
- SQLite + FTS5 storage with vector similarity search
- Multi-user isolation by user_id
Memory enables temporal compression - extract key facts instead of
carrying full conversation history (4000 tokens → 50 tokens).
Includes:
- Comprehensive test suite (71 new tests)
- Documentation (docs/memory.md)
- Benchmark examples comparing approaches
- E2E test with LLM-as-judge evaluation