Hermes Agent's long-term memory is provided by Hindsight, a persistent memory system for AI agents from Vectorize. It recalls relevant context before each LLM call and retains conversations across sessions, so Hermes keeps continuity beyond a single chat.
Vendor documentation: hindsight.vectorize.io
How it plugs into Hermes
| Mechanism | Hook | Purpose |
|---|---|---|
| Auto-recall | pre_llm_call |
Queries memories and injects the result into the system prompt |
| Auto-retain | post_llm_call |
Stores the exchange after each response |
| Explicit tools | model-invoked | hindsight_retain, hindsight_recall, hindsight_reflect |
OtherU deployment
We run Hindsight in local mode, the embedded server with its own PostgreSQL, rather than the hosted cloud API. It lives on the inference host otheru alongside Ember.
| Item | Value |
|---|---|
| Service | hindsight-api on otheru |
| Port | 9177 |
| API version | 0.8.4 |
| Mode | local (embedded server + PostgreSQL) |
Enabled features on this deployment: observations, MCP, background worker, bank config API, file upload, document import/export, LLM trace, and document-text storage. Audit log and per-bank LLM health checks are not enabled.
Endpoints
The API is bank-scoped. The main routes are:
| Route | Purpose |
|---|---|
/v1/default/banks/{bank_id}/memories/recall |
Retrieve relevant memories |
/v1/default/banks/{bank_id}/memories/list |
List stored memories |
/v1/default/banks/{bank_id}/reflect |
Higher-level reflection over memories |
/v1/default/banks/{bank_id}/entities |
Entity records and entity graph |
/v1/default/banks/{bank_id}/mental-models |
Derived mental models |
/v1/default/banks/{bank_id}/stats |
Bank statistics and time series |
/health, /version |
Health and version probes |
Configuration
Environment variables:
| Variable | Purpose |
|---|---|
HINDSIGHT_API_URL |
Endpoint (our local server, not the cloud default) |
HINDSIGHT_API_KEY |
Authentication token |
HINDSIGHT_MODE |
cloud or local, we use local |
HINDSIGHT_BANK_ID |
Memory bank identifier |
HINDSIGHT_AUTO_RECALL / HINDSIGHT_AUTO_RETAIN |
Toggle the automatic hooks |
Config file at ~/.hermes/hindsight/config.json:
| Key | Values |
|---|---|
recallBudget |
low, mid, high |
memory_mode |
hybrid, context, tools |
prefetch_method |
recall or reflect |
Setup is via hermes memory setup (select hindsight), or by setting the variables and config file manually.
Operations notes
These are lessons from running it in production, worth reading before touching the service:
- Run it as
**glovepost**, not**root**. Starting it as the wrong user has caused an extended outage. **auto_recall**must be enabled. With it off, memories are still written but never read back, so the system looks healthy while silently doing nothing useful. Recall types in use:observation,world,experience.- Memory pressure is shared with Ember. The two compete on the same host;
oom_score_adjis tuned so Ember is protected and Hindsight is sacrificed first if the kernel must choose. - A fresh start can fail on a latent
huggingface-hubdependency, check that before assuming a config fault.