feat(honcho): context injection overhaul, 5-tool surface, cost safety, session isolation
Context Injection Overhaul: - Base layer: peer.context() (representation + card) cached with 5-minute TTL - Dialectic supplement: cadence-gated, cached until next refresh - Trivial prompt skip: short inputs/slash commands skip injection - New peer guard: dialectic skipped at session start when peer has no context - Targeted warm prompt for better dialectic quality Tool Surface (5 bidirectional tools): - honcho_profile: read or update peer card - honcho_search: semantic search over context - honcho_context: full session context (summary, representation, card, messages) - honcho_reasoning: synthesized answer, reasoning_level param - honcho_conclude: create or delete conclusions (PII removal) Cost Safety: - dialectic_cadence defaults to 3 (~66% fewer LLM calls) - context_tokens defaults to uncapped (cap opt-in via config/wizard) - on_turn_start hook wired up (fixes broken cadence/injection gating) Correctness: - Explicit target= on peer context/card fetches (fixes identity blur) - honcho_search perspective fix under directional observation - Timeout config plumbing - peerName precedence over gateway user_id - skip_memory on temp agents (orphan session prevention) - gateway_session_key for stable per-chat session continuity - initOnSessionStart for eager tools-mode init - get_session_context fallback respects peer param - mid -> medium in reasoning level validation ABC changes (minimal, honcho-only): - run_agent.py: gateway_session_key param + memory provider wiring (+5 lines) - gateway/run.py: skip_memory on 2 temp agents, gateway_session_key on main agent (+3 lines) - agent/memory_manager.py: sanitize regex for context tag variants (+9 lines)
This commit is contained in:
@@ -49,33 +49,70 @@ Get an API key at [honcho.dev](https://honcho.dev).
|
||||
|
||||
## Configuration Options
|
||||
|
||||
```yaml
|
||||
# ~/.hermes/config.yaml
|
||||
honcho:
|
||||
observation: directional # "unified" (default for new installs) or "directional"
|
||||
peer_name: "" # auto-detected from platform, or set manually
|
||||
```
|
||||
Honcho is configured in `~/.honcho/config.json` (global) or `$HERMES_HOME/honcho.json` (profile-local). The setup wizard handles this for you.
|
||||
|
||||
**Observation modes:**
|
||||
- `unified` — All observations go into a single pool. Simpler, good for single-agent setups.
|
||||
- `directional` — Observations are tagged with direction (user→agent, agent→user). Enables richer analysis of conversation dynamics.
|
||||
**Key settings:**
|
||||
|
||||
| Setting | Default | Description |
|
||||
|---------|---------|-------------|
|
||||
| `sessionStrategy` | `per-directory` | `per-directory`, `per-repo`, `per-session`, or `global` |
|
||||
| `recallMode` | `hybrid` | `hybrid` (auto-inject + tools), `context` (inject only), `tools` (tools only) |
|
||||
| `contextTokens` | uncapped | Token budget for auto-injected context per turn. Set to an integer (e.g. 1200) to cap |
|
||||
| `dialecticReasoningLevel` | `low` | Base reasoning level: `minimal`, `low`, `medium`, `high`, `max` |
|
||||
| `dialecticDynamic` | `true` | When `true`, model can override reasoning level per-call via tool param |
|
||||
| `dialecticCadence` | `3` | Turns between Honcho LLM calls (higher = fewer calls) |
|
||||
| `writeFrequency` | `async` | When to flush messages to Honcho: `async` (background thread), `turn` (sync each turn), `session` (flush on end), or integer N (every N turns) |
|
||||
| `observation` | all on | Per-peer `observeMe`/`observeOthers` booleans |
|
||||
|
||||
**Session strategy** controls how Honcho sessions map to your work:
|
||||
- `per-session` — each `hermes` run gets a fresh session. Clean starts, memory via tools. Recommended for new users.
|
||||
- `per-directory` — one Honcho session per working directory. Context accumulates across runs.
|
||||
- `per-repo` — one session per git repository.
|
||||
- `global` — single session across all directories.
|
||||
|
||||
**Recall mode** controls how memory flows into conversations:
|
||||
- `hybrid` — context auto-injected into system prompt AND tools available (model decides when to query).
|
||||
- `context` — auto-injection only, tools hidden.
|
||||
- `tools` — tools only, no auto-injection. Agent must explicitly call `honcho_reasoning`, `honcho_search`, etc.
|
||||
|
||||
**Dialectic cadence** controls cost. With default `3`, Honcho rebuilds the user model every 3 turns instead of every turn — ~66% fewer LLM calls without losing model fidelity.
|
||||
|
||||
**Settings per recall mode:**
|
||||
|
||||
| Setting | `hybrid` | `context` | `tools` |
|
||||
|---------|----------|-----------|---------|
|
||||
| `writeFrequency` | flushes messages | flushes messages | flushes messages |
|
||||
| `dialecticCadence` | gates auto LLM calls | gates auto LLM calls | irrelevant — model calls explicitly |
|
||||
| `contextTokens` | caps injection | caps injection | irrelevant — no injection |
|
||||
| `dialecticDynamic` | gates model override | N/A (no tools) | gates model override |
|
||||
|
||||
In `tools` mode, the model is fully in control — it calls `honcho_reasoning` when it wants, at whatever `reasoning_level` it picks. `dialecticCadence` and `contextTokens` only apply to modes with auto-injection (`hybrid` and `context`).
|
||||
|
||||
## Tools
|
||||
|
||||
When Honcho is active as the memory provider, four additional tools become available:
|
||||
When Honcho is active as the memory provider, five tools become available:
|
||||
|
||||
| Tool | Purpose |
|
||||
|------|---------|
|
||||
| `honcho_conclude` | Trigger server-side dialectic reasoning on recent conversations |
|
||||
| `honcho_context` | Retrieve relevant context from Honcho's memory for the current conversation |
|
||||
| `honcho_profile` | View or update the user's Honcho profile |
|
||||
| `honcho_search` | Semantic search across all stored conclusions and observations |
|
||||
| `honcho_profile` | Read or update peer card — pass `card` (list of facts) to update, omit to read |
|
||||
| `honcho_search` | Semantic search over context — raw excerpts, no LLM synthesis |
|
||||
| `honcho_context` | Full session context — summary, representation, card, recent messages |
|
||||
| `honcho_reasoning` | Synthesized answer from Honcho's LLM — pass `reasoning_level` (minimal/low/medium/high/max) to control depth |
|
||||
| `honcho_conclude` | Create or delete conclusions — pass `conclusion` to create, `delete_id` to remove (PII only) |
|
||||
|
||||
## CLI Commands
|
||||
|
||||
```bash
|
||||
hermes honcho status # Show connection status and config
|
||||
hermes honcho status # Connection status, config, and key settings
|
||||
hermes honcho setup # Interactive setup wizard
|
||||
hermes honcho strategy # Show or set session strategy
|
||||
hermes honcho peer # Update peer names for multi-agent setups
|
||||
hermes honcho mode # Show or set recall mode
|
||||
hermes honcho tokens # Show or set context token budget
|
||||
hermes honcho identity # Show Honcho peer identity
|
||||
hermes honcho sync # Sync host blocks for all profiles
|
||||
hermes honcho enable # Enable Honcho
|
||||
hermes honcho disable # Disable Honcho
|
||||
```
|
||||
|
||||
## Migrating from `hermes honcho`
|
||||
|
||||
@@ -51,7 +51,7 @@ AI-native cross-session user modeling with dialectic Q&A, semantic search, and p
|
||||
| **Data storage** | Honcho Cloud or self-hosted |
|
||||
| **Cost** | Honcho pricing (cloud) / free (self-hosted) |
|
||||
|
||||
**Tools:** `honcho_profile` (peer card), `honcho_search` (semantic search), `honcho_context` (LLM-synthesized), `honcho_conclude` (store facts)
|
||||
**Tools:** `honcho_profile` (read/update peer card), `honcho_search` (semantic search), `honcho_context` (session context — summary, representation, card, messages), `honcho_reasoning` (LLM-synthesized), `honcho_conclude` (create/delete conclusions)
|
||||
|
||||
**Setup Wizard:**
|
||||
```bash
|
||||
@@ -74,10 +74,12 @@ hermes memory setup # select "honcho"
|
||||
| `workspace` | host key | Shared workspace ID |
|
||||
| `recallMode` | `hybrid` | `hybrid` (auto-inject + tools), `context` (inject only), `tools` (tools only) |
|
||||
| `observation` | all on | Per-peer `observeMe`/`observeOthers` booleans |
|
||||
| `writeFrequency` | `async` | `async`, `turn`, `session`, or integer N |
|
||||
| `writeFrequency` | `async` | When to flush messages to Honcho: `async` (background thread), `turn` (sync each turn), `session` (flush on end), or integer N (every N turns) |
|
||||
| `sessionStrategy` | `per-directory` | `per-directory`, `per-repo`, `per-session`, `global` |
|
||||
| `contextTokens` | uncapped | Token budget for auto-injected context per turn. Set to an integer (e.g. 1200) to cap |
|
||||
| `dialecticReasoningLevel` | `low` | `minimal`, `low`, `medium`, `high`, `max` |
|
||||
| `dialecticDynamic` | `true` | Auto-bump reasoning by query length |
|
||||
| `dialecticDynamic` | `true` | When `true`, model can override reasoning level per-call via tool param |
|
||||
| `dialecticCadence` | `3` | Turns between Honcho LLM calls — only applies to `hybrid`/`context` modes. In `tools` mode the model calls `honcho_reasoning` explicitly |
|
||||
| `messageMaxChars` | `25000` | Max chars per message (chunked if exceeded) |
|
||||
|
||||
</details>
|
||||
@@ -165,6 +167,7 @@ This inherits settings from the default `hermes` host block and creates new AI p
|
||||
},
|
||||
"dialecticReasoningLevel": "low",
|
||||
"dialecticDynamic": true,
|
||||
"dialecticCadence": 3,
|
||||
"dialecticMaxChars": 600,
|
||||
"messageMaxChars": 25000,
|
||||
"saveMessages": true
|
||||
|
||||
Reference in New Issue
Block a user