Honcho — Self-Hosted Deployment

Status: Active (deriver fixed 2026-08-29) Tags: Deployment Infrastructure Service Tooling Created: 2026-08-30 Related: local-llm-stack

Definition

Self-hosted Honcho (honcho-ai 2.2.0) agent-memory service on this host, backed by the local llama.cpp stack and Postgres/pgvector. API on 127.0.0.1:8003, workspace id hermes, honcho.json timeout 300.

How It Works

  • Docker-compose deploy at /media/data/honcho (config in .env, fully local — nothing leaves the network).
  • LLM: local Qwen3.8-27B via OpenAI-compatible API (http://host.docker.internal:8000/v1 inside containers).
  • Vector store: pgvector; embeddings via a local nomic-embed-text service (http://embedding:8000/v1 inside the compose network), 768 dims.
  • Deriver: DERIVER_WORKERS=1, DERIVER_FLUSH_ENABLED=true (process immediately, no token-threshold batching).
  • Test scripts (pattern reference): /home/romain/workspace/honcho_e2e_test.py (workspace→peers→session→messages→read-back), honcho_deriver_test.py (poll queue_status() until work units drain), honcho_readback.py (verify conclusions + raw DB document counts).

Key Parameters

The deriver silently produces zero observations without all three of these (the 2026-08-29 root causes, documented in /media/data/honcho/.env):

  1. EMBEDDING_MODEL_CONFIG__OVERRIDES__BASE_URL=http://embedding:8000/v1 — without a base_url, save_representation fell back to api.openai.com and timed out → documents table stayed at 0.
  2. EMBEDDING_VECTOR_DIMENSIONS=768 + one-off configure_embeddings.py run to resize the pgvector column.
  3. Qwen thinking mode off for all Honcho LLM calls: llama-server needs a real JSON boolean, so each override is set as one JSON env var: *_MODEL_CONFIG__OVERRIDES__PROVIDER_PARAMS={"extra_body":{"chat_template_kwargs":{"enable_thinking":false}}} (nested leaf env vars arrive as strings → HTTP 400). Plus DERIVER_MODEL_CONFIG__MAX_OUTPUT_TOKENS=16384 — the 2500 default left empty content on a ~2.5 tok/s reasoning model.

When To Use

  • Agent memory / user-model service for Hermes profiles on this host.
  • Reusable pattern: any self-host of Honcho (or similar LLM pipeline) on a local reasoning LLM.

Risks & Pitfalls

  • Reasoning models burn the output budget on thinking traces before the JSON answer → set high max tokens AND disable thinking (see parameters above).
  • llama.cpp image has no python3 — use bash healthchecks.
  • Re-enable thinking by deleting the PROVIDER_PARAMS lines.
  • Embedding model changes require re-running dimension configuration; stale vector dim = silent empty retrieval.

Sources

  • /media/data/honcho/.env (authoritative config + root-cause comments, 2026-08-29)
  • /home/romain/workspace/honcho_*.py test scripts