Honcho — Self-Hosted Deployment
Status: Active (deriver fixed 2026-08-29) Tags: Deployment Infrastructure Service Tooling Created: 2026-08-30 Related: local-llm-stack
Definition
Self-hosted Honcho (honcho-ai 2.2.0) agent-memory service on this host, backed by the local llama.cpp stack and Postgres/pgvector. API on 127.0.0.1:8003, workspace id hermes, honcho.json timeout 300.
How It Works
- Docker-compose deploy at
/media/data/honcho(config in.env, fully local — nothing leaves the network). - LLM: local Qwen3.8-27B via OpenAI-compatible API (
http://host.docker.internal:8000/v1inside containers). - Vector store: pgvector; embeddings via a local nomic-embed-text service (
http://embedding:8000/v1inside the compose network), 768 dims. - Deriver:
DERIVER_WORKERS=1,DERIVER_FLUSH_ENABLED=true(process immediately, no token-threshold batching). - Test scripts (pattern reference):
/home/romain/workspace/honcho_e2e_test.py(workspace→peers→session→messages→read-back),honcho_deriver_test.py(pollqueue_status()until work units drain),honcho_readback.py(verify conclusions + raw DB document counts).
Key Parameters
The deriver silently produces zero observations without all three of these (the 2026-08-29 root causes, documented in /media/data/honcho/.env):
EMBEDDING_MODEL_CONFIG__OVERRIDES__BASE_URL=http://embedding:8000/v1— without a base_url,save_representationfell back to api.openai.com and timed out → documents table stayed at 0.EMBEDDING_VECTOR_DIMENSIONS=768+ one-offconfigure_embeddings.pyrun to resize the pgvector column.- Qwen thinking mode off for all Honcho LLM calls: llama-server needs a real JSON boolean, so each override is set as one JSON env var:
*_MODEL_CONFIG__OVERRIDES__PROVIDER_PARAMS={"extra_body":{"chat_template_kwargs":{"enable_thinking":false}}}(nested leaf env vars arrive as strings → HTTP 400). PlusDERIVER_MODEL_CONFIG__MAX_OUTPUT_TOKENS=16384— the 2500 default left empty content on a ~2.5 tok/s reasoning model.
When To Use
- Agent memory / user-model service for Hermes profiles on this host.
- Reusable pattern: any self-host of Honcho (or similar LLM pipeline) on a local reasoning LLM.
Risks & Pitfalls
- Reasoning models burn the output budget on thinking traces before the JSON answer → set high max tokens AND disable thinking (see parameters above).
- llama.cpp image has no python3 — use bash healthchecks.
- Re-enable thinking by deleting the
PROVIDER_PARAMSlines. - Embedding model changes require re-running dimension configuration; stale vector dim = silent empty retrieval.
Related Concepts
- local-llm-stack — LLM + embedding backend
- cloudflare-pages-deploy — other local infra pattern on this host
Sources
/media/data/honcho/.env(authoritative config + root-cause comments, 2026-08-29)/home/romain/workspace/honcho_*.pytest scripts