Hermes Agent Cron Operations (this host)
Status: Active Tags: Tooling Infrastructure Agentic Created: 2026-08-30 Last updated: 2026-10-01 Related: web-news-scraping-stdlib, cloudflare-pages-deploy, local-llm-stack
Definition
Operational patterns, limits, and observed failure modes for scheduled Hermes Agent cron jobs running on this host (profile romain). The recurring theme: agent-driven research jobs are bounded by several hard timeouts, and jobs must be written so they never sit idle.
How It Works
Hard limits (observed failures)
| Limit | Symptom | Observed incident |
|---|---|---|
| Cron idle timeout: 600 s | TimeoutError: Cron job '...' idle for 600s (limit 600s) — job killed if no tool activity for 600 s | Daily News job, 2026-08-30 04:41 (after building the scraping toolkit, session went idle); 2026-09-04 04:42 (reported as “idle for 603 s”; last activity a completed patch call — the run had made substantial progress and left partial artifacts in the workspace) |
| Non-streaming API call: 240 s | RuntimeError: Non-streaming API call timed out after 240s with no response | Daily News job, 2026-08-22 15:07 |
| Stale-token floor (provider) | Long agent runs killed when the provider’s stale check (default 180 s on the qwen3 slug) fires mid-run | Fixed 2026-08-30: providers.custom.stale_timeout_seconds=900 in config — note the CLI splits dotted model names, so the key is provider-wide, not per-model |
| LLM API connection error | RuntimeError: Connection error. — API connection fails before the agent does any work; output .md contains only prompt + error (~4 KB), no partial state | Daily News job, 3 consecutive days: 2026-08-31, 09-01, 09-02 (05:41) |
| web_search loop guardrail | loop_web_search_cap fires after ~50 non-progressing web_search calls — agent must change strategy (direct curl / local toolkit), not retry the same query | Daily News re-run 2026-08-30 11:11 — agent then fell back to curl, wrote content, built and deployed successfully |
| Response length limit | RuntimeError: Response truncated due to output length limit — the run dies when a model response hits an output length cap; the output .md contains only prompt + error (same shape as a connection error), no partial report. First observed; root cause not yet pinned (long final report vs. provider cap — both plausible) | Daily News fire-now re-run 2026-09-05 18:26 (the same morning’s scheduled run had died on idle 601 s); site unaffected |
| Cron sandbox: no network / no web_search tool | The cron sandbox lost all external network access (all interfaces NO-CARRIER, DNS failure) and the hermes_tools Python module is not installed in the cron venv. The agent cannot research news or deploy to Cloudflare Pages. Observed on 2026-09-25 — a fundamental environment regression; likely triggered by a recent change (Python 3.14 upgrade, Hermes Agent version update, or container network config). | |
| Generic request timeout | RuntimeError: Request timed out. — a generic HTTP-level request timeout (not the 600 s idle limit, not the 240 s API timeout). The run produced no partial state. First observed 2026-09-28 — the error string is too terse to classify further; may be a network-level timeout or a provider-side timeout. |
Job-authoring best practices
- Self-contained prompts: absolute paths, explicit workdir, exact file schemas, exact verify steps. Cron sessions have no conversation history.
no_agentfor deterministic work — don’t pay LLM latency for fixed pipelines.- Pin model + provider at
hermes cron create— thecronjobtool cannot set them later. [SILENT]in the prompt for monitor-style jobs with nothing to report.- Keep total research steps few: 600 s idle + 240 s API timeouts mean a 6-category × N-query research job is near the ceiling; the 2026-08-30 news run built a local fallback (web-news-scraping-stdlib) precisely to cut hosted-search latency.
- Job output + failure reports land in
~/.hermes/profiles/<p>/cron/output/<job-id>/— check the latest.mdthere when a job “fails” to see the exact error.
Prompt-hardening patterns (validated 2026-09-03/04)
Both the Daily News (eef1a69519af) and Weekly Deals (7f0264357282) prompts were hardened after repeated timeout deaths, and verification (fire-now) runs then completed end-to-end: news 2026-09-04 01:56 (deployed, verified 200/30 cards first try) and deals 2026-09-03 23:35 (+9 deals, deployed, HTTP 200). The patterns that made the difference:
- Pre-flight validation step (STEP 0) before any research — e.g. import-check the content files and auto-repair (
fix_quotes.py) if broken, so a killed prior run’s half-written state is healed before new work starts. - Atomic file writes — write to
<file>.new, validate (ast.parse/ JSON parse), only thenmvover the original. A run killed mid-write can no longer leave a broken file behind. - Context-bloat rules — research ONE unit (category/angle) at a time and write results to disk immediately; cap search queries (~15–20); spill raw research notes to a scratch file (
/tmp/deals_research.md) instead of holding search dumps in context; keep assistant turns short. - Explicit tool fallbacks — state the goal, not a required tool (curl when
web_searchis absent).
Evidence: the same jobs had died on 240 s API timeouts (context bloat, deals 08-29) and 600 s idle (news 08-30) before hardening; after it, both verification runs finished. Caveat: hardening reduced but did not eliminate failures — the 09-04 04:42 news run still died on the idle limit (see below).
Gateway / service ops
- The gateway rewrites its systemd unit from a template on every start — direct edits to the unit file silently revert. Persist changes with drop-in files:
<unit>.service.d/*.conf. - Active log files (agent.log, gateway.log, errors.log, weixin-relay.log) must never be deleted; old
*.log*get gzipped by the daily maintenance job. - Profile topology:
default+jiayiprofiles are dormant (no gateway/creds); all WeChat/DingTalk credentials exist only in theromainprofile (WEIXIN_*in its.env; weixin adapter in the gateway, accounta8a52773).
Key Parameters
| Item | Value |
|---|---|
| Cron idle limit | 600 s |
| Non-streaming API timeout | 240 s |
| Stale timeout (fixed) | providers.custom.stale_timeout_seconds=900 |
| Output dir | ~/.hermes/profiles/romain/cron/output/<job-id>/ |
| Jobs registry | ~/.hermes/profiles/romain/cron/jobs.json |
Current jobs (snapshot 2026-09-14)
| ID | Name | Schedule (CST) |
|---|---|---|
eef1a69519af | Daily News Refresh + Deploy | daily 04:10 |
7f0264357282 | Weekly Deep Deals Research + Deploy | Sat 04:30 |
4f90ecb2a611 | Weekly Memory Compaction | Sun 03:00 |
c6def452f00f | Daily BESS News Feed Refresh | daily 05:15 |
2cd6faaa0f8a | Daily Wiki + Workspace Maintenance | daily 05:00 |
efcedb24cce0 | wiki-index-to-honcho | every 360 min |
532b734cea6f | outlook-email-poll | every 360 min |
When To Use
- Authoring or debugging any cron job on this host.
- When a cron job fails with a timeout: identify which of the three limits above fired (the error string in the output
.mdtells you). - When changing gateway service config — use drop-ins, never the unit file.
Risks & Pitfalls
- A job that does one very long terminal command (e.g. a 10+ min deploy) is fine (tool is active) — the idle limit counts time between tool completions; what kills jobs is the model stalling between tools or a hung non-streaming call.
- The stale-timeout fix is provider-wide: raising it affects every model served by
providers.custom(see local-llm-stack). - The jobs table above is a snapshot — re-check
jobs.jsonbefore relying on IDs. - A connection-error failure is the cleanest kind: nothing was executed, so a plain re-run is safe (no half-written files to reconcile). Contrast with an idle timeout, which can leave content files half-written.
- An idle kill can land mid-write even on a run that was progressing (2026-09-04 04:42 news run: killed “idle 603 s” after a
patchcall, leaving rewritten content files, a newnews_assemble.py, and partial per-card JSONs innews_cards/). The live site was unaffected (previous deploy still up), but before re-running or resuming, check the workspace’s on-disk state — atomic writes (.new→ validate →mv) make the broken-file case impossible, though partial new artifacts can still accumulate. - The maintenance job itself is not immune to the idle limit. It died on idle timeouts two days running — 2026-09-04 06:08 (“idle 600 s”) and 2026-09-05 06:02 (“idle 603 s”, last activity a completed
patch) — each time leaving wiki work on disk without the index/log/commit steps (09-05: a whole new pagerealtime-voice-assistant.md+ three user-layer summary edits, uncommitted, missing from both index.md files and both log.md files). Recovery recipe, applied 2026-09-06: ①git -C /media/data/wiki-common statusfor uncommitted/untracked pages; ② compare each layer’sindex.mdagainst the filesystem (find wiki -name '*.md'vs index entries); ③ read the interrupted run’scron/output/<id>/<latest>.mdfor what it intended; ④ append the missing index entries + log entries (log LAST, as always); ⑤ commit. Keep wiki-editing turns short (one page per tool call, no long synthesis passes between calls) so the run never sits idle >600 s mid-batch. |web_searchis not always present in cron sandboxes. Two distinct failure modes: (1) 2026-09-03 — noweb_searchbinary but network worked; agent succeeded viacurl(caching ~45 articles to/tmp/news/). (2) 2026-09-25 — complete environment isolation: nohermes_toolsmodule and no network at all (all interfaces NO-CARRIER, DNS fails). The sandbox cannot research news or deploy to Cloudflare Pages. Root cause: likely a recent environment change (Python 3.14 upgrade, Hermes Agent version update, or container network config). Prompts should not hard-require a specific search tool — state the goal and allow fallbacks (web-news-scraping-stdlib). - Job scripts can rot independently of the cron runtime — not every repeated cron failure is a timeout.
outlook-email-poll(532b734cea6f) has failed every run since 2026-09-11: firstFileNotFoundError: .outlook_credentials.json, laterImportError: cannot import name 'TENANT_ID' from 'outlook_graph'—email_poller.pyimports a constant the module never defined (code drift, exits immediately, exit 1). Diagnose by reading the newest run report incron/output/<id>/; the error string tells you whether it’s a runtime limit vs. a script bug. (Full state in users/romain wiki:summaries/outlook-email-pipeline.md.) - wiki-index-to-honcho (
efcedb24cce0) has a silent skip bug: the Honcho API caps message content at 25,000 chars and the index script does not chunk, sopengcheng-fashion-exhibition.md(109 KB) has never been indexed (HTTP 422string_too_long). Worse, the skip logic keys on mtime only — a file whose last attempt errored is skipped forever until its mtime changes. Fix (not yet applied; script lives under the default profile’s~/.hermes/cron/scripts/so the romain agent’s cross-profile guard blocks edits): chunk long files into ≤ ~24,000-char messages and change the skip condition tostatus == "indexed" AND mtime matches.
Related Concepts
- web-news-scraping-stdlib — latency-reduction fallback born from these failures
- cloudflare-pages-deploy — the deploy step shared by news/deals jobs
- local-llm-stack — the local provider these limits interact with
Sources
cron/output/eef1a69519af/2026-08-30_04-41-46.md(idle timeout) and2026-08-22_15-07-24.md(API timeout)cron/output/eef1a69519af/2026-09-04_01-56-52.md(hardened-prompt success) and2026-09-04_04-42-13.md(idle 603 s, last activity apatchcall)cron/output/7f0264357282/2026-08-29_18-00-59.md(240 s timeout, context bloat) and2026-09-03_23-35-08.md(hardened-prompt success, +9 deals)cron/output/eef1a69519af/2026-09-05_18-26-21.md(response-truncation error) and2026-09-06_05-10-04.md(scheduled success after 09-05 double failure)cron/output/2cd6faaa0f8a/2026-09-04_06-08-05.mdand2026-09-05_06-02-52.md(maintenance job idle kills, partial wiki state left behind)cron/jobs.json, profilememories/MEMORY.md+USER.md(compacted 2026-08-30)logs/agent.logcron scheduler entries, 2026-08-30 and 2026-09-04 (run tagscron_<job>_<ts>, per-call latency + tool activity)