Web News Scraping (stdlib-only toolkit)
Status: Active Tags: Tooling NLP Agentic Created: 2026-08-30 Last updated: 2026-09-14 Related: hermes-cron-operations, cloudflare-pages-deploy
Definition
A set of five dependency-free Python scripts (stdlib only: urllib, re, json, xml.etree) for researching fresh web news without any search API key. Lives in /home/romain/workspace/. Built 2026-08-30 by the Daily News cron session as an API-key-free fallback when the hosted web_search tool path is slow or unavailable.
How It Works
| Script | Input | Output |
|---|---|---|
bing_search.py | query(s) | Bing HTML web search scrape; freshness via qft=interval%3d%227%22 (24 h) / %228%22 (7 d); parses <li class="b_algo"> blocks → (title, url, snippet, age) |
bing_news.py | query(s) | Bing News vertical scrape; parses news-card divs (triple-regex fallback for attribute order) → (title, source, url, snippet, age) |
ld_headlines.py | URL(s) | fetches page HTML, extracts JSON-LD nodes of type Article/NewsArticle/ReportageNewsArticle/WebPage/BlogPosting → deduped (headline, date, url) |
ld_local.py | local HTML file(s) | same JSON-LD extraction offline (for saved/archived pages) |
rss_scan.py | RSS 2.0 / RDF (DW) / Atom file(s) | parses all three formats, multi-format date parsing, prints items newer than N hours, newest first |
rss_probe.py, rss_batch*.py | live RSS URLs | batched multi-outlet RSS probing (added 2026-09-14 by the news cron) — probe outlet feeds for freshness, then batch-dump the fresh ones to disk for the agent to summarize |
Common patterns:
- Browser
User-Agentspoofing (Chrome 126 string) on every request. - Regex-based HTML parsing, no
requests/bs4— each parse is wrapped so failures are skipped, not fatal. - Relative age strings (“N hours ago”) are captured as-is; RSS dates are normalized to UTC before the freshness cutoff.
- CLI-first: each script is runnable directly (
python3 bing_search.py 1 "query"), which makes them usable from agent terminal calls without import plumbing.
Key Parameters
| Item | Value |
|---|---|
| Location | /home/romain/workspace/{bing_search,bing_news,ld_headlines,ld_local,rss_scan}.py |
| Dependencies | Python stdlib only (no venv needed) |
| Freshness filter (Bing) | qft=interval="7" = 24 h, interval="8" = 7 days |
| RSS cutoff | rss_scan.py <hours> <files...> |
When To Use
- Research step of news/deals cron jobs when
web_searchis too slow, rate-limited, or the job must avoid hosted-tool latency (see hermes-cron-operations — the news job has died twice on API timeouts). - Verifying a story’s real outlet coverage via JSON-LD of a known article page.
- Bulk freshness check of a directory of saved feeds with
rss_scan.py.
RSS-first research mode (as of 2026-09-14)
- The Daily News cron’s research step is now RSS-first: Google/Bing/Reuters/NYT are blocked by the sandbox network, so the 09-14 run built
rss_probe.py+rss_batch*.pyand pulled all 30 cards from live RSS of these 19 outlets: France24, The Hill, UN News, TASS, RT, Global News (Canada), NPR, TechCrunch, Ars Technica, Wired, CNBC, MarketWatch, Climate Home, Inside Climate News, Electrek, Phys.org, ScienceDaily, Variety, The Hollywood Reporter. - Pattern: probe each outlet’s feed for items fresh enough, batch-dump the fresh items to disk (spill research to files — keep context lean per the cron hardening rules), then have the agent summarize per category.
- This supersedes the earlier curl-against-hub-pages fallback (09-03) as the primary research path;
web_searchremains unavailable in cron sandboxes.
Risks & Pitfalls
- Regex HTML parsing is fragile: Bing class names (
b_algo,news-card) change without notice — if a script suddenly returns 0 items, assume markup drift, not empty results. - Headline/snippet only — no full-article extraction; use
ld_headlines.pyon the article URL for structured metadata, but body text still needs a fetch+strip pass. rss_scan.pystrips<!DOCTYPE>beforeElementTreeparsing (required for feeds that carry it); other strict-XML quirks will surface as PARSE FAIL lines.- Heavy use of the same UA may get IP-throttled by Bing; keep request counts modest (the cron pattern is a handful of queries per run).
ld_headlines.pyincludesWebPagetype — expect some noise beyond real articles;ld_local.pydeliberately omitsWebPage.
Related Concepts
- hermes-cron-operations — why this toolkit exists (cron failure modes) and how it’s driven
- cloudflare-pages-deploy — the consumer pipeline (news site) this feeds
- llm.md — the agent that consumes the scraped research
Sources
/home/romain/workspace/bing_search.py,bing_news.py,ld_headlines.py,ld_local.py,rss_scan.py(created 2026-08-30 04:14–04:25 by cron sessioncron_eef1a69519af_20260830_041056)- Cron job
eef1a69519af(Daily News Refresh + Deploy) prompt + failure reportcron/output/eef1a69519af/2026-08-30_04-41-46.md