--- url: /guide/getting-started.md --- # Getting started Neurowire turns any blog, website, RSS, or Atom feed into clean, modern feeds. Point it at a URL and get back **NWF** (a compact custom format), **Atom**, **RSS 2.0**, **JSON Feed 1.1**, or **Markdown**. Bundle many sources into one **mesh**, group meshes into a **construct**, keep the history in an append-only **journal**, and render any feed, mesh, or construct into a self-contained **HTML news page**. ## Three surfaces Neurowire is one toolkit you can reach through three doors: * **Library**: the npm packages (`@neurowire/core`, `@neurowire/ingest`, `@neurowire/taps`, `@neurowire/web`). Fetch, parse, merge, and serialize feeds from your own code. See [Library usage](/guide/library). * **CLI**: the `neurowire` binary. Point it at a URL, get a terminal view or a serialized feed, bundle meshes, watch for new posts, push to Slack. See [CLI reference](/guide/cli). * **HTTP API**: a small Hono service exposing `GET /feed`, `/mesh`, and `/construct`. Run it yourself and call it over HTTP. See [HTTP API](/guide/http-api). The data model is the same across all three: every parser produces a single canonical feed shape, and every serializer consumes it. Learn the model in [The model](/concepts/model). ## 60-second quickstart ### 1. Install the CLI ::: code-group ```bash [pnpm] pnpm add -g @neurowire/cli ``` ```bash [npm] npm install -g @neurowire/cli ``` ::: ::: tip Node 24+ Neurowire is ESM-only and needs Node >= 24. See [Installation](/guide/installation) for the full matrix. ::: ### 2. Point it at a URL ```bash # a terminal view (no --format) neurowire https://blog.rust-lang.org/feed.xml # or a serialized feed to stdout neurowire https://blog.rust-lang.org/feed.xml --format atom > rust.xml ``` Neurowire detects whether the URL is a feed or an HTML page. For an HTML page it follows a declared feed link, falls back to a curated per-host recipe (a [tap](/concepts/taps)), then to heuristic auto-detection. ### 3. Bundle a mesh A [mesh](/concepts/meshes) is a named bundle of sources, fetched in parallel and merged into one newest-first feed. Drop this into `ai-news.json`: ```json { "name": "AI News", "sources": [ { "name": "Claude Blog", "url": "https://claude.com/blog" }, { "name": "Simon Willison", "url": "https://simonwillison.net/atom/everything/" } ] } ``` Then fetch it: ```bash neurowire --mesh ai-news.json --format json --limit 10 ``` ### 4. Follow it live A mesh does not have to be something you re-run. `tail` polls it and prints each new entry the moment it appears, until you stop it: ```bash neurowire tail --mesh ai-news.json --interval 60s ``` The first tick prints what is on the front page right now, one entry per line, then the command goes quiet and only speaks when something new lands: ``` 09:41:02 Claude Code 2.0 Claude Blog ยท 2026-08-27 https://claude.com/blog/claude-code-2 ``` Add `-f nwf` to stream raw [NWFJ](/formats/nwfj) records instead, which is what you want when the output is going into another tool rather than your eyes. See [Tail](/concepts/tail). ### 5. Render an HTML page Install the page generator and turn a mesh into a self-contained HTML page (all CSS inline, no external requests): ::: code-group ```bash [pnpm] pnpm add -g @neurowire/web neurowire-web --mesh ai-news.json --out index.html ``` ```bash [npm] npm install -g @neurowire/web neurowire-web --mesh ai-news.json --out index.html ``` ::: ## Sites with no feed Some sites publish a blog and ship no feed at all. A [tap](/concepts/taps) teaches Neurowire to read one: a small set of CSS selectors saying where each article sits on the listing page. You do not write them by hand, `tap wizard` proposes them from the page's structure and verifies the result before saving: ```bash # author a tap (--yes takes the top suggestion for every field) neurowire tap wizard https://example.com/blog --yes -o example.json # the site now reads like any feed neurowire https://example.com/blog --taps example.json --format atom ``` Drop the tap into `~/.config/neurowire/taps/` (where the wizard writes by default) and the `--taps` flag becomes unnecessary. No model is involved and no API key is needed, at any point. ## Pull from a peer instead of fetching If someone already runs a Neurowire node that journals these sources, you do not have to fetch them yourself. Register the peer once and pull: ```bash neurowire peers add https://hub.example.com neurowire sync --peers # ai: 1284 new, cursor 1284, 4 requests, 812.0 KB ``` Later syncs move only what is new, and cost one request when nothing is. What you get back is an ordinary [journal](/concepts/journals), so `neurowire journal cat ai -f json` and friends work on it immediately. This matters most for a laptop that is closed most of the day: a live fetch can only show you what is on the front page right now, while a node that stayed awake recorded everything you missed. See [Sync](/concepts/sync) for why, and [Federation](/guide/federation) for how to run the node. ## Where to next * [CLI reference](/guide/cli): every flag and subcommand. * [The model](/concepts/model): the canonical feed shape. * [Atom format](/formats/atom): the default serializer. * [Journals](/concepts/journals): keep an append-only archive of what a source publishes, and query it back. * [Tail](/concepts/tail): follow a source as a live stream, in the terminal or over SSE. * [Sync](/concepts/sync): pull journal deltas from a peer instead of re-fetching every source yourself. * [Recipes](/guide/recipes): practical end-to-end workflows. * [Taps](/concepts/taps): read sites that ship no feed, and keep those readers working when the site is redesigned. --- --- url: /guide/installation.md --- # Installation Neurowire is published as a set of scoped npm packages. Install only the ones you need for your use case. ## Requirements * **Node >= 24.** Every package declares `engines.node: ">=24"`. * **ESM only.** All packages are `"type": "module"`. Import them with `import`, not `require`. ## Published packages | Package | Version | Role | |---------|---------|------| | `@neurowire/core` | 0.7.0 | Canonical model, serializers (NWF, atom, rss, json, md), `validateNwf`, `mergeFeeds`, mesh/construct types. Pure, no network. | | `@neurowire/ingest` | 0.6.0 | Fetch + detect + parse, HTML auto-detect, the CSS-template engine, `fetchFeed`/`fetchMesh`/`fetchConstruct`. | | `@neurowire/taps` | 0.3.0 | Curated per-host templates for feed-less sites plus loaders. | | `@neurowire/tap-wizard` | 0.1.0 | Tap authoring and healing: candidate selectors, preview, and the verification gate. No model, no API key. | | `@neurowire/cli` | 0.7.0 | The `neurowire` binary. | | `@neurowire/web` | 0.5.0 | HTML page generator: `toHtml` plus the `neurowire-web` binary. | ::: tip Dependency direction Dependencies flow strictly one way: `core` <- `ingest` <- (`taps`, `tap-wizard`) <- (`cli`, `web`). Higher packages pull in everything below them, so you rarely install more than one for a given job. `taps` and `tap-wizard` are siblings and know nothing of each other: one carries taps, the other authors them. ::: ## Which package do I install? ### CLI users Install the CLI globally and use the `neurowire` command. It bundles core, ingest, taps, and tap-wizard, so `tap wizard`, `tap check`, and `tap heal` work out of the box with nothing else to install. ::: code-group ```bash [pnpm] pnpm add -g @neurowire/cli ``` ```bash [npm] npm install -g @neurowire/cli ``` ::: ### Page generators For self-contained HTML news pages, install the web package. It ships the `neurowire-web` binary and the `toHtml` / `toConstructHtml` / `toConstructPages` functions. ::: code-group ```bash [pnpm] pnpm add -g @neurowire/web ``` ```bash [npm] npm install -g @neurowire/web ``` ::: ### Library users Add packages to your project as dependencies (not global). Most programmatic work needs `@neurowire/ingest` (which re-exports nothing from core, so add core too for the types and serializers): ::: code-group ```bash [pnpm] pnpm add @neurowire/core @neurowire/ingest ``` ```bash [npm] npm install @neurowire/core @neurowire/ingest ``` ::: Add `@neurowire/tap-wizard` when your own code needs to author, verify, or heal taps rather than just use them. It is what the CLI's tap commands are built on, and it pulls in nothing beyond `core`, `ingest`, and cheerio: ::: code-group ```bash [pnpm] pnpm add @neurowire/tap-wizard ``` ```bash [npm] npm install @neurowire/tap-wizard ``` ::: See [`@neurowire/tap-wizard`](/reference/tap-wizard) for the API. Add `@neurowire/taps` when you want curated recipes for feed-less sites (call `registerAllTaps()`), and `@neurowire/web` when you want HTML rendering: ::: code-group ```bash [pnpm] pnpm add @neurowire/taps @neurowire/web ``` ```bash [npm] npm install @neurowire/taps @neurowire/web ``` ::: See [Library usage](/guide/library) for code examples. ## Verify the install ```bash neurowire --version ``` --- --- url: /guide/cli.md --- # CLI reference The `neurowire` binary (from `@neurowire/cli`) turns a URL, mesh, or construct into a terminal view or a serialized feed, with filtering, sorting, time windows, a live tail, a watch loop, and delivery sinks. ## Synopsis ``` neurowire [options] neurowire --mesh [options] neurowire --construct [options] neurowire validate neurowire tap wizard [-o file] [--yes] neurowire tap check [path] [--all] [--json] [--url ] neurowire tap heal [--yes] [--url ] neurowire tap doctor neurowire opml export --mesh |--construct [-o out.opml] neurowire opml import [-o mesh.json] [--name ] neurowire journal head|cat|query [options] neurowire tail [url] [options] ``` With no `--format`, Neurowire prints a colorized terminal view. With `--format` it serializes the feed to stdout (or to `--out`). ::: warning Dev vs built binary: the `--` gotcha When you run the CLI through pnpm in this repo (`pnpm cli -- `), pass flags **after** `--`, otherwise pnpm eats `-f` as its own `--filter`. The CLI also tolerates one leading `--` that pnpm/tsx may inject. The installed binary needs no `--`: ```bash # dev (in the monorepo) pnpm cli -- --mesh ai-news.json -f atom # installed binary neurowire --mesh ai-news.json -f atom ``` ::: ## Source selection | Flag | Description | |------|-------------| | `` (positional) | A website, RSS, or Atom URL. Auto-detected and normalized. | | `-m, --mesh ` | Fetch a mesh: a JSON bundle of named sources, merged into one feed. | | `-c, --construct ` | Fetch a construct: a bundle of meshes. The terminal view keeps the per-mesh grouping; `--format` flattens it into one feed. | | `-t, --template ` | Path to a JSON CSS-selector template, forcing on-page extraction instead of auto-detect (positional URL only). | | `--header ` | Extra request header for the positional URL fetch, e.g. `--header 'Authorization: Bearer ...'`. Repeatable. Split on the first colon. Mesh files set headers per source instead, with `${ENV_VAR}` references resolved at load; see [Meshes](/concepts/meshes#private-sources-per-source-headers). | `{ ref }` members in a construct (mesh references by name) are resolved from `~/.config/neurowire/meshes` or `NEUROWIRE_MESHES`. ## Output | Flag | Description | |------|-------------| | `-f, --format ` | Output format: `nwf`, `atom`, `rss`, `json`, `md`. Omit for the terminal view. | | `-o, --out ` | Write serialized output to a file instead of stdout. | ```bash neurowire https://example.com/feed.xml --format atom > feed.xml neurowire --mesh ai-news.json -f json -o ai-news.json ``` ::: tip Formats `atom` and `rss` both produce XML. `json` is JSON Feed 1.1. `md` is Markdown. `nwf` is the compact Neurowire format. See [NWF](/formats/nwf), [Atom](/formats/atom), [RSS](/formats/rss), [JSON Feed](/formats/json-feed), and [Markdown](/formats/markdown). ::: ## Shaping the output These run **before** `--format`, in this order: filters, then sort/order/limit and time windows. ### Filtering | Flag | Description | |------|-------------| | `--filter ` | Keep entries where the field matches. Repeatable. | | `--exclude ` | Drop entries where the field matches. Repeatable. | * Fields: `title`, `summary`, `source`, `author`, `tag`. * The pattern is a case-insensitive **substring** by default. * Wrap it in slashes for a case-insensitive **regex**: `/pattern/`. * Splitting is on the first colon only, so patterns may contain colons. ```bash neurowire --mesh ai-news.json --filter tag:release --exclude title:sponsored -f json neurowire --mesh ai-news.json --filter 'title:/\bv\d+\.\d+/' -f json ``` ### Sort, order, limit | Flag | Description | |------|-------------| | `--sort ` | Sort by `date`, `title`, or `source`. | | `--order ` | `asc` or `desc`. Default: newest-first for date, A-Z otherwise. | | `-n, --limit ` | Keep at most `n` entries (non-negative integer). | ```bash neurowire --mesh ai-news.json --sort date --order desc --limit 10 ``` ### Date windows | Flag | Description | |------|-------------| | `--since ` | Keep entries within this window, e.g. `24h`, `90m`, `7d`. | | `--max-age ` | Drop entries older than this (same window syntax). | | `--today` | Keep entries since midnight UTC today. | | `--this-week` | Keep entries since Monday midnight UTC. | | `--between ..` | Keep entries between two parseable dates, e.g. `2026-01-01..2026-02-01`. | ```bash neurowire --mesh ai-news.json --since 24h --sort date -f atom neurowire --mesh ai-news.json --between 2026-01-01..2026-02-01 -f json ``` ## Tail mode `neurowire tail` treats a source as a stream instead of a document: it polls forever and prints each new entry as it appears, one at a time. It takes the same sources and the same shaping flags as a normal run, applied per tick. | Flag | Description | |------|-------------| | `--interval ` | Poll interval, e.g. `30s`, `15m`, `6h`, `1d`. Default `5m`, floor `30s`. | | `--state ` | JSON file of seen entry keys, so restarts skip already-reported items. | | `-f nwf` | Stream raw [NWFJ](/formats/nwfj) journal lines instead of the terminal view. | | `--from ` | Render a remote [`GET /tail`](/guide/http-api#get-tail) SSE stream instead of polling locally. | | `--journal ` | Append everything seen to a journal, exactly as on a one-shot run. | | `--sink ` | Deliver each tick's new entries to a sink. | ```bash neurowire tail https://example.com/feed.xml neurowire tail --mesh ai-news.json --interval 60s --filter tag:release neurowire tail --mesh ai-news.json -f nwf | grep -i release neurowire tail --from https://api.example.com/tail?src=ai-news ``` Status lines (the interval, poll errors) go to stderr, so `-f nwf` pipes cleanly: the whole stream is one valid NWFJ document, header and checkpoints included, and `neurowire validate` style tooling can read it back. A tick whose fetch fails prints `[tail] error: ...` and waits for the next one instead of ending the tail. With another `-f` (`atom`, `rss`, `json`, `md`) each tick serializes just its new entries, which is the batch shape watch mode uses. `--from` expects the remote stream's default `format=json`. It reconnects on its own with exponential backoff, resuming from the last event id it saw, and reports each failed attempt on stderr rather than retrying silently. `--journal` and `--sink` still apply to what arrives; the shaping flags (`--filter`, `--exclude`, `--sort`, `--order`, `--limit`, `--since`, `--max-age`, `--between`) do not, because the server decides what the stream contains, so passing them prints a warning. ::: tip Tail is the concept page [Tail](/concepts/tail) covers what "new" means, why the interval has a floor, and how tail, watch, and a one-shot fetch differ. ::: ## Watch mode Watch is tail's batch-output sibling: the same poll loop, but each tick emits one feed of the new entries rather than a line per entry. | Flag | Description | |------|-------------| | `-w, --watch` | Enable the watch loop. Runs until the process is killed. | | `--interval ` | Poll interval, e.g. `30s`, `30m`, `6h`, `1d`. Default `5m`, floor `30s`. | | `--state ` | JSON file of seen entry keys, so restarts skip already-reported items. | Each tick re-applies your filters and refinements, prints only the new entries (in `--format` when set), and writes a `[watch] N new (M seen)` line to stderr. A failed fetch is reported as `[watch] error: ...` and retried on the next tick. ```bash neurowire --mesh ai-news.json --watch --interval 15m -f json neurowire --mesh ai-news.json --watch --state ~/.neurowire-seen.json ``` ::: tip Polling politeness Both loops clamp the interval to 30 seconds, add a little jitter so many pollers do not arrive together, and revalidate with conditional requests. Choose an interval that suits the source, not the floor. ::: ## Journals Keep an append-only archive of everything a source has published, so you can come back to it later. Entries the journal already holds are dropped, which makes journaling safe to run on a timer. | Flag | Description | |------|-------------| | `--journal ` | Append the fetched entries to the journal ``. | | `--journal-dir ` | Where journals live. Default `$NEUROWIRE_JOURNAL`, else `~/.config/neurowire/journal`. | ```bash neurowire --mesh ai-news.json --journal ai neurowire --mesh ai-news.json --journal ai --watch --interval 30m ``` Each run reports what it added: `Journaled 4 new entries to ai (head 128)`. Journals are stored as [NWFJ](/formats/nwfj) segments plus a rebuildable manifest, and are read back with the `journal` subcommands below. ## Sinks Push entries to a destination over HTTP POST. Repeatable. The destination kind is auto-detected from the URL host: Slack (`slack.com`), Discord (`discord.com`/`discordapp.com`), or a generic webhook (everything else, which receives the JSON Feed as `application/feed+json`). | Flag | Description | |------|-------------| | `--sink ` | POST entries to this destination. Repeatable. | Sinks are additive to stdout and never abort the run: a failing sink prints a one-line warning and continues. With `--watch`, only the fresh entries are delivered each tick. ```bash neurowire --mesh ai-news.json --watch --sink https://hooks.slack.com/services/... neurowire --mesh ai-news.json --sink https://discord.com/api/webhooks/... ``` ::: tip Slack and Discord message shape Slack and Discord receive a short text message: a header line then up to 10 bullet lines (title - link), with an overflow line when there are more. Discord content is capped at 2000 characters. ::: ## Taps Taps teach Neurowire to read sites with no RSS/Atom feed. The built-in taps load automatically. Add your own: | Flag | Description | |------|-------------| | `--taps ` | Load extra taps from a `.json` file or a directory. Repeatable. | | `--tap-pack ` | Register one or more themes from the optional [`@neurowire/taps-pack`](/reference/taps-pack) catalog (e.g. `gaming,space`, or `all`). Repeatable. | You can also set the `NEUROWIRE_TAPS` env var (a path or `:`-separated list), or drop `*.json` files into `~/.config/neurowire/taps/`. When custom taps load, the CLI writes `Loaded N custom tap(s)` to stderr. See [Taps](/concepts/taps). You do not have to write the selectors yourself. [`tap wizard`](#tap-wizard) authors a tap interactively, [`tap check`](#tap-check) tells you when one stops matching, and [`tap heal`](#tap-heal) repairs it after a redesign. `--tap-pack` needs `@neurowire/taps-pack` installed (`pnpm add @neurowire/taps-pack`); if it is absent the CLI prints an install hint and continues. Unknown theme keys are skipped with a warning. Example: `neurowire --tap-pack gaming https://www.pcgamer.com/rss/ -f json`. ## Global flags | Flag | Description | |------|-------------| | `-h, --help` | Show help. | | `-v, --version` | Show the version. | ## Subcommands ### tail Follow a feed, mesh, or construct as a live stream. See [Tail mode](#tail-mode) above for the flags. ```bash neurowire tail --mesh ai-news.json --interval 60s ``` ### validate Check that an NWF document is well-formed. Prints line-numbered warnings and errors to stderr; on success prints a summary, on failure exits non-zero. ```bash neurowire validate feed.nwf neurowire validate https://example.com/feed.nwf ``` ### tap wizard Author a tap step by step. The page is fetched once, then each of the seven steps shows ranked candidate selectors, a live `matched:` count, and a sample of what the current pick extracts. Type a number to accept a candidate, paste a selector to override it, press Enter to skip an optional field (or to take the top candidate on a required one). ```bash neurowire tap wizard https://example.com/blog neurowire tap wizard https://example.com/blog --yes -o ./example.json ``` | Flag | Description | |------|-------------| | `-o, --out ` | Where to write the tap. Default: `~/.config/neurowire/taps/.json`. | | `-y, --yes` | Accept every top candidate without prompting. | Nothing is written unless the template passes [verification](/concepts/taps#the-verification-gate); on a failure the failed checks print and the command exits 1. That holds for `--yes` too: it is a shortcut past the prompts, not past the gate. The written file carries a `url` hint naming the page it was authored against, so `tap check` and `tap heal` know where to look later. The template engine ignores that key. ### tap check Do taps still match the pages they were written for? Deterministic and network-cheap (one fetch per tap), so it belongs in CI: it exits 1 the day a redesign breaks a tap, instead of letting the feed quietly go empty. ```bash neurowire tap check ~/.config/neurowire/taps neurowire tap check ./example.json --json neurowire tap check --all ``` | Flag | Description | |------|-------------| | `--all` | Check every registered tap instead of a file or directory. | | `--json` | Print machine-readable results instead of the table. | | `--url ` | Force the page to check against, instead of each tap's own `url` hint. | Each tap comes back: | Status | Meaning | |--------|---------| | `healthy` | Every check passed. | | `degraded` | Passed the gate with a soft check failing, e.g. off-host links. | | `broken` | A hard check failed, or the page could not be fetched. | | `unknown` | The tap names no page, so there was nothing to check. | The exit code is 1 if any tap is broken. A tap's `host` is deliberately never turned into a page to fetch: most listing pages live at a path (`example.com/blog`), so fetching the apex would report a perfectly good tap as broken. Give a tap a `url` hint, or pass `--url`, and it gets checked; otherwise it is honestly reported as `unknown`. ### tap heal A site changed. Re-author the tap against the page as it stands today: fields whose selectors still match are kept verbatim, and only the broken ones are walked. `--yes` takes the top candidate for each broken field. ```bash neurowire tap heal ~/.config/neurowire/taps/example.com.json neurowire tap heal ./example.json --yes ``` The healed template goes through the same gate as the wizard, and the previous file is kept as `.bak`. Healing repairs the tap it was given rather than growing it: a field the tap never claimed stays unclaimed, and the tap's `host` and `feedTitle` survive. A tap under `node_modules` is code someone else ships, so its replacement is printed rather than written. ### tap doctor Inspect a feed-less page and propose a `FeedTemplate` (a tap). The proposed template prints as pretty JSON to stdout (redirect it into a taps file); a human preview of matched entries prints to stderr. Exits non-zero when nothing can be proposed. ```bash neurowire tap doctor https://example.com/blog > ~/.config/neurowire/taps/example.com.json ``` ::: tip Alias `neurowire doctor ` is accepted as a shorthand for `neurowire tap doctor `. ::: ### opml export Export a mesh or construct to OPML 2.0. Requires `--mesh ` or `--construct `. Writes to `--out` or stdout. ```bash neurowire opml export --mesh ai-news.json > ai-news.opml neurowire opml export --construct daily.json -o daily.opml ``` ### opml import Import an OPML file or URL into a mesh JSON. The mesh name comes from `--name`, else the OPML head title, else `imported`. Writes to `--out` or stdout. | Flag | Description | |------|-------------| | `-o, --out ` | Write the mesh JSON here instead of stdout. | | `--name ` | Set the imported mesh's name. | ```bash neurowire opml import subscriptions.opml -o mesh.json --name "My Reader" ``` ### journal head Print the journal's head cursor, the position to resume from next time. ```bash neurowire journal head ai ``` ### journal cat Print a journal as a feed, in any output format. With `--cursor ` it prints only the entries after that sequence number, which is how you pull a delta rather than the whole archive. A cursor older than the oldest retained entry prints a warning, since compaction means the delta cannot be complete. | Flag | Description | |------|-------------| | `--cursor ` | Only entries after this sequence number. Accepts `42` or `42.`. | | `-f, --format ` | Any output format. Omit for the terminal view. | | `-o, --out ` | Write to a file instead of stdout (with `--format`). | ```bash neurowire journal cat ai -f md neurowire journal cat ai --cursor 128 -f json ``` ### journal query Query an archive with the same `--filter`, `--exclude`, date-window, `--sort`, and `--limit` flags the fetch path uses, so a journal answers exactly what a live feed answers. Segments that provably cannot match are skipped without being read, and a note goes to stderr when that happens. ```bash neurowire journal query ai --filter tag:rust --since 30d -f md neurowire journal query ai --filter source:Anthropic --sort date --limit 20 ``` ::: tip Journals are plain text Segments are one record per line, TAB-separated, so `grep` and `awk` work directly on them. Use `journal cat ai -f json` when you want to load an archive into duckdb, sqlite, or pandas. See the [NWFJ format](/formats/nwfj). ::: ### sync Pull journal deltas from a peer running the [`nwf-sync/1`](/formats/nwf-sync) endpoints, instead of re-fetching every upstream source yourself. Entries land in the local journal store and are then read back with the ordinary `journal` commands. | Flag | Description | |------|-------------| | `--peers` | Sync every peer in `~/.config/neurowire/peers.json` instead of one URL. | | `--journal ` | Sync only this journal. Omit it to sync everything the peer publishes. | | `--token ` | Bearer token, when the peer requires one. | | `--journal-dir ` | Where the local journals live. | ```bash neurowire sync https://hub.example.com --journal ai --token secret neurowire sync --peers ``` Each line reports what moved, and the summary closes with the error count. Exits non-zero when any peer or journal failed, so a cron job notices. ``` https://hub.example.com ai: 26 new, cursor 1310, 2 requests, 4.1 KB 26 new entries from 1 peer (4.1 KB, 0 errors) ``` Cursors are recorded per peer and journal in `~/.config/neurowire/peers-state.json` (or `$NEUROWIRE_PEERS_STATE`), and only after the entries have been appended, so an interrupted sync costs one re-pull rather than a hole. Re-syncing is idempotent: the store drops entries it already holds. A line can also carry `bootstrapped from snapshot` (the cursor predated the peer's retention, so the pull restarted from what it still keeps) or `peer journal diverged, cursor reset` (the peer's journal was rebuilt or restored, so the cursor no longer meant anything and the pull started over). ### peers Manage `~/.config/neurowire/peers.json` (or `$NEUROWIRE_PEERS`). Adding a URL already on the list replaces its entry, which is how you rotate a token. The file is written `0600`, since it holds bearer tokens, and a URL with no scheme is stored as `https://`. If the file cannot be parsed, `add` and `remove` refuse rather than overwriting it. ```bash neurowire peers add https://hub.example.com --token secret neurowire peers add http://node-c.lan:8787 --journal ai neurowire peers list neurowire peers remove https://hub.example.com ``` ::: tip Set up the whole topology The [Federation guide](/guide/federation) builds a three-node hub, laptop, and offline-relay setup from scratch with these commands. ::: ## More examples ```bash neurowire https://example.com/blog neurowire --construct daily.json neurowire --construct daily.json --format atom --limit 20 neurowire --mesh ai-news.json --filter tag:release --exclude title:sponsored -f json neurowire tail --mesh ai-news.json --interval 60s --sink https://hooks.slack.com/services/... neurowire tap wizard https://example.com/blog --yes neurowire tap check ~/.config/neurowire/taps --json neurowire sync --peers ``` --- --- url: /guide/library.md --- # Library usage Neurowire is a library first. The CLI and API are thin wrappers over the same functions. This page shows the common entry points; see the [reference](/reference/core) for full symbol lists. The split is deliberate: `@neurowire/ingest` does network work (fetch, detect, parse, merge), `@neurowire/core` is pure (the model, serializers, filtering, merging), `@neurowire/web` renders HTML. ## Fetch a single feed `fetchFeed(url, options)` fetches a website, RSS, or Atom URL and normalizes it to a `NeurowireFeed`. ```ts import { fetchFeed } from '@neurowire/ingest' const feed = await fetchFeed('https://blog.rust-lang.org/feed.xml') console.log(feed.title, feed.entries.length) ``` For an HTML page, `fetchFeed` follows a declared feed link, then a registry tap, then heuristic auto-detection. Options include `template` (force a CSS-selector template), `timeoutMs`, `retries`, `backoffMs`, `maxDepth`, `signal`, and `cache`. ## Fetch a mesh `fetchMesh(mesh, options)` fetches every source in parallel and merges them into one newest-first feed (tagged by source, deduped). Sources that fail are skipped and never fatal unless all of them fail. ```ts import { fetchMesh } from '@neurowire/ingest' import type { Mesh } from '@neurowire/core' const mesh: Mesh = { name: 'AI News', sources: [ { name: 'Claude Blog', url: 'https://claude.com/blog' }, { name: 'Simon Willison', url: 'https://simonwillison.net/atom/everything/' }, ], } const feed = await fetchMesh(mesh, { limit: 20 }) ``` Pass `onSourceError` to override the default stderr warning for partial failures. ## Fetch a construct A construct is a bundle of meshes. `fetchConstruct(construct, options)` keeps the per-mesh grouping in a `FetchedConstruct` (one merged feed per mesh). `flattenConstruct` collapses it into a single feed for the serializers. ```ts import { fetchConstruct, flattenConstruct, createConfigMeshResolver } from '@neurowire/ingest' import type { Construct } from '@neurowire/core' const construct: Construct = { name: 'Daily', meshes: [ 'ai-news', // a { ref } shorthand, resolved by the resolver below { name: 'Releases', sources: [{ name: 'Claude Code', url: 'https://github.com/anthropics/claude-code/releases.atom' }] }, ], } const fetched = await fetchConstruct(construct, { resolver: createConfigMeshResolver() }) const feed = flattenConstruct(fetched) // one feed, entries tagged by mesh ``` `{ ref }` members need a `MeshResolver`. `createConfigMeshResolver()` reads `~/.config/neurowire/meshes`; you can supply any `(ref: string) => Mesh | undefined` function. Meshes are fetched with bounded `concurrency` (default 2). ## Parse an already-fetched document When you already have the bytes (no network), `ingestDocument(doc, options)` turns a `RawDocument` into a feed. Useful for tests or custom transports. ```ts import { ingestDocument } from '@neurowire/ingest' const feed = await ingestDocument({ url: 'https://example.com/blog', contentType: 'text/html', body: '...', }) ``` ## Register taps To resolve feed-less sites (e.g. `claude.com`), register the curated taps once at startup. ```ts import { registerAllTaps } from '@neurowire/taps' registerAllTaps() // built-ins plus ~/.config/neurowire/taps and NEUROWIRE_TAPS ``` See [Taps](/concepts/taps) for adding your own. ## Serialize a feed `serialize(feed, format)` from core renders a `NeurowireFeed` to a string. Formats: `nwf`, `atom`, `rss`, `json`, `md`. ```ts import { serialize, FORMATS, MEDIA_TYPES } from '@neurowire/core' const atom = serialize(feed, 'atom') const json = serialize(feed, 'json') console.log(FORMATS) // ['atom', 'rss', 'json', 'md', 'nwf'] console.log(MEDIA_TYPES.json) // 'application/feed+json; charset=utf-8' ``` The direct serializers (`toAtom`, `toRss`, `toJsonFeed`, `toMarkdown`, `toNwf`) and the NWF round-trip helpers (`fromNwf`, `validateNwf`) are also exported. See [The model](/concepts/model) and the [core reference](/reference/core). ## Archive to a journal A [journal](/concepts/journals) is an append-only archive of what a source has published. `openJournalStore` persists one per id as [NWFJ](/formats/nwfj) segments; entries the journal already holds are dropped on append, so archiving on a timer is safe. ```ts import { openJournalStore } from '@neurowire/ingest' const store = openJournalStore() // ~/.config/neurowire/journal, or { dir } const { added, head } = store.append('ai', feed.entries, { id: feed.id, title: feed.title, home: feed.home, }) console.log(`added ${added}, head is now ${head.seq}`) ``` Read it back from a cursor to get only what arrived since last time: ```ts const { entries, tooOld } = store.since('ai', head) ``` `tooOld` is true when compaction dropped the entries the cursor pointed past, so the delta cannot be complete and you should re-read the whole journal. Query an archive with the same filter, window, sort, and limit semantics a live feed uses: ```ts const { entries, scanned, skipped } = store.query('ai', { filter: { include: [{ field: 'tag', pattern: 'rust' }] }, from: Date.now() - 30 * 86_400_000, sort: 'date', limit: 20, }) ``` `skipped` lists segment files the plan ruled out from the manifest alone and never opened. To hand the result to a serializer, wrap it back into a feed: ```ts import { journalToFeed, serialize } from '@neurowire/core' const records = store.read('ai') const md = serialize(journalToFeed({ records, head: store.head('ai'), issues: [] }), 'md') ``` The format itself (`createJournalEncoder`, `parseJournal`, `readJournalSince`, `journalHead`, `verifyJournal`, `queryJournal`) is pure and lives in core, so you can journal to something other than the filesystem. See the [core](/reference/core#journal) and [ingest](/reference/ingest#journal-store) references. ## Render HTML `@neurowire/web` is the only package that emits HTML (core stays format-pure). All three functions return self-contained pages with inline CSS and no external requests. ```ts import { toHtml, toConstructHtml, toConstructPages } from '@neurowire/web' // a single feed or mesh -> one page const page = toHtml(feed) // a fetched construct -> an overview page of recap cards const overview = toConstructHtml(fetched) // a fetched construct -> an index.html plus one page per mesh const pages = toConstructPages(fetched) // [{ filename, html }, ...] ``` `toConstructHtml` accepts `{ meshHref }` to control where each recap card links. See the [web reference](/reference/web). --- --- url: /guide/http-api.md --- # HTTP API `@neurowire/api` is a small [Hono](https://hono.dev) service that exposes Neurowire over HTTP: convert a feed, serve a named mesh or construct, or build one from a posted body. It serves the feed formats only (NWF, Atom, RSS, JSON Feed, Markdown). HTML is not a feed format and lives in `@neurowire/web`. ## Running it The package ships a `neurowire-api` binary that starts the server. ::: code-group ```bash [pnpm] pnpm add @neurowire/api pnpm exec neurowire-api ``` ```bash [npm] npm install @neurowire/api npx neurowire-api ``` ::: It listens on `http://localhost:8787` by default and prints the bound URL on start. The `app` (a Hono instance) is also exported, so you can mount it in your own server or test it with `app.fetch`. ::: warning No built-in auth or rate limiting The API ships with no authentication and no rate limiting. If you expose it publicly, put it behind a proxy, gateway, or auth layer of your own. That is the operator's responsibility. The one exception is the peer sync surface: `/sync/*` takes an optional bearer token, and publishes nothing at all unless you name the journals. That is access control for a surface that hands out whole archives, not a general auth layer for the service. ::: ## Configuration All configuration is via environment variables: | Variable | Default | Purpose | |----------|---------|---------| | `PORT` | `8787` | Port the server listens on. | | `NEUROWIRE_CACHE_TTL` | `300` | Response cache TTL in seconds (matches the `Cache-Control: max-age=300` header). | | `NEUROWIRE_MESHES` | (unset) | Extra mesh directories (`:` or `,` separated), searched before `~/.config/neurowire/meshes`. | | `NEUROWIRE_CONSTRUCTS` | (unset) | Extra construct directories, searched before `~/.config/neurowire/constructs`. | | `NEUROWIRE_TAPS` | (unset) | Extra taps (path or `:`-separated list); built-ins always load. | | `NEUROWIRE_JOURNAL` | `~/.config/neurowire/journal` | Journal store the `/sync/*` routes read and `GET /tail` replays from. | | `NEUROWIRE_TAIL_HEARTBEAT_MS` | `25000` | How often `GET /tail` writes a keep-alive comment. | | `NEUROWIRE_SYNC_PUBLISH` | (unset) | Journal ids to publish over `/sync/*` (`:` or `,` separated), or `*` for all. Nothing is published without it. | | `NEUROWIRE_SYNC_TOKEN` | (unset) | Bearer token required on every `/sync/*` request. Unset means those routes are open. | | `NEUROWIRE_SYNC_CONFIG` | `~/.config/neurowire/sync.json` | File holding `{ "publish": [...], "token": "..." }`, if you prefer it to the env. | ## Caching Each `GET` route caches its serialized response in an in-memory TTL cache keyed by source and format. Upstream fetches use a separate conditional cache so a TTL miss can revalidate (ETag / Last-Modified) instead of refetching the whole document. Every response carries `Cache-Control: public, max-age=300`. ## Endpoints ### GET / Service metadata: name, version, supported formats, the endpoint map, and the names of all available meshes and constructs. ```bash curl http://localhost:8787/ ``` ### GET /healthz Liveness check. ```bash curl http://localhost:8787/healthz # {"status":"ok","service":"neurowire","version":"0.5.0"} ``` ### GET /feed Convert any feed or website URL. Query params: * `url` (required): the source URL. * `format` (optional, default `atom`): one of `nwf`, `atom`, `rss`, `json`, `md`. ```bash curl "http://localhost:8787/feed?url=https%3A%2F%2Fblog.rust-lang.org%2Ffeed.xml&format=json" ``` A missing `url` returns `400`; an unknown `format` returns `400`; an upstream failure returns `502` with `{ error, detail }`. ### GET /mesh Serve a **named** mesh, resolved from your mesh directories then the bundled `ai-news`. Query params: * `src` (required): the mesh name. * `format` (optional, default `atom`). ```bash curl "http://localhost:8787/mesh?src=ai-news&format=atom" ``` A missing `src` returns `400` (with the list of known meshes); an unknown mesh returns `404`. ### POST /mesh Build a mesh from a JSON body (no named lookup). The body is validated against the mesh schema. `format` is a query param (default `atom`). ```bash curl -X POST "http://localhost:8787/mesh?format=json" \ -H 'content-type: application/json' \ -d '{ "name": "AI News", "sources": [ { "name": "Claude Blog", "url": "https://claude.com/blog" } ] }' ``` An invalid body returns `400` with `{ error, detail }`. A `headers` key on a posted source is dropped: per-source headers come only from the server's own mesh files, never from a request body, so a caller cannot make the server send credentials. ### GET /construct Serve a **named** construct, resolved from your construct directories then the bundled `daily`. The construct is fetched and flattened into one feed. Query params: * `src` (required): the construct name. * `format` (optional, default `atom`). ```bash curl "http://localhost:8787/construct?src=daily&format=json" ``` A missing `src` returns `400` (with the list of known constructs); an unknown construct returns `404`. `format=html` is rejected like any unknown format, because the API serves feed formats only. ### POST /construct Build a construct from a JSON body. Inline meshes and `{ ref }` members are accepted; refs are resolved against the named meshes (same lookup as `GET /mesh`). The result is flattened. `format` is a query param (default `atom`). ```bash curl -X POST "http://localhost:8787/construct?format=atom" \ -H 'content-type: application/json' \ -d '{ "name": "Daily", "meshes": [ "ai-news", { "name": "Releases", "sources": [ { "name": "Claude Code", "url": "https://github.com/anthropics/claude-code/releases.atom" } ] } ] }' ``` ### GET /tail Follow a feed, mesh, or construct as a live [server-sent events](https://developer.mozilla.org/docs/Web/API/Server-sent_events) stream. The server polls the target on an interval and pushes each new entry as it appears. Query params: * One target, same as the routes above: `url=`, `src=`, or `construct=`. * `format` (optional, default `json`): `json` sends one entry object per event, `nwf` sends the [NWFJ](/formats/nwfj) journal lines for that entry. * `interval` (optional, default `300`): seconds, or a duration like `15m`. Clamped up to a floor of 60 seconds. * `journal` (optional): the id of a journal on the server, which turns on replay (see below). * `since` (optional): a journal cursor to replay from, the same value `Last-Event-ID` carries. ```bash curl -N "http://localhost:8787/tail?src=ai-news&interval=120" ``` With `format=nwf` the events carry journal lines rather than JSON, so concatenating their `data` payloads gives a valid NWFJ document. Against a journaled target the `E` line's sequence number and the SSE event id are the same number. Events: | Event | Payload | |-------|---------| | `init` | JSON: the target, the format, the effective interval, whether resume is `journal` or `live`, the journal head, and how many entries were replayed. | | `entry` | One entry: a JSON object, or its NWFJ lines when `format=nwf`. The event `id` is the entry's cursor: a journal cursor when the target is journaled, otherwise a counter local to the stream. | | (comment) | `: ping` every 25 seconds, so buffering proxies keep the connection open. | A missing target is a `400` and an unknown mesh or construct a `404`, both plain JSON: errors are decided before the stream opens, never mid-stream. An unknown `format` is a `400` too; `/tail` serves `json` and `nwf` only, not the feed formats the other routes take. Responses also carry `X-Accel-Buffering: no` for nginx. The [Tail concept page](/concepts/tail) covers the polling semantics behind all of this: what counts as new, the interval floors, and how a failed tick is handled. **One poll loop per target.** Every client following the same target shares a single upstream poll, so fifty browsers on `ai-news` cost one fetch per tick. The loop starts with the first subscriber and stops when the last one disconnects. Sharing is keyed by the target, the effective interval, and the journal, so a client that asks for a different cadence or a different journal gets its own loop rather than quietly riding someone else's. **Resume.** Pass `journal=` naming a journal that already exists on the server (created by `neurowire --journal `, see [Journals](/concepts/journals)). The route then appends what it sees to that journal, so event ids are real cursors, and a client that reconnects with `Last-Event-ID` (or `?since=`) is replayed everything after that cursor before going live. Without a journal the tail is live-only, event ids are stream-local, and `init` reports `"resume": "live"`. The route never creates a journal of its own; an unknown id simply falls back to live-only. ::: tip Rate expectations The interval floor is 60 seconds server-side, and each tick is jittered slightly so many tails on one host do not arrive together. Conditional requests mean an unchanged source usually costs a `304`, but pick an interval that suits the source rather than the floor. ::: ### GET /sync/\* Peer delta exchange over [`nwf-sync/1`](/formats/nwf-sync): `/sync/journals`, `/sync/head`, `/sync/since`, and `/sync/snapshot`, so another node can pull journal deltas instead of re-fetching every upstream source itself. These routes serve [journals](/concepts/journals), not feeds, so the `format` query does not apply to them. They are inert until you set `NEUROWIRE_SYNC_PUBLISH`. See [Federation](/guide/federation) to set up a node, [Sync](/concepts/sync) for the trust model, and the [API reference](/reference/api#sync-endpoints) for every status code. ```bash curl "http://localhost:8787/sync/head?journal=ai" # {"journal":"ai","head":1284,"hash":"9f1c0f0b8ad0f0e3"} ``` ## Bundled defaults The API ships one bundled mesh (`ai-news`) and one bundled construct (`daily`), so `?src=ai-news` and `?src=daily` work with no setup. Add your own by dropping JSON files into the mesh/construct directories (see [Meshes](/concepts/meshes) and [Constructs](/concepts/constructs)). Names must be simple identifiers; anything path-like is rejected to avoid directory traversal. --- --- url: /guide/federation.md --- # Federation A feed is a snapshot of a front page. A [journal](/concepts/journals) is the archive behind it. **Federation** is what happens when one node keeps that archive and other nodes pull deltas from it instead of re-fetching every source themselves. The wire protocol is [`nwf-sync/1`](/formats/nwf-sync): four read-only HTTP `GET`s, NWFJ segments as the payload. This guide builds the three-node topology from scratch using only documented commands. ## Why bother Three devices following the same 200 sources make 600 requests a tick against those publishers, from three IPs that together look like a small scraping operation. Route them through one node and it is 200, with the devices pulling compact deltas. That is the bandwidth argument. The one that actually matters is time. Open a laptop at 18:00 and a live fetch truthfully reports what those sites show *now*, which for a busy outlet may be the last three hours. Everything published while the lid was shut has scrolled off the page and no amount of fetching brings it back. A node that stayed awake recorded all of it, and the laptop asks for exactly the window it missed. If you run one machine following twenty feeds, just fetch them. Sync earns its keep at device count, at intermittent connectivity, or when history matters. ## The topology Peers are configured, not discovered. `A -> C -> D` is fine: D gets what C already pulled from A. Merging is idempotent by entry key, so a node peering with both A and C stores one copy, and a cycle terminates. `entry.source` travels inside the record, so provenance survives every hop. ## Node A: the hub A fetches the open web and journals what it sees. Nothing here is new: it is the ordinary fetch path with `--journal`. ```bash # One-shot, on a timer (cron, systemd, launchd): neurowire --construct daily.json --journal ai # Or keep it fresh in one long-running process: neurowire --construct daily.json --journal ai --watch --interval 30m ``` Then publish that journal. Nothing is published by default, so this is an explicit act: ```bash export NEUROWIRE_JOURNAL=/var/lib/neurowire/journal export NEUROWIRE_SYNC_PUBLISH=ai export NEUROWIRE_SYNC_TOKEN=$(openssl rand -hex 32) neurowire-api ``` Or the same thing as a file, `~/.config/neurowire/sync.json`: ```json { "publish": ["ai"], "token": "a-long-random-string" } ``` Check it from the machine itself: ```bash curl -s localhost:8787/sync/journals -H "Authorization: Bearer $NEUROWIRE_SYNC_TOKEN" # {"version":1,"journals":[{"id":"ai","title":"Daily","head":1284,...}]} ``` ::: warning The token is access control, not authentication of content The bearer token decides who may read. The hash chain in the payload detects corruption. Neither proves who wrote an entry. You sync from peers you chose to trust, over TLS your own proxy terminates. See the [trust model](/formats/nwf-sync#integrity). ::: Put A behind a reverse proxy with a real certificate. Neurowire serves plain HTTP and expects the operator to own transport security, the same [self-host stance](/guide/http-api) the rest of the API takes. ## Node B: the laptop B never fetches the open web. It records A as a peer and pulls: ```bash neurowire peers add https://hub.example.com --token a-long-random-string neurowire peers list # https://hub.example.com (token, all journals) neurowire sync --peers # https://hub.example.com # ai: 1284 new, cursor 1284, 4 requests, 812.0 KB # 1284 new entries from 1 peer (812.0 KB, 0 errors) ``` The first sync moves the archive. Every one after it moves a delta: ```bash neurowire sync --peers # https://hub.example.com # ai: 26 new, cursor 1310, 2 requests, 4.1 KB ``` And when nothing has changed, one request that returns a number: ```bash neurowire sync --peers # https://hub.example.com # ai: 0 new, cursor 1310, 1 request, 41 B ``` Peers live in `~/.config/neurowire/peers.json`; cursors live beside them in `peers-state.json`, keyed by peer and journal. A cursor is only ever written after the entries have landed, so a crash mid-sync costs one re-pull rather than a hole in the archive. A one-shot pull, without registering anything: ```bash neurowire sync https://hub.example.com --journal ai --token a-long-random-string ``` Put it on a timer, or run it when the lid opens: ```bash */15 * * * * neurowire sync --peers >> ~/.local/state/neurowire-sync.log 2>&1 ``` ### A synced journal is an ordinary journal This is the payoff of the journal being the shared substrate. Nothing about reading a pulled archive is special: ```bash neurowire journal head ai neurowire journal cat ai -f json neurowire journal query ai --filter tag:release --since 7d -f md neurowire-web --mesh ai.json --out page.html ``` ## Node C: a relay C pulls from A and republishes what it now holds. It is node B plus the two lines that make it a server: ```bash neurowire peers add https://hub.example.com --token a-long-random-string neurowire sync --peers export NEUROWIRE_SYNC_PUBLISH=ai neurowire-api ``` C's journal is not a byte copy of A's. It holds the same entries, with C's own sequence numbers and its own chain, because C encoded them itself. That is why a cursor is only meaningful against the peer it came from, and why nothing compares a local head to a remote one. ## Node D: no internet at all D peers with C over the LAN. Nothing in its configuration says the outside world exists: ```bash neurowire peers add http://node-c.lan:8787 neurowire sync --peers ``` D gets everything A collected, one hop late. Add more hops and it keeps working: the entry-key dedupe means a diamond stores one copy, and a cycle simply stops adding. ## Operating it **Retention.** A journal grows. `compact` drops the oldest segments: ```ts import { openJournalStore } from '@neurowire/ingest' openJournalStore().compact('ai', 12) // keep the newest 12 segments ``` A peer whose cursor pointed into a dropped segment gets `410 Gone` on its next pull and re-bootstraps from the snapshot automatically, reporting `bootstrapped from snapshot`. That is loud on purpose: retention is a decision, not something an archive should discover as a quiet hole. **Verifying.** Every pulled response is chain-verified before a single entry is appended. A flipped byte aborts that response with a named error and merges nothing from it. **Rotating the token.** Change `NEUROWIRE_SYNC_TOKEN` on A, restart, and update each peer with `neurowire peers add --token ` (adding an existing URL replaces its entry). **Debugging a pull.** The endpoints are plain `GET`s, so `curl` is a first-class client: ```bash curl -si "https://hub.example.com/sync/head?journal=ai" -H 'Authorization: Bearer ...' curl -s "https://hub.example.com/sync/since?journal=ai&cursor=1284" -H 'Authorization: Bearer ...' ``` ## What this is not No push, no gossip, no discovery, no DHT. No signatures. No conflict resolution, because append-only facts have nothing to conflict over. The [non-goals](/formats/nwf-sync#not-in-v1) are a fence, not an oversight: pull federation gets to prove itself first. ## See also * [Sync](/concepts/sync), the concept behind this guide: the three cases it solves, and what the hash chain does and does not prove. * [`nwf-sync/1`](/formats/nwf-sync), the protocol. * [NWFJ](/formats/nwfj), the payload format. * [Journals](/concepts/journals), the concept. * [CLI](/guide/cli#sync), the `sync` and `peers` commands. * [`@neurowire/api`](/reference/api#sync-endpoints), the server reference. --- --- url: /guide/agents.md --- # Agents (MCP) Neurowire ships an MCP server, `@neurowire/mcp`, so an LLM agent can read feeds, follow journals, and verify taps directly. Results default to NWF, the most compact format, which keeps an agent's context small. ## Claude Code plugin The plugin bundles the server and two skills (following feeds, authoring a tap), so one install gives an agent both the tools and the know-how: ```bash /plugin marketplace add starside-io/claude-plugins /plugin install neurowire@starside ``` ## Any MCP client The server speaks stdio. Point your client at: ```bash npx -y @neurowire/mcp ``` For Claude Code without the plugin: ```bash claude mcp add neurowire -- npx -y @neurowire/mcp ``` For a client configured with JSON (Claude Desktop, Cursor, and similar): ```json { "mcpServers": { "neurowire": { "command": "npx", "args": ["-y", "@neurowire/mcp"], "env": { "NEUROWIRE_MCP_ALLOW": "github.com,simonwillison.net,claude.com" } } } } ``` ## Configuration | Variable | Default | Purpose | |----------|---------|---------| | `NEUROWIRE_MCP_ALLOW` | (unset) | Comma separated hosts that caller-supplied URLs may fetch. Unset allows every host. | | `NEUROWIRE_MESHES` | (unset) | Extra mesh directories, searched before `~/.config/neurowire/meshes`. | | `NEUROWIRE_CONSTRUCTS` | (unset) | Extra construct directories, searched before `~/.config/neurowire/constructs`. | | `NEUROWIRE_JOURNAL` | `~/.config/neurowire/journal` | Where the journal tools read from. | | `NEUROWIRE_TAPS` | (unset) | Extra tap files or directories, as for the CLI. | ## What an agent can do * **Answer "what shipped this week?"** with `fetch_mesh` on the bundled `ai-news` mesh, or `query` with `since: "7d"`. * **Follow an archive over time.** Keep a journal fresh with the CLI (`neurowire --mesh ai-news.json --journal ai --watch`), then have the agent call `whats_new` and store the cursor it returns. The next call yields exactly what arrived. * **Read a site with no feed.** `propose_tap` drafts a tap and `verify_tap` gates it. The agent can iterate on selectors, but only a passing template counts, and nothing is installed. See [Taps](/concepts/taps). * **Answer questions about Neurowire itself** with `search_docs`, which reads the published [`llms.txt`](https://neurowire.starside.io/llms.txt) index. The full tool list, input fields, and result conventions are in the [`@neurowire/mcp` reference](/reference/mcp). ## Docs for agents The docs site publishes a flattened, token-efficient copy of itself for any agent, MCP or not: * [`/llms.txt`](https://neurowire.starside.io/llms.txt): the table of contents. * [`/llms-full.txt`](https://neurowire.starside.io/llms-full.txt): every page in one file. --- --- url: /guide/recipes.md --- # Recipes Practical end-to-end workflows. Each one is a short sequence of commands you can copy. They assume the `neurowire` CLI (and, where noted, `neurowire-web`) are installed. See [Installation](/guide/installation). Every recipe below is the same pipeline with different pieces switched on. Shaping always runs before serializing, so a filter narrows what a format, a sink, and a journal all receive: Source Shape Out ## Watch a site and push new posts to Slack Long-poll a mesh and deliver only the entries you have not seen yet to a Slack incoming webhook. The `--state` file makes restarts skip already-reported items. ```bash neurowire --mesh ai-news.json \ --watch --interval 15m \ --state ~/.neurowire-seen.json \ --sink https://hooks.slack.com/services/T000/B000/XXXX ``` Swap the sink URL for a Discord webhook (`https://discord.com/api/webhooks/...`) or any generic endpoint (which receives the JSON Feed as `application/feed+json`). The sink kind is auto-detected from the URL. See [CLI sinks](/guide/cli#sinks). ## Build a daily HTML news page from a construct A [construct](/concepts/constructs) bundles several meshes. Render it to a self-contained page (all CSS inline, no external requests). Single combined page (every entry tagged by its mesh): ```bash neurowire-web --construct daily.json --combined --out public/index.html ``` Multi-page "repo of feeds" (an overview plus one page per mesh) into a directory: ```bash neurowire-web --construct daily.json --out public/ ``` Limit to recent items with `--since` or `--today`: ```bash neurowire-web --construct daily.json --combined --since 24h --out public/index.html ``` ## Migrate subscriptions via OPML Import an OPML export from another reader into a Neurowire mesh, then export it back out if you need to. Import (the mesh name comes from `--name`, else the OPML title): ```bash neurowire opml import subscriptions.opml -o my-reader.json --name "My Reader" ``` Use the resulting mesh like any other: ```bash neurowire --mesh my-reader.json --format json --limit 20 ``` Export a mesh (or construct) back to OPML 2.0: ```bash neurowire opml export --mesh my-reader.json > my-reader.opml ``` See [opml subcommands](/guide/cli#opml-export). ## Add a tap for a feed-less site A [tap](/concepts/taps) teaches Neurowire to read a site with no RSS/Atom feed. Let `tap doctor` propose one, save it, then use it. ```bash # propose a template and save it where Neurowire looks for taps neurowire tap doctor https://example.com/blog \ > ~/.config/neurowire/taps/example.com.json # now the site resolves like any feed neurowire https://example.com/blog --format atom ``` You can also load a tap ad hoc with `--taps ` or via the `NEUROWIRE_TAPS` env var. See [tap doctor](/guide/cli#tap-doctor). For a page the one-shot proposal gets wrong, and for keeping the tap alive afterwards, see [Author a tap and keep it working](#author-a-tap-and-keep-it-working) below. ## Archive a mesh and research it later Front pages scroll away. Keep everything a mesh publishes in an append-only [journal](/concepts/journals), then query the archive months later. Archive on a timer (cron, a systemd timer, a CI schedule). Re-running adds only what is new, so the interval does not have to be precise: ```bash neurowire --mesh ai-news.json --journal ai ``` Or let a single long-running watch loop do both jobs, archiving every tick while it notifies: ```bash neurowire --mesh ai-news.json \ --watch --interval 30m \ --state ~/.neurowire-seen.json \ --journal ai \ --sink https://hooks.slack.com/services/T000/B000/XXXX ``` Then research it with the same flags a live fetch takes: ```bash # everything tagged rust in the last 30 days, as Markdown neurowire journal query ai --filter tag:rust --since 30d -f md # one source, newest first neurowire journal query ai --filter source:Anthropic --sort date --limit 20 ``` Pull just the delta since a position you recorded earlier: ```bash cursor=$(neurowire journal head ai) # ...later... neurowire journal cat ai --cursor "$cursor" -f json ``` To take the archive somewhere else, dump it as JSON and load it into duckdb, sqlite, or pandas: ```bash neurowire journal cat ai -f json > ai-archive.json ``` See [CLI journals](/guide/cli#journals). ## Pipe a live tail into other tools `neurowire tail -f nwf` streams [NWFJ](/formats/nwfj) records to stdout as entries arrive: one line per record, TAB-separated, nothing else on the stream. Status output goes to stderr, so the pipe stays clean. Watch for something specific as it lands. Records are plain lines, so `grep` works, and `--line-buffered` keeps it from sitting on a buffer while it waits for more: ```bash neurowire tail --mesh ai-news.json -f nwf \ | grep --line-buffered -i 'release' ``` ::: warning A grepped stream is not a document Filtering by line drops the header and the dictionary lines the entry records point at, which makes the result unparseable as NWFJ. Grep it to look at it; keep the whole stream when you want to read it back. ::: To keep a file you can query later, write the whole stream and watch a copy: ```bash neurowire tail --mesh ai-news.json -f nwf \ | tee -a ~/feeds/ai-news.nwfj \ | grep --line-buffered '^E' ``` A long-running tail can also archive into a proper [journal](/concepts/journals) as it goes, which is the better option when you want cursors, segments, and `journal query` rather than one growing file: ```bash neurowire tail --mesh ai-news.json --interval 5m --journal ai ``` Follow a [self-hosted API](/guide/http-api#get-tail) instead of polling the sources yourself, so one machine does the fetching for all your terminals: ```bash neurowire tail --from 'http://localhost:8787/tail?src=ai-news' ``` See [CLI tail mode](/guide/cli#tail-mode) and the [Tail concept](/concepts/tail). ## Share one archive across two machines A laptop that is closed most of the day cannot fetch what it missed: those entries scrolled off the front page hours ago. Let an always-on machine do the fetching and journaling, then pull the deltas when the laptop wakes. ### On the hub (a VPS, a home server, anything always on) Journal the mesh on a timer. Re-running adds only what is new, so the interval is not critical: ```bash # crontab: every 30 minutes */30 * * * * neurowire --mesh ai-news.json --journal ai ``` Then publish that journal over [`nwf-sync/1`](/formats/nwf-sync). Nothing is published unless you name it, and the token is optional but wanted on anything reachable from the internet: ```bash export NEUROWIRE_SYNC_PUBLISH=ai export NEUROWIRE_SYNC_TOKEN=$(openssl rand -hex 32) neurowire-api ``` Put it behind a reverse proxy with a real certificate. Neurowire serves plain HTTP and leaves transport security to you. ### On the laptop Register the hub once: ```bash neurowire peers add https://hub.example.com --token ``` Then sync on wake, or on a short timer. Being up to date costs one request that returns a number, so checking often is cheap: ```bash neurowire sync --peers # ai: 26 new, cursor 1310, 2 requests, 4.1 KB ``` The result is an ordinary journal, so everything in [Archive a mesh and research it later](#archive-a-mesh-and-research-it-later) applies unchanged: ```bash neurowire journal query ai --filter tag:rust --since 7d -f md ``` The laptop never fetches the sources, so it never has to keep taps current, and it sees exactly the corpus the hub saw rather than a different snapshot per machine. ::: tip Add a third machine and nothing changes A node that pulled from the hub can publish `/sync` itself, and a third machine can peer with *that*, even if it never touches the internet. Merges are deduplicated by entry key, so a diamond stores one copy and a cycle terminates. See [Federation](/guide/federation) for the three-node walkthrough and [Sync](/concepts/sync) for the trust model: the hash chain detects corruption, it does not prove authorship, so peer with nodes you trust. ::: ## Convert any feed to RSS 2.0 Normalize any source (RSS, Atom, JSON Feed, or an HTML page) and re-emit it as RSS 2.0. ```bash neurowire https://simonwillison.net/atom/everything/ --format rss > willison.rss ``` The same `--format rss` works on a mesh or a (flattened) construct: ```bash neurowire --mesh ai-news.json --format rss --limit 25 > ai-news.rss ``` See [RSS format](/formats/rss). ## Author a tap and keep it working A [tap](/concepts/taps) is the only part of Neurowire that depends on how someone else's HTML is laid out, so it is the only part that rots. This is the full lifecycle: author it once, let CI tell you the day it breaks, repair it in a minute. ### 1. Author it `tap wizard` fetches the page once and walks the seven tap fields, showing ranked candidate selectors and a sample of what the current pick extracts: ```bash neurowire tap wizard https://example.com/blog ``` Type a number to accept a candidate, paste a selector to override it, press Enter to skip an optional field. For a site the heuristics already handle, skip the prompts entirely: ```bash neurowire tap wizard https://example.com/blog --yes ``` Either way the tap is written only if it passes the [verification gate](/concepts/taps#the-verification-gate), so `--yes` cannot leave you with a tap that quietly scrapes the navigation bar. By default it lands in `~/.config/neurowire/taps/.json`, where Neurowire picks it up automatically: ```bash neurowire https://example.com/blog --format atom ``` ### 2. Check it in CI `tap check` re-runs the same gate against the live page. It is one fetch per tap, no model, and no API key, so it is cheap enough to run on a schedule: ```bash neurowire tap check ~/.config/neurowire/taps ``` ``` healthy example.com (24 items) degraded linkblog.example (31 items) link-host: 4/31 on linkblog.example 1 tap(s): 1 healthy, 1 degraded, 0 broken, 0 unknown ``` The exit code is 0 while every tap is `healthy` or `degraded`, and 1 as soon as one is `broken`, which is all a CI step needs: ```yaml # .github/workflows/taps.yml on: schedule: [{ cron: '0 6 * * *' }] jobs: taps: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-node@v4 with: { node-version: 24 } - run: npm install -g @neurowire/cli - run: neurowire tap check ./taps ``` Add `--json` when something downstream should read the result rather than a human: ```bash neurowire tap check ./taps --json | jq '.taps[] | select(.health == "broken") | .host' ``` ::: tip Give each tap a page to check `check` never guesses a listing page from a tap's `host`, because fetching `https://example.com/` to check a tap written for `https://example.com/blog` would report a healthy tap as broken. Taps written by the wizard carry a `url` hint for this. A tap without one is reported `unknown` and skipped; add the key by hand, or pass `--url`. ::: ### 3. Heal it after a redesign When the check goes red, `tap heal` re-authors the tap against the page as it stands today. Selectors that still match are kept, and only the broken fields are walked: ```bash neurowire tap heal ~/.config/neurowire/taps/example.com.json ``` ``` broken Article block (was article.post-card) 1) article.tile-block > 1 keep Title h2 keep Date time verify 24 items ยท items 24 extracted, 3 needed ยท titles 24/24 non-empty ยท ... wrote ~/.config/neurowire/taps/example.com.json (previous kept as ...json.bak) ``` `--yes` takes the top candidate for each broken field, which is usually right when only the class names changed. The healed tap goes through the same gate, and the file it replaces is kept as `.bak`. Re-run `tap check` to confirm, and CI goes green again. See [tap wizard](/guide/cli#tap-wizard), [tap check](/guide/cli#tap-check), and [tap heal](/guide/cli#tap-heal). --- --- url: /concepts/model.md --- # The canonical model Neurowire has exactly one in-memory representation of a feed: the `NeurowireFeed`. Every parser (RSS, Atom, RDF, JSON Feed, HTML auto-detect, taps) produces this shape, and every serializer (NWF, atom, rss, json, md) consumes it. Nothing else in the system reads the upstream format directly. This single model is the contract that lets any input format turn into any output format. The model lives in [`packages/core/src/model.ts`](https://github.com/neurowire/neurowire) and is defined with [zod](https://zod.dev) schemas, so the same definition both types the code and validates unknown input at runtime. ## Design: list metadata only, no article bodies Neurowire deliberately models a feed as a **list of articles**, not the articles themselves. An entry carries title, link, dates, a short summary, authors, and tags. It does not carry the full HTML article body. This keeps the model small, the outputs compact, and fetching cheap (one listing page or feed, not N article pages). If you want the full text, follow the entry's `link`. ::: tip This is why the `nwf` format can be so terse and why a mesh of dozens of sources stays small. The model never holds article content. ::: ## `NeurowireFeed` | Field | Type | Required | Notes | |-------|------|----------|-------| | `id` | `string` | yes | Stable identifier for the feed (usually the source URL). | | `title` | `string` | yes | Human-readable feed title. | | `home` | `string` | no | The website the feed represents (the "alternate" link). | | `self` | `string` | no | The canonical URL of the feed document itself. | | `updated` | `string` | yes | Last-updated timestamp (ISO 8601). | | `authors` | `Person[]` | no | Feed-level authors. | | `generator` | `{ name, version? }` | no | What produced the feed. Neurowire stamps its own here. | | `entries` | `NeurowireEntry[]` | yes | The articles, newest first by convention. | ## `NeurowireEntry` | Field | Type | Required | Notes | |-------|------|----------|-------| | `id` | `string` | yes | Stable per-entry identifier (see synthetic ids below). | | `title` | `string` | yes | Article title. | | `link` | `string` | yes | URL of the article. | | `published` | `string` | no | First-published timestamp (ISO 8601). | | `updated` | `string` | no | Last-updated timestamp (ISO 8601). | | `summary` | `string` | no | Short description or excerpt. Not the full body. | | `authors` | `Person[]` | no | Per-entry authors. | | `tags` | `string[]` | no | Categories or labels. | | `source` | `{ name?, url? }` | no | Where this entry came from. Used by meshes to tag entries by source. | ## `Person` | Field | Type | Required | |-------|------|----------| | `name` | `string` | yes | | `url` | `string` | no | | `email` | `string` | no | ## The zod schemas The schemas mirror the tables above. `EntrySchema` and `FeedSchema` are exported alongside their inferred types, and `parseNeurowireFeed(data)` validates an unknown value (throwing a `ZodError` on bad input). ```ts export const EntrySchema = z.object({ id: z.string(), title: z.string(), link: z.string(), published: z.string().optional(), updated: z.string().optional(), summary: z.string().optional(), authors: z.array(PersonSchema).optional(), tags: z.array(z.string()).optional(), source: z .object({ name: z.string().optional(), url: z.string().optional() }) .optional(), }) export const FeedSchema = z.object({ id: z.string(), title: z.string(), home: z.string().optional(), self: z.string().optional(), updated: z.string(), authors: z.array(PersonSchema).optional(), generator: z.object({ name: z.string(), version: z.string().optional() }).optional(), entries: z.array(EntrySchema), }) ``` ## Stable synthetic entry ids (content hashing) Many sources, especially scraped HTML listings, give an entry no real GUID. To keep deduplication and round-trips stable across every format, Neurowire derives a deterministic id from the entry's content when none is present. The logic lives in `packages/core/src/id.ts`: * `hashHex(input)` is a pure FNV-1a (64-bit) hash returned as a fixed 16-char lowercase hex string. It uses no `node:crypto`, so core stays portable. * `stableId(link, title)` hashes `\`${link}\n${title}\``and returns`urn:nwf:<16-char hex>\`. During ingestion, an entry that already has a real id keeps it. An entry with an empty id gets `stableId(link, title)` stamped in. The same `(link, title)` always produces the same urn, so the same article keeps the same id across fetches and across output formats. ``` urn:nwf:9f1c4b2a6d3e0f87 ``` ::: tip Because the id is derived from the link and title, an article that is re-fetched (or merged into a mesh) is recognized as the same entry and deduplicated, even when the upstream source never gave it an id. ::: --- --- url: /concepts/output-formats.md --- # Output formats Once a source is parsed into the canonical [model](/concepts/model), Neurowire can serialize it to any of its output formats. The format system is small and centralized in [`packages/core/src/serialize/index.ts`](https://github.com/neurowire/neurowire): one list of formats, one media-type map, one extension map, and one dispatch function. ## `serialize(feed, format)` The single entry point. It takes a `NeurowireFeed` and a `Format` and returns the serialized string. ```ts import { serialize } from '@neurowire/core' const atom = serialize(feed, 'atom') const json = serialize(feed, 'json') const md = serialize(feed, 'md') ``` Internally it is a switch over the format that delegates to the per-format serializer (`toAtom`, `toRss`, `toJsonFeed`, `toMarkdown`, `toNwf`), each of which is also exported directly if you want to call it without the dispatch. ## The formats | Format | Page | Media type | Extension | |--------|------|------------|-----------| | `nwf` | [/formats/nwf](/formats/nwf) | `text/x-neurowire; charset=utf-8` | `nwf` | | `atom` | [/formats/atom](/formats/atom) | `application/atom+xml; charset=utf-8` | `xml` | | `rss` | [/formats/rss](/formats/rss) | `application/rss+xml; charset=utf-8` | `xml` | | `json` | [/formats/json-feed](/formats/json-feed) | `application/feed+json; charset=utf-8` | `json` | | `md` | [/formats/markdown](/formats/markdown) | `text/markdown; charset=utf-8` | `md` | The `json` format is [JSON Feed 1.1](https://www.jsonfeed.org/version/1.1/). The `nwf` format is Neurowire's own compact line-oriented format (see [/formats/nwf](/formats/nwf)). ## The registries Three exported constants keep formats consistent across the CLI, API, and web packages: * `FORMATS`: the readonly tuple `['atom', 'rss', 'json', 'md', 'nwf']`, and the `Format` type derived from it. * `MEDIA_TYPES`: the `Content-Type` string for each format (used by the API). * `EXTENSIONS`: the file extension for each format (used by the CLI and page generator). ```ts import { FORMATS, MEDIA_TYPES, EXTENSIONS, isFormat } from '@neurowire/core' ``` ## `isFormat(value)` A type guard that narrows an arbitrary string to `Format`. Use it to validate user input (a `--format` flag, a query parameter) before calling `serialize`. ```ts if (isFormat(input)) { return serialize(feed, input) // input is now typed as Format } ``` ## What is deliberately not here ::: warning HTML is not a core format HTML is intentionally **not** a feed serializer and is **not** in `FORMATS`. Rendering a feed or mesh into a self-contained HTML news page lives in [`@neurowire/web`](/formats/html) (`toHtml`, `toConstructHtml`). Keeping HTML out of core keeps `@neurowire/core` pure and dependency-light (zod only), with no DOM or presentation concerns. ::: ::: tip OPML is a subscription list, not a feed OPML describes a **list of subscriptions** (which feeds to follow), not the contents of a feed. It is therefore not one of the feed serializers dispatched by `serialize`. See its dedicated page for import and export. ::: ::: tip NWFJ is storage, not rendering A [journal](/concepts/journals) stores a feed's history rather than rendering its current state, so `nwfj` is not in `FORMATS` either. It has its own media type (`application/x-nwf-journal`) and extension (`nwfj`), exported as `JOURNAL_MEDIA_TYPE` and `JOURNAL_EXTENSION`. Reading a journal back yields an ordinary `NeurowireFeed` (`journalToFeed`), which every serializer above already handles. See [the NWFJ format](/formats/nwfj). ::: --- --- url: /concepts/fetching.md --- # Fetching Every source Neurowire reads, a feed, a website, a mesh member, goes through one HTTP client: `fetchDocument`, in [`packages/ingest/src/fetch.ts`](https://github.com/neurowire/neurowire). It is built for fetching feeds from untrusted URLs on a schedule, so it handles timeouts, retries, conditional requests, and redirect-based SSRF in one place. ```ts import { fetchDocument, createMemoryCache } from '@neurowire/ingest' const doc = await fetchDocument('https://example.com/feed.xml', { timeoutMs: 15000, retries: 2, cache: createMemoryCache(), }) ``` `fetchDocument` returns a `RawDocument`: the final `url` (after redirects), the `contentType`, the `body`, optional `etag` / `lastModified` validators, and a `notModified` flag set when the body came from cache via a 304. ## Document detection Once a document is fetched, `detectKind(contentType, body)` classifies it before any parser runs: the `Content-Type` first, then a body sniff when the type is unhelpful (many hosts serve feeds as `text/plain` or `application/octet-stream`). | Kind | Matched by | |------|-----------| | `nwf` | `text/x-neurowire`, or an `NWF1` first line | | `atom` | `application/atom+xml`, or a ` void \| Promise` | none | Per-hop guard (throw to block). See SSRF below. | | `headers` | `Record` | none | Extra request headers, e.g. `authorization` for a private feed. See caller headers below. | | `delay` | `(ms, signal) => Promise` | setTimeout-based | Injectable sleep (tests drive it with fake timers). | ## Retry policy Each attempt runs the full manual redirect loop. On failure the request is retried up to `retries` times. The classification is deliberate: **Retried** (a fresh attempt might recover): * Network-level failures (DNS, connection reset), which `fetch` surfaces as a `TypeError`. * Internal timeout aborts (the per-attempt deadline). * Upstream `5xx` responses. * `429 Too Many Requests`. **Not retried** (a retry cannot help): * `4xx` other than `429`. * An invalid or unsupported-protocol URL. * An SSRF rejection from the `validate` guard. * Too many redirects (more than 5 hops). * A caller-driven abort via `signal` (rejects immediately, no further attempts). ### Honoring `Retry-After` A `429` carries its `Retry-After` header, parsed (delta-seconds or an HTTP date) into milliseconds. When present, the retry loop waits that long instead of the computed backoff, capped at 30 seconds so a hostile or huge `Retry-After` cannot stall the process forever. ### Jittered backoff When there is no `Retry-After`, the wait is full-jitter exponential backoff: `base * 2 ** attempt`, multiplied by a random factor in `[0.5, 1.0)`, capped at 30 seconds. Jitter spreads out retries so many sources failing at once do not all retry in lockstep. ## Conditional cache (304 handling) Pass a `ConditionalCache` and `fetchDocument` will make conditional requests. After a successful `200`, it stores the body together with the response's `ETag` and `Last-Modified`. On the next fetch of that URL it sends: * `If-None-Match` from the stored `ETag`. * `If-Modified-Since` from the stored `Last-Modified`. If the server replies `304 Not Modified`, the cached body is returned with `notModified: true`. A `304` is a success and is never retried. The cache is written only on a final `200`. Conditional headers apply only to the originally requested URL (the first hop), not to redirect targets, and they are re-sent on every retried attempt. `createMemoryCache()` returns a simple `Map`-backed `ConditionalCache`. The cache is always injected by the caller; the library keeps no global state, so you control its lifetime and scope. ```ts const cache = createMemoryCache() await fetchDocument(url, { cache }) // 200, body cached await fetchDocument(url, { cache }) // may return { notModified: true } ``` ## Redirects and the per-hop SSRF guard Redirects are followed **manually**, one hop at a time, up to 5 hops (`Location` is resolved against the current URL). This exists so the optional `validate` guard runs against **every** target, not just the URL the caller passed in. This closes the classic SSRF bypass where a public URL `302`-redirects to an internal address (`http://169.254.169.254/`, `http://localhost/`, a private RFC 1918 range). `validate` is called for the initial URL and for each redirect target before the request is made. Throwing from `validate` blocks that hop, and the rejection is **not** retried. ::: warning If you fetch user-supplied URLs, supply a `validate` that rejects private, loopback, and link-local addresses. Because redirects are re-validated per hop, a public URL cannot redirect its way into your internal network. ::: Only `http:` and `https:` are allowed; any other protocol throws an unsupported-protocol error. ## Request headers Every request sends a stable identifying `User-Agent`: ``` Neurowire/0.1 (+https://github.com/neurowire/neurowire) ``` and an `Accept` header that prefers feed formats, then JSON, then HTML: ``` application/atom+xml, application/rss+xml, application/feed+json, application/json;q=0.9, text/html;q=0.8, */*;q=0.5 ``` ## Caller headers and redirects `FetchOptions.headers` adds request headers, which is how a private feed gets its `Authorization`. Header names are lowercased. A caller header overrides the default `User-Agent` and `Accept`, but never the conditional `If-None-Match` / `If-Modified-Since` headers: those belong to the cache, and a caller value for them is ignored. Because redirects are followed manually, each hop builds its headers afresh, and credential headers get the same treatment browsers give them: `authorization`, `proxy-authorization`, and `cookie` are sent only while the hop's origin (scheme, host, and port) matches the origin of the URL you asked for. A `302` to another host, a scheme change, or a port change drops them; every other caller header still goes along. Without this, a token meant for `api.example.com` would ride along to whatever host the redirect names. `fetchFeed` applies the same rule to a discovered feed link: a page on one origin that advertises its feed on another gets that feed fetched without credentials. Headers are not part of the conditional-cache key (the key is the URL), so do not share one `ConditionalCache` between callers with different credentials for the same URL. ## Cancellation The caller's `signal` is composed with the internal per-attempt timeout signal, so either can cancel an in-flight request. A caller abort is distinguished from a timeout: the timeout is retryable, the caller abort is not. --- --- url: /concepts/taps.md --- # Taps Many sites publish a blog but ship no RSS or Atom feed. A **tap** wiretaps such a site: it is a per-host recipe of CSS selectors (a `FeedTemplate`) that turns the site's listing page into a Neurowire [feed](/concepts/model). The template engine lives in [`packages/ingest/src/html/template.ts`](https://github.com/neurowire/neurowire); the curated taps and their loaders live in [`@neurowire/taps`](https://github.com/neurowire/neurowire); the authoring and healing tools live in [`@neurowire/tap-wizard`](/reference/tap-wizard). ## The life of a tap A tap is the one part of Neurowire shaped by someone else's HTML, so it is the one part that rots. The concept therefore covers four things, not one: | Stage | What it answers | Where | |-------|-----------------|-------| | **Shape** | What is a tap, and how does the engine read one? | [A tap is a `FeedTemplate`](#a-tap-is-a-feedtemplate) | | **Author** | Which selectors describe this page? | [the wizard](#authoring-a-tap-with-the-wizard), [`tap wizard`](/guide/cli#tap-wizard) | | **Verify** | Is this set of selectors actually good? | [the verification gate](#the-verification-gate) | | **Maintain** | Does it still match, and how do I fix it? | [checking and healing](#checking-and-healing-taps), [`tap check`](/guide/cli#tap-check) / [`tap heal`](/guide/cli#tap-heal) | None of it involves a model, at any stage. Structure proposes, a human confirms, and a deterministic verifier decides. Agent-driven authoring, when it arrives, will drive these same primitives as tools rather than replace them. ## A tap is a `FeedTemplate` A tap is a set of CSS selectors describing where each article and its fields live on the listing page. Selectors for fields other than `item` and `title` are looked up **within** each matched item. | Field | Required | What it selects | |-------|----------|-----------------| | `host` | no | Hostname this tap applies to, e.g. `blog.example.com`. Used to match by host in the registry. | | `feedTitle` | no | Overrides the feed title (otherwise the page `` is used). | | `item` | yes | Each article row. The other field selectors run inside each match. | | `title` | yes | The title text within an item. | | `link` | no | The link (its `href` is read). **Omit it when the matched `item` element is itself the `<a>`.** | | `date` | no | The date. Reads `[datetime]` first, then the element's text. | | `summary` | no | The summary text. | | `author` | no | The author name. | | `tags` | no | Tag elements (each one's text becomes a tag). | ### How `applyTemplate` extracts entries For each element matching `item`, the engine reads the `title` text and resolves the link. When `link` is omitted, the item element itself is treated as the anchor (its `href`, or the first `<a>` inside it, is used). An item with no title or no resolvable link is skipped. Dates are normalized, relative links are resolved against the source URL, and entries get [stable synthetic ids](/concepts/model#stable-synthetic-entry-ids-content-hashing) if the source gives none. ```json { "host": "blog.example.com", "feedTitle": "Example Blog", "item": "article.post", "title": "h2", "link": "a.permalink", "date": "time", "summary": "p.excerpt" } ``` ## Resolution order When Neurowire ingests a page (`ingestDocument`), it tries to produce a feed in this order, taking the first that yields entries: 1. **Explicit template.** A `FeedTemplate` passed directly by the caller always wins. 2. **Discovered feed link.** A declared `<link rel="alternate">` RSS/Atom/JSON feed on the page is followed (the highest-fidelity result). 3. **Registry tap (by host).** A curated per-host tap from the registry. This beats heuristic auto-detect. 4. **Heuristic auto-detect.** On-page extraction (JSON-LD, then semantic HTML). If none of these extracts a feed, ingestion throws. ::: tip A real RSS/Atom feed always wins over a tap. Taps only matter for sites that have no feed at all. ::: ## Adding your own taps Users register custom taps from three sources, applied in order (later sources win on a host collision): 1. The drop-in directory `~/.config/neurowire/taps/` (or `$XDG_CONFIG_HOME/neurowire/taps`). A missing default directory is silently ignored. 2. The `NEUROWIRE_TAPS` environment variable: a path, or a `:` / `,` separated list of paths. 3. The CLI `--taps <path>` flag. A path may be a single JSON file or a directory of `*.json` files (each loaded in sorted order). Each file holds one tap object or an array of them, and every tap is validated against the schema. An explicitly requested path that is missing or invalid throws (unlike the optional default directory). ```bash neurowire https://example.com/blog --taps ~/my-taps/example.json NEUROWIRE_TAPS=~/taps-a:~/taps-b neurowire https://example.com/blog ``` ## Authoring a tap with `tap doctor` You do not have to write selectors by hand. `proposeTemplate(html, url)` inspects a feed-less page and **proposes** a `FeedTemplate`: it finds repeated item-like containers (sibling `article` / `li` / class-patterned `div`s that each hold a heading and an `<a href>`, or a grid of bare `<a>` cards), picks the most consistent selector, and derives `title` / `link` / `date` selectors relative to it. The candidate is validated by actually running `applyTemplate`, so a proposal is returned only when it extracts at least one entry. The CLI exposes this as `tap doctor <url>`. It prints the proposed template to stdout and a match count plus sample titles to stderr, so you can pipe the template straight to a file: ```bash neurowire tap doctor https://example.com/blog > ~/.config/neurowire/taps/example.com.json ``` The proposal includes `template`, `matched` (entry count), and `sampleTitles` (up to 5), so you can eyeball the result before saving. ## Authoring a tap with the wizard `tap doctor` is one shot: it proposes, you take it or you do not. The **wizard** is the same heuristics turned into a loop, so a page the one-shot proposal gets wrong is still authorable without DOM-spelunking. ``` $ neurowire tap wizard https://example.com/blog step 1/7 Article block (required) Pick the block that repeats once for each article on the page. 1) article.post-card 2) li.entry 3) div.card > 1 step 2/7 Title (required) The headline text inside each block. 1) h2 2) h2.post-title > 1 preview Rust 1.94 released https://example.com/posts/rust-194 Announcing the new API https://example.com/posts/new-api step 3/7 Link (optional, Enter to skip) > verify 24 items ยท items 24 extracted, 3 needed ยท titles 24/24 non-empty ยท ... wrote ~/.config/neurowire/taps/example.com.json ``` Type a number to accept a candidate, paste a selector to override it, press Enter to skip an optional field. **One fetch per session:** every pick is re-applied against the document already in memory, which is what keeps the loop instant and the publisher unbothered. See [`tap wizard`](/guide/cli#tap-wizard) for the flags, and [`createTapSession`](/reference/tap-wizard#the-session) to drive the same walkthrough from your own code. Nothing about this involves a model. Candidates come from structure, ranked by how a page is actually built: | Field | Where the candidates come from | |-------|--------------------------------| | `item` | `article`/`li`/`div`/`section` elements containing a link, keyed as `tag.first-class` and ranked by how often that selector repeats (3 to 300 times). `is-` / `has-` / `js-` state classes are skipped, since those flip at runtime, as are classes CSS cannot address unescaped (`md:flex`, `w-1/2`). | | `title` | Headings and `[class*=title\|headline]` inside the first matched item. | | `link` | `a[href]` inside the item. | | `date` | `time`, `[datetime]`, `[class*=date\|time\|published]`. | | `summary` | `p`, `[class*=summary\|excerpt\|dek]`. | | `author` | `[class*=author\|byline]`, `[rel=author]`. | | `tags` | `[class*=tag\|category\|label]`, `[rel=tag]`. | Whatever `proposeTemplate` (the `tap doctor` heuristic) returns is seeded as candidate zero for each field, so the existing answer is always the default and the alternatives sit right behind it. The full ranking rules are in [`suggestCandidates`](/reference/tap-wizard#candidate-suggestion). What each pick extracts is previewed by running the real engine, not a lookalike: [`previewTemplate`](/reference/tap-wizard#preview) calls `ingestDocument` with the candidate template, so what the walkthrough shows you is exactly what a fetch will produce. ### The verification gate No tap is written without passing [`verifyTemplate`](/reference/tap-wizard#the-verification-gate), on the interactive path and on `--yes` alike. The same gate decides the verdicts `tap check` reports, so authoring and monitoring cannot disagree about what a healthy tap is. The checks are deterministic and give the same verdict every run: | Check | Hard | What it catches | |-------|------|-----------------| | `items` | yes | Fewer than 3 extracted entries: the selector did not really find the list. | | `titles` | yes | Fewer than 80% of matched items carrying title text. A rate, not a clean sweep, since the engine simply skips a title-less item and one promo card sharing the article class is not a broken tap. | | `links` | yes | A link that did not resolve to an absolute `http(s)` URL. | | `link-host` | no | Links pointing off-host. Common enough on link blogs to be a warning, not a failure. | | `unique-links` | yes | Every item yielding the same href, i.e. a selector that grabbed one shared anchor per item. | | `dates` | no | A claimed `date` selector that rarely parses. | | `ancestor` | yes | More than half the matched items sitting inside `nav`, `footer`, or `aside`, or items scattered instead of sharing a parent. This is what stops a tap from "working" by scraping the navigation bar, while a lone "related posts" card in a sidebar is tolerated. | A template passes when no hard check failed and the share of passing checks clears the threshold (0.75 by default). A tap that clears the gate with a soft check failing is reported as **degraded** rather than healthy. ## Checking and healing taps Taps rot: sites redesign, classes get renamed, and a broken tap fails silently by producing an empty feed. `tap check` is the smoke alarm. ```bash neurowire tap check ~/.config/neurowire/taps # a file or a directory neurowire tap check --all --json # every registered tap ``` Each tap is fetched once and run through the same gate, then classified: | Verdict | Meaning | Exits 1 | |---------|---------|---------| | `healthy` | Every check passed. | no | | `degraded` | Passed the gate with a soft check failing, e.g. off-host links. | no | | `broken` | A hard check failed, or the page could not be fetched. | **yes** | | `unknown` | The tap names no page, so there was nothing to check. | no | That exit code is the point: it is what makes `tap check` worth putting on a CI schedule, and what stops a redesign from silently emptying a feed. See [the CI recipe](/guide/recipes#author-a-tap-and-keep-it-working) for a workflow file. A tap file may carry an optional `url` key naming the listing page to check. It is not part of the template schema (the engine ignores it), it just tells `check` and `heal` where to look. The wizard writes the hint for you. A tap with no hint is reported `unknown` rather than checked. The `host` is not turned into a page: most listing pages live at a path, so fetching `https://example.com/` to check a tap written for `https://example.com/blog` would report a healthy tap as broken. A check that lies is worse than a check that abstains. ```json { "host": "example.com", "item": "article.post-card", "title": "h2", "url": "https://example.com/blog" } ``` When a tap does break, [`tap heal <path>`](/guide/cli#tap-heal) re-authors it against the live page: each old selector is re-applied first, the ones that still match are kept, and only the broken fields are walked. A class rename, which is what most redesigns amount to, usually needs one answer. Healing repairs the tap it was given rather than growing it. A field the tap never claimed stays unclaimed, and the tap's `host` and `feedTitle` survive, since those are the author's decisions and not the site's. The replacement goes through the same gate as the wizard, and the file it replaces is kept as `<path>.bak` (written once, so a second heal cannot bury the hand-written original). Bundled taps in `@neurowire/taps` and `@neurowire/taps-pack` are code rather than user files, so a tap under `node_modules` has its replacement printed instead of written. Nothing in this loop calls a model or needs an API key. Heuristics propose, a human confirms, and a deterministic verifier decides. ## Bundled taps `@neurowire/taps` ships four curated taps, registered with `registerTaps()` / `registerAllTaps()`: | Tap | Host | Feed title | |-----|------|-----------| | `claudeBlog` | `claude.com` | Claude Blog | | `cursorBlog` | `cursor.com` | Cursor Blog | | `deepmindBlog` | `deepmind.google` | Google DeepMind Blog | | `mistralNews` | `mistral.ai` | Mistral AI News | `cursorBlog` omits `link` because each post card is itself an `<a href="/blog/...">`, so the matched item element is the anchor. ## Where to next * [tap wizard / check / heal](/guide/cli#tap-wizard): every flag and exit code. * [Author a tap and keep it working](/guide/recipes#author-a-tap-and-keep-it-working): the whole lifecycle as a runnable recipe, including a CI workflow. * [`@neurowire/tap-wizard`](/reference/tap-wizard): the library behind the commands, for driving the same walkthrough from your own code. * [`@neurowire/taps`](/reference/taps): the bundled taps and the loaders that register them. * [`@neurowire/taps-pack`](/reference/taps-pack): the optional themed catalog of ready-made sources. --- --- url: /concepts/meshes.md --- # Meshes A **mesh** is a named bundle of sources that fetch in parallel and merge into one feed. Point a mesh at several blogs, releases pages, and RSS feeds, and get back a single newest-first feed in any [output format](/concepts/output-formats). The `Mesh` type lives in [`@neurowire/core`](https://github.com/neurowire/neurowire) (`packages/core/src/model.ts`); fetching and merging live in `fetchMesh` (`packages/ingest/src/mesh.ts`). ## Shape A mesh is a `name` plus a list of `sources`, each a display `name` and a `url` (a feed URL or a website Neurowire can ingest), with optional per-source request `headers`: ```ts export const MeshSourceSchema = z.object({ name: z.string(), url: z.string(), headers: z.record(z.string(), z.string()).optional(), }) export const MeshSchema = z.object({ name: z.string(), sources: z.array(MeshSourceSchema) }) ``` ```json { "name": "AI News", "sources": [ { "name": "Claude Code Releases", "url": "https://github.com/anthropics/claude-code/releases.atom" }, { "name": "Claude Blog", "url": "https://claude.com/blog" }, { "name": "Simon Willison", "url": "https://simonwillison.net/atom/everything/" } ] } ``` `parseMesh(data)` validates an unknown value into a `Mesh`. ## How fetching and merging work `fetchMesh(mesh, options)` fetches every source **in parallel** (`Promise.allSettled`) and then merges the results into one feed via `mergeFeeds`: * Each entry is **tagged by source** (the source name carries through), so a rendered page or feed can show where each item came from. * Entries are **deduplicated** (the [stable entry ids](/concepts/model#stable-synthetic-entry-ids-content-hashing) make the same article from two sources collapse into one). * The merged feed is sorted **newest-first**. * An optional `limit` keeps only the newest N merged entries. ## Partial-failure semantics A mesh is resilient by design. A source that fails to fetch is **logged and skipped**, never fatal: * Each failure goes through `onSourceError(source, error)`, which by default writes a one-line warning to stderr with the source name and a short reason. * The mesh still merges and returns whatever sources succeeded. * It throws **only when every source fails**: `Mesh "<name>": no sources could be fetched`. You can pass a custom `onSourceError` handler to silence or redirect those warnings. ```ts import { fetchMesh } from '@neurowire/ingest' const feed = await fetchMesh(mesh, { limit: 50, onSourceError: (source, err) => log.warn(`skip ${source.name}: ${err}`), }) ``` `FetchMeshOptions` also forwards the shared fetch tuning (`signal`, `cache`, `timeoutMs`, `retries`, `backoffMs`) to every source. A single `cache` is shared across all sources in the mesh. See [Fetching](/concepts/fetching) for what those do. ## Private sources: per-source headers A source can carry request `headers`, sent only when fetching that source. That is how a mesh reads a private feed: a GitHub token for a private repository's releases feed, a bearer token for an internal API, a basic-auth header for a password-protected RSS endpoint. ```json { "name": "Internal", "sources": [ { "name": "Private releases", "url": "https://api.github.com/repos/acme/private/releases", "headers": { "authorization": "Bearer ${GITHUB_TOKEN}" } }, { "name": "Public blog", "url": "https://acme.example/blog" } ] } ``` Three rules keep this safe: * **Mesh files never hold a literal token.** Header values may reference `${ENV_VAR}`; every loader that reads mesh JSON from local config (the CLI, the API's named meshes, the MCP catalog) resolves those references at load time and throws if a referenced variable is unset or empty, naming the mesh, source, and header. A missing secret fails at startup, not as a confusing `401` later. * **Headers stay on their origin.** Credential headers (`authorization`, `proxy-authorization`, `cookie`) are sent to the source URL and its same-origin redirects only. A redirect or a discovered feed link on another origin is fetched without them, so a token meant for one host never reaches a host the page chose. See [Fetching](/concepts/fetching#caller-headers-and-redirects). * **Only local config can set headers.** Meshes supplied by a remote caller (`POST /mesh`, `POST /construct`, and the MCP `fetch_mesh` / `fetch_construct` inline inputs) are parsed with `PublicMeshSchema`, which drops `headers`. A caller cannot make a server send credentials, and the `${ENV_VAR}` substitution never runs on their input, so they cannot read the server's environment through a URL they control. Headers are not part of the conditional-cache key and are never written to the partial-failure warning on stderr. ## Named meshes from config Meshes can be stored as JSON files and referred to by name. The config directories are searched in order: 1. Directories in the `NEUROWIRE_MESHES` env var (`:` or `,` separated). 2. `~/.config/neurowire/meshes/` (or `$XDG_CONFIG_HOME/neurowire/meshes`). Files are matched as `<name>.mesh.json` then `<name>.json`. Names must be simple identifiers; anything path-like is rejected to prevent directory traversal. ## Bundled `ai-news` So that `?src=ai-news` works with no setup, the API ships one built-in mesh, `ai-news` ("AI News"), with three sources: Claude Code Releases, the Claude Blog, and Simon Willison's feed (see `packages/api/src/meshes.ts`). User mesh directories take precedence over the bundled defaults. ::: tip A mesh groups sources into one flat feed. To group several **meshes** into a repo of feeds (with the per-mesh grouping preserved), use a [construct](/concepts/constructs). ::: --- --- url: /concepts/constructs.md --- # Constructs A **construct** is a named bundle of [meshes](/concepts/meshes): a "repo" of feeds grouped into sections. Where a mesh merges sources into one flat feed, a construct keeps the per-mesh grouping, so you can render a multi-section page (a "Daily Brief" with a Models section, a Releases section, and so on) or flatten the whole thing into one feed for the standard serializers. The `Construct` type lives in [`@neurowire/core`](https://github.com/neurowire/neurowire) (`packages/core/src/model.ts`); fetching lives in `fetchConstruct` / `flattenConstruct` (`packages/ingest/src/construct.ts`). ## Shape A construct is a `name` plus a list of `meshes`. Each member is either an **inline mesh** (self-contained, with its own sources) or a **reference** to a mesh resolved elsewhere by name. ```ts export const ConstructMemberSchema = z.union([ z.string().transform((ref) => ({ ref })), // bare string is shorthand for { ref } ConstructRefSchema, // { ref: "..." } MeshSchema, // an inline mesh ]) export const ConstructSchema = z.object({ name: z.string(), meshes: z.array(ConstructMemberSchema), }) ``` A bare string is shorthand for `{ ref: string }`, so a published list of mesh names stays terse: ```json { "name": "Daily Brief", "meshes": ["ai-news", "security"] } ``` Or mix inline meshes with references: ```json { "name": "Daily Brief", "meshes": [ { "name": "Models", "sources": [ { "name": "Claude Blog", "url": "https://claude.com/blog" }, { "name": "Simon Willison", "url": "https://simonwillison.net/atom/everything/" } ] }, { "ref": "security" } ] } ``` `parseConstruct(data)` validates an unknown value into a `Construct`. ## Resolving references: `MeshResolver` A construct only carries references; the **lookup** is supplied by the caller, so core and ingest stay free of any filesystem or registry assumptions. ```ts export type MeshResolver = (ref: string) => Mesh | undefined ``` `fetchConstruct` takes an optional `resolver`. It is required only when the construct actually has `{ ref }` members: an inline mesh passes through unchanged, but a reference with no resolver (or one that resolves to nothing) throws. `createConfigMeshResolver()` (in `packages/ingest/src/mesh-config.ts`) returns a resolver backed by the mesh config directories: explicit dirs, then `NEUROWIRE_MESHES`, then `~/.config/neurowire/meshes/`. Drop a published mesh pack there and a construct can reference it by name. ```ts import { fetchConstruct, createConfigMeshResolver } from '@neurowire/ingest' const fetched = await fetchConstruct(construct, { resolver: createConfigMeshResolver(), }) ``` ## `fetchConstruct` vs `flattenConstruct` These are two views of the same data. **`fetchConstruct(construct, options)`** returns a `FetchedConstruct`: the construct name plus one merged feed **per mesh**, with grouping preserved. ```ts interface FetchedConstruct { name: string parts: { mesh: Mesh; feed: NeurowireFeed }[] } ``` This is what the grouped HTML rendering uses (a per-mesh section per part). **`flattenConstruct(fetched, options)`** collapses a `FetchedConstruct` into one `NeurowireFeed`, tagging every entry with the mesh it came from. The grouping is dropped (NWF, Atom, and JSON Feed cannot express it), so this is the path the feed serializers and the API use. ```ts const fetched = await fetchConstruct(construct, { resolver }) const feed = flattenConstruct(fetched, { limit: 50 }) const atom = serialize(feed, 'atom') ``` ## Concurrency control A construct of many meshes could otherwise open every source of every mesh at once (dozens of connections), and that burst makes slow hosts time out and drop whole meshes. So meshes are fetched with **bounded concurrency** (`concurrency`, default `2`). Sources **within** a single mesh are still fetched together in parallel. ## Partial-failure semantics Failures are non-fatal at both levels: * A source that fails is logged via `onSourceError` and skipped (same as a plain mesh). * A mesh whose every source fails is logged via `onMeshError` (default: a one-line stderr warning) and skipped entirely. * `fetchConstruct` throws only when **no mesh** could be fetched: `Construct "<name>": no meshes could be fetched`. `FetchConstructOptions` also forwards the shared fetch tuning (`signal`, `cache`, `timeoutMs`, `retries`, `backoffMs`) and a `limit` (newest N entries within each mesh). See [Fetching](/concepts/fetching). ## Named constructs from config Constructs can be stored as JSON and referenced by name. The directories are searched in order: `NEUROWIRE_CONSTRUCTS` (`:` / `,` separated), then `~/.config/neurowire/constructs/` (or `$XDG_CONFIG_HOME/neurowire/constructs`). Files are matched as `<name>.construct.json` then `<name>.json`. Names must be simple identifiers; path-like names are rejected. ## Bundled `daily` So that `?src=daily` works with no setup, the API ships one built-in construct, `daily` ("Daily Brief"), with two meshes: a "Models" mesh (Claude Blog, Simon Willison) and a "Releases" mesh (Claude Code Releases). See `packages/api/src/constructs.ts`. User construct directories take precedence over the bundled defaults. --- --- url: /concepts/journals.md --- # Journals A **journal** is an append-only archive of what a source has published. Everything else in Neurowire deals in snapshots: a feed is whatever a site is showing right now, and yesterday's entries are gone once they fall off the page. A journal keeps them. That changes what you can ask. A feed answers "what is on the front page". A journal answers "what did this source publish in March", "what is new since I last looked", and "how much did they write about Rust this year". The format is [NWFJ](/formats/nwfj), the append-only sibling of [NWF](/formats/nwf). The encoder and decoder live in [`@neurowire/core`](/reference/core#journal); the on-disk store lives in [`@neurowire/ingest`](/reference/ingest#journal-store). ## Why NWF could not just be appended to An `nwf` document puts its complete dictionaries near the top and stores each entry's date as a delta from a feed-level timestamp. Adding one entry would mean rewriting the dictionaries and the header, which breaks append-only writing and invalidates any position recorded into the file. NWFJ keeps the same cell grammar and escaping and changes only what append-only forces: | | `nwf` (snapshot) | `nwfj` (journal) | |---|---|---| | Dictionaries | complete, near the top | grow one line at a time, before first use | | Timestamps | delta from the feed header | absolute, and `published` and `updated` each keep their own cell | | Entry position | implicit (line order) | an explicit `seq`, unique for the life of the journal | | Integrity | none | an optional chain, checkpointed | ## Cursors Every entry gets a sequence number, starting at 1 and never reused. A **cursor** is that number, optionally with the chain value it was taken at: ``` 128 the 128th entry 128.9f1c0f0b8ad0f0e3 ...and the chain value there, so a reader can verify it ``` A cursor is how you resume. Record the head after a read, hand it back next time, and you get exactly what arrived in between: ```bash neurowire journal head ai # 128 neurowire journal cat ai --cursor 128 -f json ``` Because sequence numbers are global to the journal, a cursor keeps working after the store rotates to a new segment file. ## Segments A journal is not one growing file. The store writes size-capped segments (`ai.00001.nwfj`, `ai.00002.nwfj`, ...) and rotates past 5 MB. Each segment repeats the header and starts its dictionaries fresh, so any segment can be read on its own, and dictionaries cannot grow without bound. Compaction drops whole old segments. A cursor pointing into a dropped segment is reported as **too old** rather than silently returning a partial answer, so the caller knows to re-read from the start instead of quietly missing entries. ## The format is its own index Journals stay queryable at size without a database beside them, because two things the format already writes double as an index: * **`seq` is the primary index.** It is monotonic and dense, and each segment covers a contiguous range, so finding a cursor means picking one segment, never scanning the archive. * **Dictionaries are skip filters.** A segment declares every author, tag, and source its entries can reference. A query for `tag:rust` can therefore rule out a segment whose tag dictionary holds no match, without opening it. The vocabulary is written anyway to make the segment self-contained, so pruning is free. The store keeps a `<id>.manifest.json` sidecar recording each segment's sequence range, date range, dictionaries, and entry keys. It is a **cache, never a second source of truth**: delete it, corrupt it, or edit a segment behind its back, and it is rebuilt by rescanning. ## Querying: no new language A journal is queried with the same flags a live fetch uses, because it runs the same code. `queryJournal` composes `filterEntries` and `selectEntries` from core, so an archive and a live feed cannot disagree about what a filter means: ```bash neurowire journal query ai --filter tag:rust --since 30d --sort date -f md ``` Three levels of access, in order of how much structure you need: 1. **`grep`.** Segments are one record per line, TAB-separated. `grep -c '^E' ai.00001.nwfj` counts entries. Good enough for a quick look. 2. **`journal query`.** Resolves interned references and escaping correctly, prunes segments, applies filters and windows. 3. **Anything else.** `journal cat ai -f json` hands the archive to duckdb, sqlite, or pandas. The journal stays the source of truth; those are disposable views. ## The chain Each record folds into a running hash, and `C` lines checkpoint it. `verifyJournal` recomputes the chain and reports any checkpoint that disagrees, which catches corruption, truncation, and reordering. ::: warning It is a checksum, not a signature The chain uses core's FNV-1a hash, the same one behind `stableId`, so that `@neurowire/core` stays portable and free of `node:crypto`. It detects damage; it does not prove authorship, because anyone who edits a journal can recompute it. The header carries a version cell so a future revision can introduce a cryptographic digest and signed checkpoints. ::: ## Journaling is opt-in Nothing is archived unless you ask. Add `--journal <id>` to any fetch: ```bash neurowire --mesh ai-news.json --journal ai neurowire --mesh ai-news.json --journal ai --watch --interval 30m ``` Entries the journal already holds are dropped on append, which is what makes this safe to run on a timer: re-fetching the same front page adds nothing. Journals live in `$NEUROWIRE_JOURNAL`, else `~/.config/neurowire/journal`, or wherever `--journal-dir` points. ::: tip Journals and watch state are different things `--state` remembers entry keys so a watch loop does not re-report or re-deliver the same item. It is a set of keys with no content and no history. `--journal` keeps the entries themselves, in order, with cursors. Use `--state` to avoid duplicate notifications; use `--journal` to keep a record you can come back to. They compose: a watch loop can do both. ::: ## What a journal is not ::: warning Not an output format A journal stores a feed's history rather than rendering it, so `nwfj` is deliberately **not** in `FORMATS` and `serialize()` does not know about it. Reading one back gives you an ordinary `NeurowireFeed` (`journalToFeed`), which every serializer already handles. ::: It is also not a database: no indexes beside the files, no query language. Exchanging journal deltas between machines is a separate concern, layered on top rather than built in: see [Sync](/concepts/sync). See the [NWFJ format](/formats/nwfj) for the line grammar, and the [CLI journal commands](/guide/cli#journals) for day-to-day use. A journal is also what makes a live [tail](/concepts/tail) resumable: cursors let a reader that dropped off the stream be replayed exactly what it missed. --- --- url: /concepts/tail.md --- # Tail A feed is a snapshot: whatever a source is showing right now. Everything in Neurowire can be used that way, one fetch at a time, and for most jobs that is the right shape. **Tail** is the other posture: keep the source open and be told what is new, the way `tail -f` follows a file. Nothing about the data changes. A tail is the same canonical [model](/concepts/model), the same [taps](/concepts/taps), the same [meshes](/concepts/meshes) and [constructs](/concepts/constructs). What changes is who is waiting for whom. ## One engine, three surfaces The polling loop lives in one place, [`pollFeed`](/reference/ingest#poll-engine) in `@neurowire/ingest`. Everything that follows a source consumes it, so none of them can drift on cadence, dedupe, or what happens when a fetch fails. | Surface | What it emits | Where | |---------|---------------|-------| | `neurowire tail` | one terminal line per entry as it arrives, or raw [NWFJ](/formats/nwfj) with `-f nwf` | [CLI tail mode](/guide/cli#tail-mode) | | `neurowire --watch` | one feed per tick, in `--format`, containing only that tick's new entries | [CLI watch mode](/guide/cli#watch-mode) | | `GET /tail` | server-sent events, one per entry, fanned out to every connected client | [HTTP API](/guide/http-api#get-tail) | The engine itself is an async generator and owns no I/O. You hand it a `load` function (a `fetchFeed`, a `fetchMesh`, a whole filter pipeline) and it hands back the entries that are new: ```ts import { fetchMesh, pollFeed } from '@neurowire/ingest' for await (const { fresh } of pollFeed(() => fetchMesh(mesh), { intervalMs: 60_000 })) { for (const entry of fresh) console.log(entry.title, entry.link) } ``` ## What "new" means New means "not in the seen-set", and the seen-set is keyed by `entryKey`: the entry's id, or its link when it has none. That is the same key [watch state](/guide/cli#watch-mode) and [journal dedupe](/concepts/journals) use, so the three agree by construction. Two consequences worth knowing: * **Order does not matter.** A source that reorders its front page, or backfills an old post, produces no noise. Only a key the loop has not reported before is new. * **The first tick is not special.** It runs immediately, with an empty seen-set, so a fresh tail reports the whole front page and then goes quiet. Seed the set (`--state` on the CLI, a journal on the server) when you want a restart to pick up where it left off instead. ::: tip Shaping happens before dedupe On the CLI, each tick fetches, then applies your `--filter`, `--since`, `--sort`, and `--limit` flags, and only then diffs against the seen-set. An entry that a filter excluded was never reported, so it is still new if a later tick lets it through. That is what makes `tail --mesh ai.json --filter tag:release` behave the way you would expect over hours. ::: ## Politeness is built in, not left to you Following a source means asking for it repeatedly, so the loop is deliberately conservative: | Guard | Value | Why | |-------|-------|-----| | Interval floor | 30s on the CLI, 60s on the API | A client cannot talk the loop into hammering an upstream. | | Jitter | up to +10% of the interval, added, never subtracted | Many tails on one host stop arriving in lockstep. The interval stays a floor. | | Conditional requests | `ETag` / `Last-Modified` per source | An unchanged source costs a `304` rather than a body. See [Fetching](/concepts/fetching). | | Shared loops | one poll per distinct target on the API | Fifty browsers following `ai-news` cost one upstream fetch per tick, not fifty. | Pick an interval that suits the source rather than the floor. A blog that posts weekly does not need a 30 second tail, and the floor is a safety rail, not a recommendation. ## A failed tick is not a failed tail A long-running loop that dies on the first flaky response is not much use. `load` throwing is isolated to its tick: the error is reported, the loop waits, and the next tick tries again. That sits on top of the retry and backoff [fetching](/concepts/fetching) already does inside a single tick, so a transient upstream problem is usually handled before the loop ever sees it. ``` [tail] error: Upstream responded 503 Service Unavailable for https://example.com/feed.xml ``` Cancellation is explicit and immediate: an `AbortSignal` ends the generator between ticks and mid-sleep, which is how the API tears a loop down the moment its last subscriber disconnects. ## Composing with journals A tail and a [journal](/concepts/journals) answer the two halves of the same question. The tail says what is arriving; the journal says what arrived. Put them together and a reader can drop off the wire without losing anything. **A tail stream is a journal being written.** `tail -f nwf` does not invent a second streaming encoding: it emits NWFJ, the same append-only format the on-disk store writes, because that format is already designed to be appended to one entry at a time. The output is a complete document, header, dictionary growth, and checkpoints included, so anything that reads an archive reads the stream. ```bash neurowire tail --mesh ai-news.json -f nwf > live.nwfj ``` **Cursors make resume lossless.** Give the API a journal that already exists and the route writes through it, so each event's SSE `id` is a real [cursor](/concepts/journals#cursors). A client that reconnects with `Last-Event-ID` is replayed everything after that cursor and then put back on the live stream: ```bash curl -N "http://localhost:8787/tail?src=ai-news&journal=ai" # ...connection drops at id 812, reconnect... curl -N -H 'Last-Event-ID: 812' "http://localhost:8787/tail?src=ai-news&journal=ai" ``` Without a journal the tail is live-only: event ids are stream-local, resume has nothing to replay from, and the `init` event says so rather than pretending otherwise. The route never creates a journal of its own, so replay is something an operator turns on, not something a client can ask for. ## Tail, watch, or a plain fetch | You want | Use | Because | |----------|-----|---------| | A feed, a file, a page, right now | a one-shot fetch | Nothing needs to stay running. Put it on cron if you want it repeated. | | To see posts appear while you work | `neurowire tail` | Line-per-entry output, readable as it scrolls. | | To pipe new entries into another tool | `neurowire tail -f nwf` | A stream of parseable records with no terminal formatting in it. | | A batch of new entries per interval, serialized | `neurowire --watch -f json` | One feed document per tick is what a downstream script wants to parse. | | To notify a channel | `--watch` or `tail` with `--sink` | Delivery of only the new entries. See [Sinks](/concepts/sinks). | | To fan one source out to many readers | `GET /tail` | One upstream poll, many subscribers, resume for free with a journal. | ::: tip Cron is often the right answer A tail holds a process open. If you only need a page rebuilt every morning or an archive topped up every hour, a one-shot fetch on a timer is simpler, survives reboots, and cannot leak a process. Reach for a tail when something is actually waiting on the output. ::: ## What tail is not ::: warning Not a delivery guarantee The loop reports what a source is showing when it looks. A post that appears and is deleted between two ticks is never seen, and there is no acknowledgement, no redelivery, and no ordering promise beyond the order the source lists. Journal-backed resume closes the gap for a client that disconnects; it does not turn polling into a message queue. ::: It is also polling by design. There is no WebSub or webhook ingestion, so nothing pushes to Neurowire, and there is no WebSocket transport: SSE covers the fan-out case over plain HTTP, through proxies, with reconnection already specified. Exchanging journal deltas between machines, so a second node pulls from a peer instead of re-fetching every upstream site, is [planned separately](https://github.com/starside-io/neurowire/blob/main/docs/plans/12-nwf-sync.md). See [CLI tail mode](/guide/cli#tail-mode) for the flags, [`GET /tail`](/guide/http-api#get-tail) for the event stream, and the [poll engine reference](/reference/ingest#poll-engine) for the library API. --- --- url: /concepts/sync.md --- # Sync A [journal](/concepts/journals) makes NWF a format you can append to. **Sync** makes it a protocol you can speak. One node fetches the open web and journals what it finds; other nodes ask it "what happened after position N" and get back exactly that. Nothing about the payload is new. Journals already number every entry, checkpoint a hash chain, and answer "everything after this cursor" locally. Sync is that same question asked over HTTP, which is why the protocol is four read-only `GET`s and no more: [`nwf-sync/1`](/formats/nwf-sync). The interesting part was already built. ## Why not just fetch the sources directly? The obvious objection: every machine can already run `fetchMesh` itself. Three cases answer it, and only the first is about bandwidth. ### A team following the same sources Eight developers each running the same 200-source construct make 1,600 requests a tick against a couple of hundred publishers, from eight IPs that together look like a small scraping operation. Some of those hosts rate-limit. A few will eventually block. Every developer also keeps taps current locally, so a site redesign breaks eight times and gets fixed eight times. Point them at one node instead and the publishers see 200 requests from one polite poller. The team reads one identical corpus rather than eight slightly different snapshots, and a broken [tap](/concepts/taps) is fixed once. ### The laptop that is closed all day This is the case direct fetching **cannot** solve, at any request budget. A feed is a snapshot of a front page. Open a laptop at 18:00 and a live fetch truthfully reports what those sites are showing *now*, which for a busy outlet may be the last three hours. Everything published while the lid was shut has already scrolled off. Fetching harder does not recover it, because the data is no longer on the page. A node that stayed awake recorded every entry as it appeared. The laptop asks for `?cursor=<where I left off>` and gets the missed window, in order, with nothing duplicated. The archive is the point; sync is just how it travels. ### Research that has to be reproducible Analyzing six months of coverage means every machine and every rerun must see the same corpus. Direct fetches give a different snapshot per machine and per hour, so a result cannot be checked by a colleague, or by yourself next week. A synced journal is addressed by sequence number and chain-verified. Cite "journal `ai`, seq 1 to 48210, chain `9f1c...`" and anyone who syncs that journal reproduces the analysis exactly, then keeps pulling deltas as it grows. The too-old-cursor error matters here too: it makes retention an explicit, loud event rather than a hole an archive quietly develops. ## Cursors and the short-circuit The steady state is one request that returns a number: ``` GET /sync/head?journal=ai -> { "head": 1284 } ``` A device on a train wakes up, asks a question worth about forty bytes, sees that its recorded cursor is already 1284, and goes back to sleep. Entries move only when the cursor is actually behind. That asymmetry is what makes polling on a short interval reasonable: being up to date is nearly free, so you can check often. Peer B asks Node A answers ### A cursor belongs to one peer A sequence number is meaningful only inside one journal on one node. When node B appends entries pulled from node A, B's store assigns **B's** numbering and computes **B's** chain. B is not a byte copy of A; it is a journal that happens to hold the same entries. So a client never compares its local head against a peer's. It records the peer's cursor per `(peer url, journal id)`, and it writes that cursor only after the entries have actually landed. A crash in between costs one re-pull, which the dedupe absorbs. The reverse order would silently lose entries, which is the failure you would never notice. ::: tip A cursor can also stop meaning anything A peer's journal can be rebuilt, restored from a backup, or have its directory repointed. Its head then restarts low while your cursor still sits past it. Left alone, every later sync would report "0 new" forever with a clean exit code, which is the worst kind of failure: silent and successful-looking. A head that moved backwards, or a chain hash that disagrees at the recorded position, is treated as divergence, and the cursor resets. ::: ## Merging cannot go wrong Entries are deduplicated by [entry key](/reference/core#entrykey) on append, the same rule that makes `--journal` safe to run on a timer. Three properties fall out of that one mechanism: * Pulling the same delta twice adds nothing the second time. * A node peering with both A and B, where B already forwarded A's entries, stores one copy. * A cycle of peers terminates instead of growing without bound. Which means topology is configuration, not code. `A -> B -> C` works: C gets what B already pulled from A. A diamond works. A relay that never touches the internet, pulling from a node on the LAN, works. Nobody has to design the graph carefully, because there is no arrangement of peers that produces duplicates or a loop. Provenance survives every hop too. `entry.source` travels inside the record, so a story that reached you through two relays still names the outlet that published it, not the peer it arrived through. ## The trust model, stated plainly Two mechanisms, and they protect against different things. Neither is a signature. | | What it does | What it does not do | |---|---|---| | **Hash chain** | detects corruption, truncation, and reordering of a journal, in transit or at rest | prove who wrote an entry: anyone who edits a journal recomputes the chain | | **Bearer token** | decides who may read a node's journals | say anything about the content behind it | Every pulled response is chain-verified **before** a single entry is appended, so a flipped byte aborts that response and merges nothing from it. That is a real guarantee, and it is a guarantee about damage, not about authorship. ::: warning Nothing here authenticates content origin `nwf-sync/1` does not sign anything. A peer that wants to hand you fabricated entries can, and the chain will verify perfectly, because the peer computed it. **You sync from peers you chose to trust**, over TLS you terminate yourself. Signed checkpoints are the designed-for v2 slot: the protocol carries a version header and NWFJ carries a version cell precisely so that upgrade has somewhere to go. ::: The same reasoning that keeps the chain non-cryptographic in [journals](/concepts/journals#the-chain) applies here: it uses core's FNV-1a hash so `@neurowire/core` stays portable and free of `node:crypto`. Sync inherits that property rather than working around it. ## Retention is a decision, not an accident Journals grow, and compaction drops whole old segments. A peer whose cursor pointed into a dropped segment does not get a partial answer: it gets a `410 Gone` naming the oldest entry still retained, and re-bootstraps from a snapshot of what the node still keeps. That is deliberately loud. A quiet fallback would let an archive develop holes that nobody discovers until someone tries to reproduce a result. The bootstrap costs one full transfer per peer and is then over, and the dedupe means the overlap adds nothing. ## Publishing is opt-in Nothing is exposed unless you say so. A node names the journal ids it publishes, and an unpublished id answers exactly the same `404` as one that does not exist, so a node's journal list is precisely what it published and no more. This matches the shape of the rest of Neurowire: journaling is opt-in, taps are explicit, meshes are files you wrote. Sync adds an operational dependency (a node someone has to run), so it asks to be turned on rather than assuming. ## When sync is the wrong tool One machine following twenty feeds should just fetch them. Sync earns its keep at **device count**, at **intermittent connectivity**, or when **history matters**, and is overhead otherwise. It is also not a general replacement for fetching. Somebody still has to be the node that talks to the open web, holds the taps, and notices when a site redesign breaks one. Sync moves that work to one place; it does not remove it. ## What sync is not * **Not push, gossip, or discovery.** Peers are configured URLs. No DHT, no broadcast, no automatic mesh formation. * **Not a conflict resolver.** Journals are append-only facts, so there is nothing to conflict over. There is no merge strategy to choose because no two nodes can disagree about what already happened. * **Not identity.** See [the trust model](#the-trust-model-stated-plainly). * **Not a new format.** The payload is [NWFJ](/formats/nwfj) segments exactly as they sit on disk. A synced journal is an ordinary journal: `journal cat`, `journal query`, mesh rendering, and the page generator all consume it with no new code. ## Next * [Federation](/guide/federation): build the three-node topology from scratch. * [`nwf-sync/1`](/formats/nwf-sync): the wire protocol, endpoint by endpoint. * [CLI sync and peers](/guide/cli#sync): the day-to-day commands. * [Journals](/concepts/journals): the archive underneath all of this. --- --- url: /concepts/sinks.md --- # Sinks A **sink** delivers a feed's entries somewhere: a Slack channel, a Discord channel, or any generic endpoint that accepts a JSON Feed. Where the rest of Neurowire reads sources, a sink writes results. Sinks are an I/O and delivery concern, so they live in the CLI, not the pure library: [`packages/cli/src/sinks.ts`](https://github.com/neurowire/neurowire). They pair naturally with the CLI's watch mode (see [CLI](/guide/cli)), which re-fetches on an interval and pushes only newly-seen entries to a sink. ## Auto-detection by host You give a sink a destination URL; the kind is detected from its host (`sinkKind(url)`): | Host contains | Sink kind | |---------------|-----------| | `slack.com` | `slack` | | `discord.com` or `discordapp.com` | `discord` | | anything else | `webhook` | There is no flag to set the kind; the URL decides. ## Payload per sink ### Slack Slack incoming-webhook body: a single `text` field. The text is a header line (`<feed title>: <N> new`) followed by up to 10 bullet lines, one per entry as `โ€ข <title> - <link>`, with an overflow line (`โ€ฆand N more`) when there are more. ```json { "text": "AI News: 3 new\nโ€ข Title one - https://...\nโ€ข Title two - https://..." } ``` The title-to-link separator is a hyphen with spaces, never an em-dash. ### Discord Discord webhook body: a single `content` field built from the same header-and-bullets text, then **capped at Discord's 2000-character limit** (truncated if longer). ```json { "content": "AI News: 3 new\nโ€ข Title one - https://..." } ``` ### Generic webhook A generic webhook gets the full [JSON Feed 1.1](/formats/json-feed) of the entries as the request body, sent with `Content-Type: application/feed+json`. ## `deliver(url, feed)` The one entry point. It picks the kind, builds the right body and content type, and POSTs: ```ts import { deliver } from '@neurowire/cli' const ok = await deliver('https://hooks.slack.com/services/...', feed) ``` `deliver` returns a `boolean` and **never throws**: a non-2xx response or a network error writes a one-line warning to stderr (`[sink] <host> failed: ...`) and returns `false`. That makes it safe to call from a long-running watch loop without a failed delivery crashing the process. ## No built-in retry or dedup ::: warning A sink fires once per call. It has **no built-in retry** and **no deduplication**. A transient failure is logged and dropped; it is not retried. ::: Deduplication is the watch loop's job, not the sink's. Use the CLI's `--state` flag with watch mode so each tick delivers only entries not seen on a previous tick. Without it, a watch run would re-deliver the whole feed every interval. See [CLI watch mode](/guide/cli) for wiring a source, an interval, `--state`, and one or more sinks together. --- --- url: /formats/nwf.md --- # NWF (Neurowire Feed) `nwf` is Neurowire's own compact, line-oriented feed format. It round-trips the model and stays small by interning repeated values, relativizing links, and storing dates as deltas. The serializer, parser, and validator live in `packages/core/src/serialize/nwf.ts`. | | | |---|---| | Format key | `nwf` | | Media type | `text/x-neurowire; charset=utf-8` | | Extension | `nwf` | | Functions | `toNwf(feed)`, `fromNwf(text)`, `validateNwf(text)` | | Types | `NwfIssue`, `NwfValidation` | ```ts import { toNwf, fromNwf, validateNwf } from '@neurowire/core' const text = toNwf(feed) // serialize const back = fromNwf(text) // parse (round-trip) const result = validateNwf(text) // validate with diagnostics ``` ## Layout Lines are LF-separated, cells within a line are TAB-separated. The first line is the magic header `NWF1`. Each subsequent line begins with a one-letter kind: ``` NWF1 magic + version F id title home self updatedEpoch authorRefs feed header A author0 author1 ... authors dictionary (interned) T tag0 tag1 ... tags dictionary (interned) S source0 source1 ... sources dictionary (interned; meshes) B https://blog.example.com/posts/ shared link prefix E id delta link authorRefs tagRefs title summary sourceRef one entry per line ``` How it stays compact: * **Interned dictionaries.** Authors (`A`), tags (`T`), and sources (`S`) are listed once and referenced by integer index. `authorRefs` and `tagRefs` are comma-separated indices into `A` / `T`; `sourceRef` is a single index into `S`. * **Relative links.** `B` holds the longest shared link prefix (trimmed to a path boundary, keeping scheme plus host). Entry links that start with the base are stored as `~rest`, expanded back at parse time. The `B` line is only emitted when a useful common base exists. * **Delta timestamps.** Each entry stores its date (`updated` else `published`) as a delta in seconds before the feed's `updatedEpoch`, or `-` when it has no date. * **Escaping.** Text cells escape backslash, TAB, CR, and LF. Within an author or source cell, sub-fields (name / url / email) are joined by the ASCII Unit Separator (0x1f), which never appears in feed text. `nwf` round-trips the list essentials: feed `id`/`title`/`home`/`self`/`updated`/`authors`, and entry `id`/`title`/`link`/`updated`/`summary`/`authors`/`tags`/`source`. It does not carry `generator`. The `sourceRef` column is appended last, so older NWF1 documents that omit it still parse. ## Annotated sample ``` NWF1 โ† magic + version F https://example.com/feed Example Blog https://example.com/ https://example.com/feed.nwf 1782640800 0 A Ada Lovelace โ† author index 0 T intro release โ† tag index 0, 1 B https://example.com/posts/ โ† shared link prefix E https://example.com/posts/hello 86400 ~hello 0 0 Hello, world A first post. ``` Reading the `E` line: id `https://example.com/posts/hello`, dated 86400 seconds (one day) before the feed's `updated`, link `~hello` (expands to `https://example.com/posts/hello`), author index `0` (Ada Lovelace), tag index `0` (intro), title `Hello, world`, summary `A first post.`, and no `sourceRef`. ::: info Round-trip `fromNwf(toNwf(feed))` reconstructs the canonical model. `fromNwf` throws if the document is missing the `NWF1` header. ::: ## Validation `validateNwf(text)` checks the document and returns line-numbered diagnostics. It validates the header, the single required feed (`F`) line, line kinds, cell counts, the `updatedEpoch` integer, entry `delta` format (`-` or an integer), empty ids/links, and that every dictionary reference is in range. When there are no errors it also parses the feed. ```ts interface NwfIssue { line: number // 1-based line number message: string } interface NwfValidation { valid: boolean errors: NwfIssue[] warnings: NwfIssue[] feed?: NeurowireFeed // present only when there are no errors } ``` Errors include: a missing/wrong header, missing or duplicate `F` line, non-integer `updatedEpoch`, an `E` line with fewer than 7 cells, an empty entry id, a malformed `delta`, and any author/tag/source reference that is not an integer or points past the dictionary. Warnings include: a `B` line with no base URL, an entry with an empty link, and a feed with no entries. ## The `validate` command The CLI ships a `validate` subcommand that runs `validateNwf` over a file or URL and exits non-zero when the document is not well-formed: ```bash neurowire validate feed.nwf neurowire validate https://example.com/feed.nwf ``` In dev: `pnpm validate <url>`. ## Publishing one, and reading it back NWF is not write-only. `detectKind` recognizes it by `Content-Type` (`text/x-neurowire`) or by its `NWF1` first line, so an NWF file served at a URL is a source like any feed: ```bash neurowire https://example.com/feed.nwf # terminal view neurowire https://example.com/feed.nwf -f md # convert to anything neurowire https://example.com/feed.nwf --watch # follow it ``` That also makes it a [mesh](/concepts/meshes) member, so one machine can publish a static file and others merge it with live feeds: ```json { "name": "Mixed", "sources": [ { "name": "Mirror", "url": "https://example.com/feed.nwf" }, { "name": "Live", "url": "https://blog.example.com/feed.xml" } ] } ``` A malformed document fails with the line number `validate` would print, for example `Invalid nwf document: line 4: entry has fewer than 7 cells`. The feed's `self` is preserved when the document carries one, and otherwise set to the URL it was fetched from, so a mirrored copy still records where it came from. For archives rather than snapshots, the journal protocol is [`nwf-sync/1`](/formats/nwf-sync). ## Usage CLI: ```bash neurowire https://example.com/feed.xml -f nwf ``` API: ```bash curl 'http://localhost:8787/feed?url=https://example.com/feed.xml&format=nwf' ``` See [Output formats](/concepts/output-formats) and [the CLI guide](/guide/cli). --- --- url: /formats/nwfj.md --- # NWFJ (Neurowire Feed Journal) `nwfj` is Neurowire's append-only sibling of [NWF](./nwf). Where an `nwf` document is a *snapshot* of a feed, a journal is a *log*: entries are appended over time, every entry carries a sequence number, and a reader can resume from a cursor instead of re-reading the whole file. The encoder, decoder, and query helpers live in `packages/core/src/journal.ts`; the segmented on-disk store lives in `packages/ingest/src/journal-store.ts`. | | | |---|---| | Media type | `application/x-nwf-journal` | | Extension | `nwfj` | | Version | `1` | | Core functions | `createJournalEncoder`, `resumeJournalEncoder`, `parseJournal`, `readJournalSince`, `journalHead`, `verifyJournal`, `queryJournal`, `journalToFeed` | | Store | `openJournalStore({ dir })` | A journal is deliberately **not** an output format: it is not a way to render a feed, it is a way to store its history. `serialize()` and the `FORMATS` registry are untouched. ## Why NWF1 cannot simply be appended to An `nwf` document puts its complete `A` / `T` / `S` dictionaries near the top and stores each entry's date as a delta from the feed-level `updated`. Appending one entry would mean rewriting the dictionaries and the header timestamp, which breaks append-only semantics and invalidates every cursor pointing into the file. NWFJ keeps the same cell grammar and escaping, and changes only what append-only forces it to change. ## Layout Lines are LF-separated, cells within a line are TAB-separated, and text cells use exactly the same escaping as `nwf` (backslash, TAB, CR, LF; sub-fields joined by the ASCII Unit Separator, 0x1f). ``` J 1 <journalId> <createdEpoch> journal header, one per segment F <feedId> <title> <home> <self> feed identity, re-emitted when it changes A+ <index> <author> dictionary growth, interned on first use T+ <index> <tag> S+ <index> <source> B <baseUrl> link prefix, may be re-declared E <seq> <updated> <published> <id> <link> <authorRefs> <tagRefs> <title> <summary> <sourceRef> C <seq> <hash> checkpoint, chain value at seq ``` What changes relative to NWF1, and why: * **Dictionaries grow incrementally.** `A+`, `T+`, and `S+` each declare one interned value with an explicit index, emitted immediately before the first `E` line that references it. A reader builds the same lookup tables `nwf` has, just lazily. Indices are per segment and always dense, starting at 0. * **Timestamps are absolute epoch seconds**, not deltas from a header that would go stale as the log grows. Both dates are kept in their own cell (`updated`, `published`), and either may be `-` when absent, so a journal round-trips dates losslessly rather than collapsing them into one field. * **Every `E` line carries a sequence number.** `seq` starts at 1 and increases by one per entry for the life of the journal, across segment rotations. It is the primary index. * **Checkpoints are optional.** A `C` line records the running chain value at a given `seq`. Readers may verify or ignore it. Unknown line kinds are reported as issues by the validator and skipped by the parser, so a future version can add record types without breaking older readers. ## Cursors A cursor is a sequence number plus an optional chain hash, written `"42"` or `"42.9f1c0f0b8ad0f0e3"`. `journalHead(text)` returns the cursor at the end of a journal, and `readJournalSince(text, cursor)` returns exactly the records after it. Because `seq` is global to the journal and never reused, a cursor survives segment rotation: the store maps a `seq` back to a segment through its manifest. ```ts import { journalHead, readJournalSince } from '@neurowire/core' const head = journalHead(text) // { seq: 128, hash: '...' } const { records } = readJournalSince(text, head) // [] until something new is appended ``` ## The chain Each record line after the header folds into a running chain value: `chain = hashHex(chain + "\n" + line)`, seeded from the journal id. `C` lines record that value and are themselves excluded from the fold. `verifyJournal(text)` recomputes the chain and reports any checkpoint that disagrees, along with the line number. The chain uses core's FNV-1a `hashHex`, the same 64-bit hash behind `stableId`. It detects corruption, truncation, and accidental reordering. **It is a checksum, not a cryptographic digest**, and it does not prove authorship: anyone who edits a journal can recompute it. Keeping it here preserves core's portability promise (no `node:crypto`, no dependencies beyond zod). The `J` line carries a version cell so a future revision can introduce a cryptographic digest and signed checkpoints where a stronger guarantee is needed. ## The format is its own index Two properties give journals indexed access with no index files beside them: * **`seq` is the primary index.** It is monotonic and dense, and each segment covers a contiguous seq range, so finding a cursor means picking a segment and scanning it, never scanning the whole journal. * **Dictionaries are per-segment skip filters.** A segment's `A+` / `T+` / `S+` lines declare every author, tag, and source any entry in that segment can reference. A query filtering on `tag:rust` can therefore skip a segment whose tag dictionary holds no match without decoding a single `E` line. The vocabulary is written anyway to make the segment self-contained, so using it for pruning is free. The store keeps a `<id>.manifest.json` sidecar recording, per segment, its seq range, its date range, its dictionaries, and its entry keys. The manifest is fully derivable by scanning the segments, so a missing or corrupt manifest is rebuilt rather than fatal. It is a cache, never a second source of truth. ## Querying No query language is invented. The existing filter engine is pointed at journal records, so an archive answers exactly what a live feed answers: ```ts import { queryJournal } from '@neurowire/core' const entries = queryJournal(records, { filter: { include: [{ field: 'tag', pattern: 'rust' }] }, from: Date.now() - 30 * 86_400_000, sort: 'date', limit: 20, }) ``` `queryJournal` composes `filterEntries` and `selectEntries` from core, so journal results and live-feed results are identical for the same spec by construction. The store's `query(id, spec)` adds planning on top: it prunes segments by date range and dictionary vocabulary, then decodes only the survivors. From the CLI: ```bash neurowire journal query ai --filter tag:rust --since 30d -f md ``` Because records are one line each and TAB-separated, crude research also works with ordinary text tools: ```bash grep -c '^E' ~/.config/neurowire/journal/ai.00001.nwfj ``` When research outgrows the built-ins, `neurowire journal cat ai -f json` hands the whole archive to duckdb, sqlite, or pandas. The journal stays the source of truth and those databases are disposable views. ## Segments and rotation The store writes `<id>.<nnnnn>.nwfj` segments and rotates to a new one past a size cap (5 MB by default). Each segment repeats the `J` header and starts its dictionaries fresh, so any segment can be read on its own, and rotation naturally bounds dictionary growth. Compaction drops whole old segments; a cursor pointing into a dropped segment is reported as too old, so the caller can fall back to reading from the oldest retained segment instead of silently missing entries. ## Annotated sample ``` J 1 ai 1782640800 F https://example.com/feed Example Blog https://example.com/ B https://example.com/ T+ 0 release E 1 - 1782640800 post-1 ~posts/one 0 First post S+ 0 Example E 2 - 1782727200 post-2 ~posts/two Second post A summary 0 C 2 9573a6ab74a4ba94 ``` Line by line (cells are TAB-separated, and consecutive tabs mean an empty cell): | Line | Reading | |------|---------| | `J` | Version 1, journal id `ai`, created 2026-06-28. | | `F` | The feed identity. The trailing empty cell is `self`, which this feed did not declare. | | `B` | The link prefix, taken from the feed's `home`. | | `T+` | The tag `release` is interned as index 0, right before the entry that first uses it. | | `E` (seq 1) | No `updated` (`-`), published at 1782640800, id `post-1`, link `~posts/one` (relative to `B`), no authors, `tagRefs` = `0`, title `First post`, no summary, no source. | | `S+` | The source `Example` is interned as index 0, again just before its first use. | | `E` (seq 2) | Same shape, with a summary and `sourceRef` = `0`. It declares no tags, so its `tagRefs` cell is empty. | | `C` | A checkpoint: the chain value after sequence 2. | Notice that the second entry reuses nothing from a dictionary it does not need, and that neither entry repeats the feed identity: `F` and `B` were emitted once and stay in effect until something changes. ## Round-trip guarantees A journal round-trips the same list essentials `nwf` does (feed `id` / `title` / `home` / `self`, and entry `id` / `title` / `link` / `updated` / `summary` / `authors` / `tags` / `source`), plus `published` as its own cell. It does not carry `generator`, and it does not store feed-level authors: a journal spans many feed identities over its lifetime, so per-entry authorship is the only attribution that stays meaningful after a merge. --- --- url: /formats/nwf-sync.md --- # NWF Sync (peer delta exchange) `nwf-sync/1` is the protocol two Neurowire nodes speak to exchange journal deltas. One node aggregates sources into an [NWFJ journal](./nwfj); other nodes pull "everything after my cursor" instead of re-fetching every upstream site themselves. It is deliberately boring transport over a novel payload: plain pull-only HTTP `GET`s, no push, no gossip, no discovery. The interesting part is the journal, which already numbers every entry and checkpoints a hash chain, so a delta is just "the records after seq N". | | | |---|---| | Protocol name | `nwf-sync` | | Version | `1` | | Version header | `NWF-Sync-Version: 1` on every `/sync/*` response | | Payload media type | `application/x-nwf-journal` (NWFJ) | | Server | `packages/api/src/sync.ts`, mounted at `/sync/*` | | Client | `packages/ingest/src/sync.ts` (`pullJournal`, `syncPeers`) | ## Endpoints | Endpoint | Returns | |----------|---------| | `GET /sync/journals` | JSON: the published journals, each with `id`, `title`, `head`, `updated`, `entries` | | `GET /sync/head?journal=<id>` | JSON `{ journal, head, hash? }`, the cheap poll target | | `GET /sync/since?journal=<id>&cursor=<c>` | One NWFJ segment containing records after `c`; `204` when up to date, `410` when `c` is too old | | `GET /sync/snapshot?journal=<id>&cursor=<c>` | The same as `since`, except a too-old cursor is clamped to the oldest retained segment instead of failing | Every response, including errors, carries `NWF-Sync-Version: 1`. A client that sees a version it does not implement must not merge the body. An **absent** header is tolerated (a proxy may have stripped it, and every other check still applies); a header naming another version is not, because v2 is free to change what a segment or a cursor means. ### `GET /sync/journals` ```json { "version": 1, "journals": [ { "id": "ai", "title": "AI News", "head": 1284, "hash": "9f1c0f0b8ad0f0e3", "entries": 1284, "segments": 3, "bytes": 812004, "updated": "2026-08-27T09:11:04.000Z" } ] } ``` Only published journals are listed (see [Publishing](#publishing)). `updated` is the newest entry timestamp known to the manifest, omitted when no entry is dated. `head` is the journal's last sequence number, `0` for an empty journal (an explicitly published id is listed even before it holds anything). `title` comes from an `F` line, which is re-emitted only when identity changes and so sits at the top of the newest segment in every ordinary journal. Servers should read a bounded prefix rather than decoding the whole segment: on an open node this endpoint is reachable without a token, and a full decode per published journal per request is a free amplifier. The title is informational, and omitting it is better than paying megabytes to produce it. ### `GET /sync/head` ```json { "journal": "ai", "head": 1284, "hash": "9f1c0f0b8ad0f0e3" } ``` This is the steady-state request: a client that is already current asks one question worth a few dozen bytes and goes back to sleep. `hash` is the chain value at `head` when the journal's last segment ends in a checkpoint at that sequence number; it is absent for an empty journal. ### `GET /sync/since` | Query | Required | Default | Description | |-------|----------|---------|-------------| | `journal` | yes | - | The journal id. | | `cursor` | no | `0` | The last sequence number the client already holds, in this server's numbering. Accepts `42` or `42.<hash>`; the hash half is ignored by the server (clients verify the chain themselves). | Responses: | Status | Meaning | |--------|---------| | `200` | An NWFJ segment. Body media type `application/x-nwf-journal`. | | `204` | The cursor is at or past the head. No body. | | `400` | `journal` missing, or `cursor` is not a non-negative integer (optionally followed by `.<hash>`). | | `401` | A token is configured and the request did not present it. | | `404` | The journal is not published (or does not exist, see [Publishing](#publishing)). | | `410` | The cursor predates the oldest retained segment. Body points at `/sync/snapshot`. | A `200` carries these response headers: | Header | Description | |--------|-------------| | `NWF-Sync-Version` | `1`. | | `NWF-Sync-Journal` | The journal id. | | `NWF-Sync-Head` | The journal's head sequence number, so a client learns it without a second request. | | `NWF-Sync-Range` | `<firstSeq>-<lastSeq>`, the sequence range of the segment in the body. | | `NWF-Sync-Segment` | The segment's index within the journal. | | `NWF-Sync-Complete` | `1` when this segment reaches the head, `0` when more segments remain. | ### `GET /sync/snapshot` Identical to `since`, including its headers, with one difference: a cursor older than the oldest retained segment is clamped to that segment rather than answered with `410`. It is therefore both the bootstrap path for a brand new client (`cursor=0`) and the recovery path after a `410`. `snapshot` never returns `410`. It still returns `204` when the cursor is at or past the head. ## One segment per response `since` and `snapshot` return **exactly one segment**, never a concatenation, even when the client is many segments behind. The client loops on `NWF-Sync-Complete: 0`, passing the previous response's `lastSeq` as the next cursor. This is a correctness requirement, not a throughput choice. NWFJ dictionary indices (`A+`, `T+`, `S+`) are per segment and dense from `0`, and the hash chain reseeds from the journal id at each `J` header line. Two segments glued together therefore decode wrongly (the second segment's `E` lines resolve their author, tag, and source references against the first segment's dictionaries) and verify wrongly (the chain never restarts). Serving whole segments keeps every response a standalone, self-describing, independently verifiable NWFJ document. It also keeps the server honest about work: a response is a file read, not a decode-and-re-encode. Re-encoding a delta would renumber the dictionary and break chain continuity with what the server actually stores, so the bytes on the wire are the bytes on disk. The consequence a client must handle: **a `200` normally contains records at or before the cursor**, because the segment holding `cursor + 1` also holds everything before it in that segment. Clients drop records with `seq <= cursor` when merging. Appending them anyway is harmless, since the store dedupes by entry key, but the drop keeps the reported counts truthful. A live hub appends to its newest segment while serving. Servers must therefore stream exactly the byte length the response's `NWF-Sync-Range` was measured from, so a body can never run past the range the same response declared. Clients should be forgiving in the same direction: a body ending **short** of the declared range is a truncated transfer and must abort the merge, but a body ending **past** it is a benign race, and those extra records are verified like any others. ## Cursors are per peer, not global A sequence number is meaningful only inside one journal on one node. When node B appends entries pulled from node A, B's store assigns B's own sequence numbers and computes B's own chain: B is not a byte copy of A, it is a journal that happens to hold the same entries. So a client must **not** compare its local head against a peer's head. It records the peer's cursor per `(peer url, journal id)`, by default in `~/.config/neurowire/peers-state.json`: ```json { "version": 1, "peers": { "https://hub.example.com\tai": { "seq": 1284, "hash": "9f1c0f0b8ad0f0e3" } } } ``` State is written **only after the append succeeds**. A crash between the pull and the write costs one re-pull, which the entry-key dedupe absorbs; the reverse order would silently lose entries. ### When a cursor stops meaning anything A peer's journal can be rebuilt, restored from a backup, or have its store directory repointed. Its head then restarts low while the client still holds a cursor from the old one, and every sync answers "nothing new" forever with a clean exit code. That failure is silent, which makes it the worst kind. Two signals catch it, and a client must check both against the recorded cursor: * the peer's head is **lower** than the recorded `seq`, or * the peer's head equals the recorded `seq` but its `hash` differs from the recorded one. Either means this is not the journal the cursor points into. The client resets to `seq 0` and pulls again, which costs bandwidth and nothing else: the entry-key dedupe adds only what is genuinely new. The reset cursor must be written even when the rebuilt journal turns out to be empty, or the stale one survives to stall the next sync too. ## Merging is idempotent The store drops entries whose [entry key](/reference/core#entrykey) it already holds. Three consequences fall out for free: * Pulling the same delta twice adds nothing the second time. * A node that peers with both A and B, where B already forwarded A's entries, stores one copy. * A cycle of peers terminates instead of growing without bound. `entry.source` travels inside the record, so provenance survives every hop: an entry that reached you through two relays still names the outlet that published it, not the peer it arrived through. ## Integrity Every `200` body is a complete NWFJ segment with its own `C` checkpoint lines. A client runs `verifyJournal(text)` on the body **before** appending anything from it. A mismatch aborts that response's merge with a named error and appends nothing from it; segments already merged in the same pull stay merged, because each was verified on its own and the store's dedupe makes re-pulling them free. A client should also reject a body whose `J` header names a different journal id than it asked for, and a body that carries entries but no checkpoint at all (a vacuous verification). **The chain is a checksum, not a signature.** It uses core's FNV-1a `hashHex`, so it detects corruption, truncation, and reordering in transit or at rest. It does not prove who wrote the journal: anyone who edits a segment can recompute the chain. Nothing in `nwf-sync/1` authenticates content origin. You sync from peers you chose to trust, over TLS you terminate yourself. The `J` version cell and this protocol's version header are the upgrade path: signed checkpoints are the designed-for v2 slot. ## Authentication Optional and static. When the server has a token configured (`NEUROWIRE_SYNC_TOKEN`, or `token` in `~/.config/neurowire/sync.json`), every `/sync/*` request must present it: ``` Authorization: Bearer <token> ``` Anything else gets `401` with `WWW-Authenticate: Bearer`. The token is compared in constant time. When no token is configured the endpoints are open, which is the deliberate self-host stance: the operator owns transport security and network exposure, and Neurowire does not pretend to. The auth check runs before the publish check, so an unauthenticated caller cannot use `404` versus `200` to enumerate which journals a node holds. ## Publishing Nothing is published by default. A node exposes journals explicitly: * `NEUROWIRE_SYNC_PUBLISH`: a `:`- or `,`-separated list of journal ids, or `*` for every journal in the store directory. * `~/.config/neurowire/sync.json` (or `$NEUROWIRE_SYNC_CONFIG`): ```json { "publish": ["ai", "rust"], "token": "a-long-random-string" } ``` The env var wins over the file. Journals are read from the usual store directory (`$NEUROWIRE_JOURNAL`, else `~/.config/neurowire/journal`). An unpublished id and a nonexistent id both answer `404` with the same body. The distinction is deliberately invisible: a node's journal list is exactly what it published. ## The handshake, end to end ``` B: GET /sync/head?journal=ai -> { head: 1284 } local cursor for (A, ai) is 1284, stop. one request, no body. B: GET /sync/head?journal=ai -> { head: 1310 } B: GET /sync/since?journal=ai&cursor=1284 -> 200, segment 3, range 1201-1310, complete 1 verify chain, drop seq <= 1284, append 26 entries, write cursor 1310. B: GET /sync/since?journal=ai&cursor=12 -> 410 { snapshot: "/sync/snapshot?journal=ai" } B: GET /sync/snapshot?journal=ai&cursor=0 -> 200, segment 1, complete 0 B: GET /sync/snapshot?journal=ai&cursor=800 -> 200, segment 2, complete 0 B: GET /sync/snapshot?journal=ai&cursor=1200-> 200, segment 3, complete 1 ``` ## Not in v1 * No push, no gossip, no peer discovery, no DHT. Peers are configured URLs. * No signing or identity. The hash chain is integrity, not authenticity. * No conflict resolution beyond the idempotent entry-key merge. Journals are append-only facts, so there is nothing to conflict. * No quotas, rate limits, or abuse controls beyond the bearer token. * No range within a segment. The unit of transfer is a segment, and retention is tuned with the store's segment size and compaction. ## See also * [Sync](/concepts/sync), the concept: why delta exchange beats fan-out fetching, and the trust model. * [NWFJ](./nwfj), the payload format. * [Journals](/concepts/journals), the archive underneath it. * [Federation](/guide/federation), the three-node setup guide. * [`@neurowire/api`](/reference/api), the server reference. * [`@neurowire/ingest`](/reference/ingest), the client reference. --- --- url: /formats/atom.md --- # Atom Atom 1.0 is Neurowire's default XML feed format. The serializer lives in `packages/core/src/serialize/atom.ts` and is exported as `toAtom`. | | | |---|---| | Format key | `atom` | | Media type | `application/atom+xml; charset=utf-8` | | Extension | `xml` | | Function | `toAtom(feed)` | ```ts import { toAtom } from '@neurowire/core' const xml = toAtom(feed) ``` `toAtom` is a pure function: it takes a [`NeurowireFeed`](/concepts/model) and returns an Atom 1.0 document as a string. No network, no DOM. ## Field mapping The canonical model maps to Atom elements as follows. ### Feed level | Model field | Atom output | Notes | |---|---|---| | `feed.id` | `<id>` | Text-escaped. | | `feed.title` | `<title>` | Text-escaped. | | `feed.updated` | `<updated>` | Coerced to RFC 3339. Falls back to the Unix epoch (`1970-01-01T00:00:00.000Z`) when missing or unparseable. | | `feed.home` | `<link rel="alternate" href="..."/>` | Emitted only when present. | | `feed.self` | `<link rel="self" href="..."/>` | Emitted only when present. | | `feed.authors[]` | `<author>` blocks | Each with `<name>`, optional `<uri>` (from `url`), optional `<email>`. | | `feed.generator` | `<generator version="...">name</generator>` | The `version` attribute is emitted only when set. | | `feed.entries[]` | `<entry>` blocks | One per entry, in order. | ### Entry level | Model field | Atom output | Notes | |---|---|---| | `entry.id` | `<id>` | Text-escaped. | | `entry.title` | `<title>` | Text-escaped. | | `entry.link` | `<link rel="alternate" href="..."/>` | Attribute-escaped. | | `entry.updated` / `entry.published` | `<updated>` | Uses `updated`, else `published`, else the feed's `updated`. Coerced to RFC 3339. | | `entry.published` | `<published>` | Emitted only when present. | | `entry.authors[]` | `<author>` blocks | Same shape as feed authors. | | `entry.tags[]` | `<category term="..."/>` | One per tag. | | `entry.summary` | `<summary type="text">...</summary>` | Emitted only when present. | ::: info Escaping Text content escapes `&`, `<`, and `>`. Attribute values escape those plus `"`. There is no other transformation: the model is assumed to hold plain text. ::: ::: tip Dates All timestamps run through an RFC 3339 coercion step (`Date.parse` then `toISOString`). Unparseable values fall back: the feed `updated` falls back to the epoch, and an entry's date falls back to the feed `updated`. ::: ## Sample output ```xml <?xml version="1.0" encoding="utf-8"?> <feed xmlns="http://www.w3.org/2005/Atom"> <id>https://example.com/feed</id> <title>Example Blog 2026-06-28T10:00:00.000Z Ada Lovelace https://example.com/ada Neurowire https://example.com/posts/hello Hello, world 2026-06-27T09:30:00.000Z 2026-06-27T09:30:00.000Z Ada Lovelace A first post. ``` ## Usage CLI: ```bash neurowire https://example.com/feed.xml -f atom ``` API: ```bash curl 'http://localhost:8787/feed?url=https://example.com/feed.xml&format=atom' ``` See [Output formats](/concepts/output-formats) for the full format list and [the CLI guide](/guide/cli) for flags. --- --- url: /formats/json-feed.md --- # JSON Feed Neurowire serializes to JSON Feed 1.1. The serializer lives in `packages/core/src/serialize/jsonfeed.ts`. | | | |---|---| | Format key | `json` | | Media type | `application/feed+json; charset=utf-8` | | Extension | `json` | | Functions | `toJsonFeed(feed)`, `toJsonFeedObject(feed)` | | Type | `JsonFeedDocument` | ```ts import { toJsonFeed, toJsonFeedObject } from '@neurowire/core' import type { JsonFeedDocument } from '@neurowire/core' const text = toJsonFeed(feed) // a JSON string, trailing newline const doc: JsonFeedDocument = toJsonFeedObject(feed) // the plain object ``` `toJsonFeedObject` builds the document object; `toJsonFeed` is `JSON.stringify(..., null, 2)` of that object with a trailing newline. Use the object form when you want to embed or post-process the feed; use the string form when you want bytes to write or send. ## Field mapping ### Feed level | Model field | JSON Feed key | Notes | |---|---|---| | (constant) | `version` | Always `"https://jsonfeed.org/version/1.1"`. | | `feed.title` | `title` | Always present. | | `feed.home` | `home_page_url` | Emitted only when present. | | `feed.self` | `feed_url` | Emitted only when present. | | `feed.authors[]` | `authors[]` | Emitted only when non-empty. Each author is `{ name, url? }`. | | `feed.entries[]` | `items[]` | One per entry, in order. | The feed `id` and `generator` are not represented in the JSON Feed output. ### Item (entry) level | Model field | JSON Feed key | Notes | |---|---|---| | `entry.id` | `id` | Always present. | | `entry.link` | `url` | Always present. | | `entry.title` | `title` | Always present. | | `entry.summary` | `summary` | Emitted only when present. | | `entry.published` | `date_published` | RFC 3339. Omitted when missing or unparseable. | | `entry.updated` | `date_modified` | RFC 3339. Omitted when missing or unparseable. | | `entry.authors[]` | `authors[]` | Emitted only when non-empty. Each is `{ name, url? }`. | | `entry.tags[]` | `tags[]` | Emitted only when non-empty (an array of strings). | ::: info Authors An author maps to `{ name }`, plus `url` only when the model `Person.url` is set. The `email` field of a `Person` is not carried into JSON Feed. ::: ::: tip Dates `date_published` and `date_modified` run through `Date.parse` then `toISOString`. Unparseable or missing values are simply omitted, rather than substituted with a fallback. ::: ## Sample output ```json { "version": "https://jsonfeed.org/version/1.1", "title": "Example Blog", "home_page_url": "https://example.com/", "feed_url": "https://example.com/feed.xml", "authors": [ { "name": "Ada Lovelace", "url": "https://example.com/ada" } ], "items": [ { "id": "https://example.com/posts/hello", "url": "https://example.com/posts/hello", "title": "Hello, world", "summary": "A first post.", "date_published": "2026-06-27T09:30:00.000Z", "date_modified": "2026-06-27T09:30:00.000Z", "authors": [{ "name": "Ada Lovelace" }], "tags": ["intro"] } ] } ``` Feed-level metadata is emitted before `items`, matching the JSON Feed convention. ## Usage CLI: ```bash neurowire https://example.com/feed.xml -f json ``` API: ```bash curl 'http://localhost:8787/feed?url=https://example.com/feed.xml&format=json' ``` See [Output formats](/concepts/output-formats) and [the CLI guide](/guide/cli). --- --- url: /formats/markdown.md --- # Markdown Neurowire can render a feed as a human-readable Markdown digest. The serializer lives in `packages/core/src/serialize/markdown.ts` and is exported as `toMarkdown`. | | | |---|---| | Format key | `md` | | Media type | `text/markdown; charset=utf-8` | | Extension | `md` | | Function | `toMarkdown(feed)` | ```ts import { toMarkdown } from '@neurowire/core' const md = toMarkdown(feed) ``` This is a presentation-oriented digest, meant to be read or pasted into a document. It is not round-trippable like [NWF](/formats/nwf): it carries the list essentials in a friendly layout, not every model field. ## Field mapping ### Feed header * `feed.title` becomes the top-level heading: `# Title`. * A metadata line follows, joining the present items with `ยท`: * `Updated ` when `feed.updated` is set (formatted as a `YYYY-MM-DD` date). * `[Home](url)` when `feed.home` is set. * `[Feed](url)` when `feed.self` is set. ### Each entry * `### [title](link)` heading linking the title to `entry.link`. * A metadata line (joined by `ยท`) when any of these are present: * The date, `entry.updated` else `entry.published`, formatted `YYYY-MM-DD`. * The authors, joined by `, ` (names only). * The tags, each wrapped in backticks (`` `tag` ``), space-separated. * The `entry.summary` on its own line when present. ::: tip Dates Dates are sliced to the `YYYY-MM-DD` day (first 10 characters of the ISO string). An unparseable date string is passed through unchanged rather than dropped. ::: The feed `id`, `self`/`home` beyond the header links, entry `id`, and `source` are not shown in the digest. ## Sample output ```markdown # Example Blog Updated 2026-06-28 ยท [Home](https://example.com/) ยท [Feed](https://example.com/feed.xml) ### [Hello, world](https://example.com/posts/hello) 2026-06-27 ยท Ada Lovelace ยท `intro` A first post. ### [Second post](https://example.com/posts/second) 2026-06-26 ยท Grace Hopper Another update with no tags. ``` ## Usage CLI: ```bash neurowire https://example.com/feed.xml -f md ``` API: ```bash curl 'http://localhost:8787/feed?url=https://example.com/feed.xml&format=md' ``` See [Output formats](/concepts/output-formats) and [the CLI guide](/guide/cli). --- --- url: /formats/rss.md --- # RSS RSS 2.0 output, new in `@neurowire/core` 0.7.0. The serializer lives in `packages/core/src/serialize/rss.ts` and is exported as `toRss`, alongside the `toRfc822` date helper. | | | |---|---| | Format key | `rss` | | Media type | `application/rss+xml; charset=utf-8` | | Extension | `xml` | | Functions | `toRss(feed)`, `toRfc822(iso)` | ```ts import { toRss, toRfc822 } from '@neurowire/core' const xml = toRss(feed) ``` `toRss` produces a valid RSS 2.0 document. It declares only the namespaces it actually uses: `xmlns:atom` appears only when the feed has a `self`, and `xmlns:dc` only when at least one entry has an author. ## The `toRfc822` helper RSS dates use the RFC 822 format. `toRfc822(iso)` converts an ISO-ish date string to an RFC 822 date in GMT, for example `Wed, 02 Oct 2024 13:00:00 GMT`. It returns `undefined` when the input is unparseable, so callers can omit the element rather than emit a bad date. ```ts toRfc822('2024-10-02T13:00:00Z') // "Wed, 02 Oct 2024 13:00:00 GMT" toRfc822('not a date') // undefined ``` ## Field mapping ### Channel level | Model field | RSS output | Notes | |---|---|---| | `feed.title` | `` | Text-escaped. | | `feed.home` / `feed.self` / `feed.id` | `<link>` | First present of `home`, `self`, then `id`. | | `feed.title` | `<description>` | RSS requires a non-empty description, so it falls back to the title. | | `feed.self` | `<atom:link href="..." rel="self" type="application/rss+xml"/>` | Emitted only when `self` is set (and the `atom` namespace is then declared). | | `feed.updated` | `<lastBuildDate>` | RFC 822 via `toRfc822`. Omitted when unparseable. | | `feed.generator` | `<generator>` | Text is `name` plus ` version` when a version is set. | | `feed.entries[]` | `<item>` blocks | One per entry, in order. | ### Item (entry) level | Model field | RSS output | Notes | |---|---|---| | `entry.title` | `<title>` | Text-escaped. | | `entry.link` | `<link>` | Text-escaped. | | `entry.id` | `<guid isPermaLink="...">` | `isPermaLink="true"` only when `id` equals `link` (synthetic ids are not resolvable URLs, so they are `false`). | | `entry.published` / `entry.updated` | `<pubDate>` | RFC 822 via `toRfc822`, using `published` else `updated`. Omitted when unparseable. | | `entry.authors[]` | `<dc:creator>` | One per author (name only). Triggers the `dc` namespace. | | `entry.tags[]` | `<category>` | One per tag. | | `entry.summary` | `<description>` | Emitted only when present. | | `entry.source` | `<source url="...">name</source>` | Emitted only when `source.url` is set (RSS requires the `url` attribute). The element text is `source.name`, falling back to the url. | ::: info Namespaces declared on demand The root `<rss>` element starts with `version="2.0"`. `xmlns:atom="http://www.w3.org/2005/Atom"` is added only when the feed has a `self`; `xmlns:dc="http://purl.org/dc/elements/1.1/"` only when some entry has authors. This keeps the document minimal. ::: ## Sample output ```xml <?xml version="1.0" encoding="utf-8"?> <rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/"> <channel> <title>Example Blog https://example.com/ Example Blog Sun, 28 Jun 2026 10:00:00 GMT Neurowire 0.1.0 Hello, world https://example.com/posts/hello https://example.com/posts/hello Sat, 27 Jun 2026 09:30:00 GMT Ada Lovelace intro A first post. Example Blog ``` ## Usage CLI: ```bash neurowire https://example.com/feed.xml -f rss ``` API: ```bash curl 'http://localhost:8787/feed?url=https://example.com/feed.xml&format=rss' ``` See [Output formats](/concepts/output-formats) and [the CLI guide](/guide/cli). --- --- url: /formats/opml.md --- # OPML OPML 2.0 is a subscription-list interchange format, used to move feed lists between readers. Neurowire can export a mesh or construct to OPML and import an OPML file into a mesh. ::: info Not a feed format OPML is **not** a feed serializer. It is not one of the `serialize()` formats (`nwf`, `atom`, `rss`, `json`, `md`) and it does not carry articles. It is lossy interchange: it lists which sources a mesh or construct subscribes to, so you can hand that list to a feed reader, or build a mesh from a reader's export. ::: | Direction | Function | Package | |---|---|---| | Export a mesh | `meshToOpml(mesh)` | `@neurowire/core` (`packages/core/src/opml.ts`) | | Export a construct | `constructToOpml(construct)` | `@neurowire/core` | | Import to a mesh | `opmlToMesh(xml, name?)` | `@neurowire/ingest` (`packages/ingest/src/opml.ts`) | ## Export ```ts import { meshToOpml, constructToOpml } from '@neurowire/core' const opml = meshToOpml(mesh) const repoOpml = constructToOpml(construct) ``` ### `meshToOpml` Serializes a mesh to a flat OPML 2.0 list: one `` per source. Each source outline carries `type="rss"`, `text` and `title` set to the source name, and both `xmlUrl` and `htmlUrl` pointing at the source URL. The OPML `head/title` is the mesh name. ### `constructToOpml` Serializes a construct with two-level nesting: one category `` per mesh, each wrapping its source outlines. This mirrors the construct's per-mesh grouping. A reference member (`{ ref }`) carries no sources of its own, so it emits an empty category outline (named after the ref); a mesh with zero sources likewise emits an empty category outline. ## Import ```ts import { opmlToMesh } from '@neurowire/ingest' const mesh = opmlToMesh(xml, 'My Reader') ``` `opmlToMesh` parses an OPML 2.0 subscription list into a [`Mesh`](/concepts/meshes): * Every `` that has an `xmlUrl` becomes a source `{ name, url }`. The name is the outline's `text`, else its `title`, else the URL's host (falling back to the raw URL). * Nested categories are walked recursively and **flattened**: a mesh has no per-source grouping, so two-level structure (as exported from a construct) collapses to one list. Outlines without an `xmlUrl` (bare categories) are skipped, but their children are still visited. * The mesh name is the `name` argument (trimmed), else the OPML `head/title`, else `"imported"`. * The result is validated against the core `Mesh` schema (`MeshSchema.parse`), so a malformed result throws. `opmlToMesh` throws on malformed XML or when the document has no `` root. ## Sample OPML A mesh exported with `meshToOpml`: ```xml AI News ``` A construct exported with `constructToOpml` (two levels): ```xml Daily ``` ## CLI The `opml` subcommand group has `export` and `import`: ```bash # Export a mesh or construct to OPML 2.0 (stdout, or -o ) neurowire opml export --mesh ai-news.json > ai-news.opml neurowire opml export --construct daily.json -o daily.opml # Import an OPML file or URL into a mesh JSON neurowire opml import subscriptions.opml -o mesh.json --name "My Reader" ``` `opml export` needs `--mesh ` or `--construct `. `opml import` takes a file path or `http(s)` URL; the mesh name comes from `--name`, else the OPML `head/title`, else `"imported"`. See [the CLI guide](/guide/cli) and [Meshes](/concepts/meshes). --- --- url: /formats/html.md --- # HTML page `@neurowire/web` renders a feed, mesh, or construct into a self-contained HTML news page. The renderer lives in `packages/web/src/render.ts`. ::: info Not a core format HTML is deliberately **not** a core feed format. It is presentation, not a serializer, so it lives in `@neurowire/web` (`toHtml`) rather than in `@neurowire/core`. This keeps core format-pure and dependency-light. HTML is not part of `serialize()` and the API does not serve it. ::: | Function | Output | |---|---| | `toHtml(feed)` | One self-contained HTML page for a feed. | | `toConstructHtml(construct, options?)` | A construct overview page (one recap card per mesh). | | `toConstructPages(construct)` | A full set: an `index.html` overview plus one page per mesh. | ```ts import { toHtml, toConstructHtml, toConstructPages } from '@neurowire/web' const page = toHtml(feed) const overview = toConstructHtml(fetchedConstruct) const pages = toConstructPages(fetchedConstruct) // [{ filename, html }, ...] ``` ## The self-contained invariant Every page is a single `.html` file that works offline and in CI: * **No external requests.** All CSS is inline in a `